Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.
Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
- HOST_INTERPRETER_KILL_REJECTION is one constant in cron/lifecycle_guard;
terminal_tool_guards and the execute_code lifecycle guard both use it,
so an image-name kill inside a cell now names the proc_* / explicit-PID
route instead of the generic "cannot restart or stop the gateway" text.
One invariant test with the generic path as control.
- gateway/restart.is_supervised_gateway_launch: the callee already maps
None to os.environ; pass environ straight through.
execute_code ran the same unbounded _is_supervised_gateway_process() probe
ahead of every cell, so the wedge #111922 bounds in terminal_tool still hung
an execute_code call (and its cron slot) forever: share the cell's deadline
and fail closed with a retryable error when the probe renders no verdict.
Moving the terminal pre-exec guard onto a deadline worker made it blind to
/stop, which keys on the tool thread's ident: record the acting-for tid in a
contextvar (copied into the worker by run_bounded_sync) so is_interrupted()
on the worker honours the tool thread's bit too.
Floor the guard's share of the deadline at 30s so a short command timeout
does not turn the guard's own cold-start cost (imports, git probes under
load) into a refusal — tests/tools/test_terminal_error_redaction.py was red
on the branch for exactly that.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
Four copies of the 40/60 head/tail algorithm with a near-identical notice
(terminal_tool_result, mcp_tool_content, code_execution_tool,
environments/base_output) collapse into tools/tool_output_truncate.py, so the
ratio and the `... [<LABEL> TRUNCATED - N <unit> omitted out of T total] ...`
marker are defined once. execute_code keeps byte mode + spill path and only
shares the notice/split. Visible change: the terminal notice now uses
thousands separators like the other three (`9,000 chars` not `9000 chars`).
kanban_specify._truncate: comment claimed escape stripping the body never did;
comment now says what the plain clamp is for.
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.
tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:
- tools/process_registry.py, tools/environments/{modal,singularity}.py:
`_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
call time (same seam as `tools/skills_tool._skills_dir`, so the existing
`monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
the `_X_resolved` flags and the lifecycle reset keep their shape.
tools/file_tools.py drops its private `file_read_max_chars` memo and reads
the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
startup) is ignored in favour of the routed profile's config.yaml. Both
sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
the static schema text is profile-neutral and `dynamic_schema_overrides=`
rebuilds the `display_hermes_home()` / create-dir hint per
`get_definitions()`, so a routed profile's model sees its own paths.
Refs #95685.
Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
Slim adaptation of #83772 to the current schema and failure-hint table.
Generated helpers are module exports on every execution path, not globals.
Correct schema, recovery hints and CLI tip rather than injecting names or
changing the execution boundary. Two registry-driven invariants reproduce
both misleading instructions on main and execute the corrected guidance.
Additional tool fix discovered during campaign #104904.
Original diagnosis and correction: @yuzilongleif-collab (#83772).
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
The prompt backend-probe and execute_code each kept a private (key, default) table for
container_config and drifted from the terminal tool's: the probe omitted docker_network, so
under docker_network: false it started a bridge-networked container; execute_code omitted
docker_extra_args, docker_forward_env and docker_env, so a sandbox created from that path
lost the operator's settings. Both now call terminal_tool_backends._container_config_from_config.
Salvaged from #62023 by @sagitario-jpn (the probe docker_network fix), widened to the whole class.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
* fix(execute_code): limits line teaches spillover — big stdout is saved, not lost (follow-up to #97043)
* fix(execute_code): drop the editorializing tail from the limits line (maintainer review)
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)
* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens
* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim
* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)
* fix(terminal): background param referenced the unadvertised legacy arg name
* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape
* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
Per-site decisions:
1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
via a function-local import; keeps the swallow-everything fail-open
contract and the (proc) -> None signature (agent/shell_hooks.py imports
it by name; _kill_git_process_tree alias preserved). The old body is kept
verbatim as _legacy_kill_process_tree and used as fallback when the
delegation import/call fails. A final proc.kill() is retained on the
happy path so Popen bookkeeping sees the exit (matches old behavior).
2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
(delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
body sent SIGTERM then SIGKILL with zero grace between them; the shared
primitive sends SIGKILL only. With no grace period the observable effect
is identical, and the psutil descendant sweep now also reaches
agent-browser's setsid'd daemon grandchild, which killpg alone missed.
tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
the legacy fallback (its assertions describe the fallback's internals).
3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
kill when escalate=True) — expressed as two delegated calls:
kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
body terminated children before the parent; the shared primitive
signals the group atomically (child is a session leader via
start_new_session=True) plus an identity-aware descendant sweep —
strictly wider coverage, same signals.
4. gateway/status.py: KEPT BOTH SITES.
- terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
incompatible with the shared primitive — it must RAISE OSError with
taskkill's stderr on non-zero exit (callers branch on that), falls back
to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
fail-soft primitive would invert the error contract.
- reap_gateway_children (~l2029): NOT migrated. It operates on a
pre-snapshotted child list from a parent that is already dead
(psutil.Process(pid) on the parent would fail), and every signal is
wrapped in identity/ownership checks the primitive lacks: is_running()
identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
The coupling is the feature; migrating would delete the safety logic.
5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
Dev tooling that intentionally kills by CAPTURED pgid because the direct
child is usually already reaped (psutil/pid-based primitive cannot find
it), and it avoids the psutil import on the test-runner hot path. Its
docstring already documents why psutil is the wrong tool there.
New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
execute_code lacked the lifecycle guard entirely, and Python argv-list
forms (subprocess.run([...])) separated command words with brackets and
commas the shell-shaped pattern could not see. Mirror the terminal_tool
guard in execute_code (ownership-gated per #92560) and strip argv-list
punctuation in the token-join re-scan. Salvaged from PR #68289 by
@arcimun, adapted to the ownership gate and current guard structure.
The PYTHONPATH/PATH sanitization suite was written POSIX-centric and
failed on real Windows 11 (reproduced natively: 4 failures before this
change). Fix the tests to express the true per-platform contract:
- test_other_major_version_site_packages_preserved /
test_make_run_env_injects_hermes_bin_dir: build inputs with
os.pathsep instead of hardcoded ':'.
- test_make_run_env_appends_homebrew_on_minimal_path: split on
os.pathsep, neutralise Git Bash dir prepending, and assert the
documented Windows passthrough (_append_missing_sane_path_entries is
a no-op off POSIX) instead of the Homebrew append.
- test_make_run_env_real_launchd_path_gains_homebrew: mark
macos_only per repo OS-marker policy (the regression is the macOS
launchd PATH; the merge is a passthrough on Windows).
- test_configured_home_alias_matches_launcher_output: create the
configured-home link via a helper that falls back to an unprivileged
directory junction (cmd /c mklink /J) when symlink creation raises
WinError 1314, and skips with a clear reason if no mechanism exists.
Also correct a stale comment in execute_code: the child is not always
the same Python as Hermes (project mode can select an external venv),
so the strip is about compatibility, not redundancy.
Adversarial review of the previous two commits (and #78917 itself)
found three ownership-boundary issues; this commit addresses them:
1. Repo direct-child over-strip (Finding A)
No launcher injects <repo>/tools or another direct child as an
independent PYTHONPATH entry - audited all four producers (Electron
electron-main.mjs, gateway/run.py::_ensure_windows_gateway_venv_imports,
cron/scheduler.py::_windows_cron_python_invocation,
tui_gateway/host_supervisor.py). The depth<=1 rule deleted user paths
that merely live under the repo directory; only the EXACT repo root is
now stripped.
2. Windows junction/symlink alias (Finding B)
The gateway launcher renders Hermes-owned paths under the configured
HERMES_HOME spelling (gateway_windows.py::_preserve_hermes_home_path),
which may be a junction to another drive, so it differs lexically from
the resolved repo root. _hermes_repo_root_aliases now carries both the
resolved and unresolved spellings; both are recognized as Hermes-owned.
3. Stale abstraction rename (Phase 4)
_strip_mismatched_site_packages -> _strip_hermes_owned_pythonpath:
the cross-version heuristic is gone, so the old name misdescribes the
behavior (ownership-based, not version-based).
Tests: direct-child now preserved; junction alias stripped (lexical pair
monkeypatched); Windows-only real-semantics test added (POSIX test remains
a safety test); mixed-ordering, duplicate-Hermes, and no-scrub PYTHONHOME
contract tests added. Full file: 52 passed / 16 failed (identical failure
set to base, all isolation-venv environment issues).
The Desktop Electron process injects the Hermes venv's site-packages path
(e.g. .../python3.11/site-packages) into PYTHONPATH so the Python 3.11
backend can import its packages. When this PYTHONPATH leaks into terminal
subprocesses running a different Python version (e.g. Python 3.13), 3.11
C extension modules appear on sys.path ahead of the correct 3.13 versions
and crash with ImportError (PIL _imaging, cryptography, etc.).
Replace the existing blunt pop of PYTHONPATH from _ACTIVE_VENV_MARKER_VARS
with a surgical Hermes-venv-aware filter:
- Parse each PYTHONPATH entry by path
- Strip only paths under ~/.hermes/hermes-agent/venv/.../site-packages
- Preserve the Hermes source root (needed for import hermes_cli)
- Preserve all user-set PYTHONPATH entries
The same filter is applied in all three env builders:
- _make_run_env (foreground terminal commands)
- _sanitize_subprocess_env (background/PTY spawns)
- PTY env builder
This preserves env_passthrough semantics and never silently discards the
user's own PYTHONPATH configuration.
/simplify-code findings on the full PR diff:
- _is_usable_python had the same sticky-failure bug the previous commit
fixed in _python_environment_prefix: lru_cache pinned a transient
probe failure (fork pressure, timeout) as False forever, silently
locking project mode to sys.executable. Both probes now share a
success-only bounded dict cache via _cache_probe_result() with FIFO
eviction at _PROBE_CACHE_MAX (the old < cap guard stopped caching new
entries instead of evicting, re-probing entry 33+ on every call).
- The hermes-root-omitted logger.info fired on every external-env call
in project mode; now deduped once per interpreter path per process
(matching the tirith/mcp warn-once convention).
- Regression test: _is_usable_python probe failures are retried, not
cached (mutation-verified).
Follow-up to the salvaged #81201 commits:
- Short-circuit _uses_hermes_python_environment when the child IS the
running interpreter (path or realpath match). The default strict-mode
path no longer spawns a probe subprocess at all, and a flaky probe of
sys.executable can never drop the hermes root from PYTHONPATH
(protects the test_repo_root_modules_are_importable invariant). The
realpath leg also covers uv-style venvs whose bin/python resolves to
the same binary.
- Stop caching failed probes: _python_environment_prefix now uses a
success-only dict cache instead of lru_cache, so one transient
timeout under load no longer sticks for the process lifetime.
- Deduplicate the subprocess probe scaffolding shared with
_is_usable_python into _probe_python().
- Log once when the hermes root is omitted so import-behavior changes
are diagnosable from user reports.
- Tests: fail the composition tests loudly if execute_code never
reaches Popen (was vacuously passing on exceptions); assert the
staging dir is literally first in PYTHONPATH (was truthiness only);
add guards for probe-failure retry and the no-probe short-circuit.
Review follow-up on the salvaged handler: a non-string 'code' (int,
dict, list) reached code.strip() and surfaced as a generic
'Tool execution failed: AttributeError' — the same unrecoverable shape
the salvage exists to eliminate. Add an isinstance guard beside the
'command' check that names the received type and shows the correct
call form; narrow the docstring to what the handler actually does.
Regression test drives int/dict/list through registry.dispatch and
asserts no AttributeError leaks (mutation-checked: removing the guard
fails 3 subtests).
Two bugs reported on the docker terminal backend (desktop app, sandboxed
profiles with container_persistent: false):
1. A NEW chat's container inherited the PREVIOUS session's workspace,
bind-mounted rw at /workspace, because the mount source was the
process-global TERMINAL_CWD env var (written by the workspace picker,
outliving its session) and all sessions shared one 'default' container.
2. Every command failed with exit 126 because the desktop gateway recorded
the HOST launch directory as the session cwd, and each command was
prefixed with 'cd /Users/<user>/...' inside the container.
Fixes (class-wide, single owners):
- container_persistent: false + docker now keys containers PER SESSION:
fresh container per chat, removed at session close/idle. delegate_task
children share the parent's container via an explicit alias registry.
container_persistent: true keeps the documented ONE-long-lived-container
contract unchanged.
- _resolve_task_host_cwd() is the single owner of the cwd->/workspace mount
policy across all four env-creation sites; under isolation it refuses
process-global cwd sources and mounts only the session's own attached
workspace (tui_gateway now tags overrides with cwd_source).
- _resolve_command_cwd() gains the same host-path guard the env-creation
sites already had (#50636/#54447 sibling site): a recorded host cwd is
discarded on container backends instead of cd-ing every command into a
nonexistent path.
E2E-tested against real Docker: distinct containers per session, no stale
mount in a fresh session, no exit 126 from host cwd records, containers
removed at session teardown.
json_parse used json.loads(strict=False), which relaxes control
characters but rejects a leading UTF-8 BOM (U+FEFF). Windows CLI
tools and some files prepend a BOM, causing JSONDecodeError on
otherwise valid JSON output.
Strip a leading BOM before calling json.loads when the input is
a string with a U+FEFF prefix.
Original PR by @woxinwuhen713-bit (#57870).
Co-Authored-By: Claude <noreply@anthropic.com>
Every tool schema ships on every API call. The terminal schema was
5,641 chars (~1,410 tokens) and execute_code 2,842 (~710) — the two
largest core tools, padded with repeated war stories and triple-stated
rules. Schema token audit across 88 tools: ~33k tokens total.
This trims prose while preserving every hard rule (each still stated
exactly once):
- terminal description 2,324 -> 1,233 chars: tool-redirect lines
collapsed to one sentence; background/notify guidance deduplicated
(was stated in desc + 2 params); PTY/pager rules merged.
- background/notify_on_complete/watch_patterns params 692/508/1,114 ->
~330/250/490 chars: kept the mutual-exclusion contracts, the
rate-limit consequence, and the bounded-vs-long-lived distinction;
dropped narrative repetition.
- execute_code description tightened (helper docs inlined to one line
each; when-to-use kept).
Net: terminal schema 5,641 -> 3,386 chars, execute_code 2,842 -> 2,522
— ~700 tokens saved on EVERY request with the terminal+code toolsets.
One test updated (pinned a removed phrase; now pins the rule's new
phrasing).
The top execute_code failure shapes in production (state.db mining)
are sandbox-contract confusions, not logic bugs: importing tools that
aren't in the sandbox from hermes_tools (23x in one window, incl.
importing the built-in helpers json_parse/shell_quote/retry), importing
third-party packages absent from the sandbox interpreter (matplotlib
6x), and indexing tool-result dicts as strings. The stderr traceback
alone sends models into re-diagnosis loops.
Failed scripts (exit != 0) now carry one actionable 'hint' field:
- unavailable hermes_tools import -> lists the tools that ARE
importable in this session + points to normal tool calls otherwise;
- built-in helper import -> 'no import needed, call it directly';
- ModuleNotFoundError -> 'sandbox has stdlib only; use terminal() with
the project venv for third-party packages';
- string-indexing errors -> 'tool functions return dicts, do not
json.loads them'.
Bounded 4KB stderr scan, first match wins, never raises; successful
scripts and unknown failures are untouched.
Production mining (state.db, 28.5k read_file calls in the recent
window) shows 74.3% of reads truncated — nearly all by the 500-line
default, not the char budget (22,443 line-limit vs 13 char-budget
truncations). That churn produced 12,229 redundant re-reads, and after
a truncated result the most common next move was fleeing to terminal
cat/sed (4,608 times) — the pagination contract was not trusted.
Median truncated file is 2,422 total lines, so a 2000-line default
makes 44% of today's truncated first-reads complete in one call while
the unchanged ~100K-char budget still caps worst-case result size
(same ceiling as before: 500 lines x 2000-char line cap = the same
100K). Schema max was already 2000.
Touchpoints: DEFAULT_READ_LIMIT + both read_file signatures
(file_operations.py), read_file_tool + schema text (file_tools.py),
execute_code sandbox stub docs (code_execution_tool.py), 3 tests
pinning the old default.