execute_code ran the same unbounded _is_supervised_gateway_process() probe
ahead of every cell, so the wedge #111922 bounds in terminal_tool still hung
an execute_code call (and its cron slot) forever: share the cell's deadline
and fail closed with a retryable error when the probe renders no verdict.
Moving the terminal pre-exec guard onto a deadline worker made it blind to
/stop, which keys on the tool thread's ident: record the acting-for tid in a
contextvar (copied into the worker by run_bounded_sync) so is_interrupted()
on the worker honours the tool thread's bit too.
Floor the guard's share of the deadline at 30s so a short command timeout
does not turn the guard's own cold-start cost (imports, git probes under
load) into a refusal — tests/tools/test_terminal_error_redaction.py was red
on the branch for exactly that.
The salvaged commit put `_pre_exec_block` behind the command's
`run_bounded_sync` deadline but let a timed-out guard fall through into
execution. The gateway-lifecycle, dangerous-workdir and self-repo checks
apply unconditionally (`force=True` cannot bypass them), so a guard that
never rendered a verdict must not let the command run unguarded: return
the terminal error envelope (`status: error`, "did not finish ... Retry
the call") instead, mirroring how the bounded `env.execute` path reports
its own expiry as a result rather than continuing.
Tests trimmed to the two invariants: a wedged guard returns a bounded
error without executing; a completed guard keeps its verdict (pass ->
execution, rejection -> its own blocked result).
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
Port from Kilo-Org/kilocode#13224: fixed waits belong in foreground terminal calls, while background mode is reserved for independently running processes.
A multiplexed dashboard can call _ensure_terminal_env_bridged while a
secondary profile's HERMES_HOME override is active. The one-shot latch then
wrote that profile's docker policy into process-global os.environ and poisoned
later unscoped launch-profile tool calls (#107422).
Skip the ambient bridge whenever a context-local home override is set —
ambient env is launch-profile authority only; routed profiles must use
terminal_scope (same rule as env_loader._reapply_terminal_config_bridge).
(cherry picked from commit 2050efb24fdc9d54a282b24d0042b90f47486c5a)
A message typed while the agent runs (CLI busy_input_mode=interrupt, gateway
priority redirect, ACP redirect) goes through AIAgent.redirect(). During tool
execution redirect() degrades to steer(), whose delivery rides the tool result
— so a long foreground command (a `sleep 285` CI poller, a build) parked the
user's message until it exited. The UI printed "Redirected current turn" while
nothing happened for minutes.
redirect() now also asks the tool workers to YIELD (tools/interrupt.request_yield).
The local terminal backend's wait loop honours it: the drain thread is stopped, the
still-running Popen is adopted by the process registry as a notify_on_complete
background session (ProcessRegistry.adopt_local — output so far seeds the buffer,
the registry reader continues from the pipe), and the tool returns immediately with
status "yielded_to_background" + session_id. The command is never killed; the
completion notification arrives as usual and process(poll/wait/log/kill) work on it.
Non-local backends and internal env.execute() consumers pass no yield_handler and
are unaffected; a stale yield bit is cleared with the interrupt bit per worker tid.
DockerEnvironment.cleanup() runs docker stop + docker rm -f on a daemon thread and
keeps the handle on the env. The idle reaper pops the env out of _active_environments
BEFORE calling cleanup(), and the atexit drain only iterated that registry — so a
detached env's worker was unreachable and died with the interpreter, leaving a
stopped (or, if the exit came fast enough, still running) labeled container behind
while the log said the environment was cleaned.
Every teardown worker is now also recorded in a module-level set in docker.py;
_atexit_cleanup joins that set after the registry pass via
DockerEnvironment.wait_for_all_teardowns (re-snapshotting each pass, since the
reaper can start a worker while the drain runs). Finished workers are dropped from
the set on each drain so it cannot grow across a long gateway life.
Mechanism from #86344 by @PRATHAMESH75; this is the slim redo — one shared set
plus one static drain hooked into the existing terminal_tool atexit, instead of a
second atexit registration in docker.py.
On hosts where Docker ships as a snap (Ubuntu cloud images / Azure VMs), the
snap's AppArmor confinement turns two sandbox hardening flags into a dead
container at start: `--init` fails with "exec /sbin/docker-init: operation not
permitted" and `--security-opt no-new-privileges` then fails every exec the
same way ("exec /usr/bin/sleep: operation not permitted"). This is snapd
LP#1908448 — not probeable from the client, and docker_extra_args cannot remove
flags we add.
`terminal.docker_snap_compat: true` drops exactly those two flags; cap-drop ALL,
the tmpfs hardening, PID limits and the privdrop caps are unchanged, and a
warning is logged at container start. Bridged everywhere the other docker_*
keys are (CLI env map, gateway env map, `hermes config set` sync, terminal_tool
env read, the shared container_config shaper, DEFAULT_CONFIG).
_HOST_CWD_PREFIXES only listed C:\ and C:/, so a D:\ or E:/ working directory on a
Windows host reached docker run -w and the container started in a directory that does
not exist. The two prefix scans now go through one _is_host_cwd predicate that matches
any drive letter with either slash.
Independent review: an over-cap timeout on `cmd &` was promoted, so the
tracked shell exited at once while the payload ran untracked (the exact
thing the '&' guidance exists to prevent); and the note promised a
notification even on finite sessions where async delivery is disabled and
the spawn had already cleared notify_on_complete. The detachment guidance
now runs before the promotion decision regardless of timeout, and the note
reads the spawn's actual notify_on_complete: notification wording when kept,
poll-only wording when the session cannot receive one.
Tests (2 new): '&' and nohup with an over-cap timeout still refuse; the note
matches the delivery capability.
`Foreground timeout Ns exceeds the maximum of 600s` was the second most
frequent tool-layer error in the 1,393-agent refactor run: 454 refusals
(285 asked 900 s, 41 asked 620, 20 asked 1800), 251 of them test-suite
invocations. Every retry was mechanical: lower the timeout (167), split into
sleeps (82) or re-send with background=true (86). The schema invited it
("set high for long tasks", max 600) while the suites take 10-30 min.
An over-cap foreground timeout is a bounded job the caller wants to wait
for. terminal_tool now runs it as a tracked background process with
notify_on_complete=true and says so in the result
(`promoted_from_foreground`, naming the requested and cap seconds, with
"do NOT re-run it"); the schema text describes the new behaviour. The
`&`/nohup/setsid and long-lived-server guidance stays a refusal: those need
the command itself rewritten, which the tool cannot do safely.
Live (real local backend): main refuses `sleep 1; echo LIVE_OK` at
timeout=900; branch returns a proc_* session with notify_on_complete and
the note, and the command runs once.
Tests: the promotion test runs a real command through the real registry and
asserts the result shape plus that it executed exactly once in the
background; `&` still refuses; schema text updated. Terminal/process suites
(27 files) 378 passed.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
The referenced-script walk in cron/lifecycle_guard.py capped each file
(1 MiB) and the recursion depth (8) but not the walk: a command referencing
hundreds of scripts, or one enormous shlex token, held the GIL for minutes
on every gateway terminal call (#78398).
Add a per-walk _LifecycleScanBudget (bytes, lines, longest line, unique
paths, remote reads) charged BEFORE any text reaches shlex, and cap each
referenced read at the remaining byte budget so an oversized file is never
read whole. Exhaustion fails closed (the existing contract for one oversized
file) and is logged at WARNING so operators can tell it from a genuine
lifecycle block. Limits are sized so real wrapper graphs never hit them:
a 200-script benign graph is allowed and a restart hidden behind it is
still caught.
tools/terminal_tool.py gates its optional launchctl pre-scan (which also
tokenizes) on the same budget; the full guard still runs afterwards.
Redesigned from #83821 by @Riccardo-Vecchi, which introduced the budget
idea but blocked benign wide graphs (64-path cap) and bundled a suffix
classification change that is left out here.
Refs #78398
Follow-up on the #42278 salvage (#98222): `a && b & &>/dev/null c` is valid
bash but rewrote to `a && { b & } &>/dev/null c`, which binds the redirect to
the brace group and orphans `c`. Treat a suffix starting with `&>` as command
text and insert the `;` separator. Adds bash -n coverage for that shape, for
the `;;` case-arm terminator (must NOT gain a separator), a multi-line mixed
form, and separator idempotence.
`_rewrite_compound_background` rewrites `A && B &` into `A && { B & }` to
avoid the subshell-wait process leak. When another statement follows the
backgrounded compound on the same line (`A && B & C`), the trailing `&`
was the only separator between the compound and `C`. The rewrite consumed
that `&` into the brace group and produced `A && { B & } C`. A brace group
must be terminated by `;`, `&`, `|`, a newline, or `)`/`}` before the next
command, so the result is a bash syntax error and the entire command fails
to run — neither `A`, `B`, nor `C` execute. The rewrite is applied by
default to every foreground command, so a valid command is silently
mangled into one that errors out.
This restores a `;` separator after the closing `}` when the suffix
resumes with command text. Only spaces and tabs are stripped before the
check; a newline already terminates the group, and an existing separator
(`;`, `&`, `|`, `)`, `}`) is left untouched, so previously-correct rewrites
are unchanged.
## What does this PR do?
Fixes a rewrite in `_rewrite_compound_background` that turned a valid
foreground command of the form `A && B & C` into the bash syntax error
`A && { B & } C`, causing the whole command to fail. A `;` is now inserted
after the brace group whenever a further statement follows on the same
line, while leaving commands that already end in a separator/newline
untouched.
## Related Issue
N/A
## Type of Change
- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- [ ] ✨ New feature (non-breaking change that adds functionality)
- [ ] 🔒 Security fix
- [ ] 📝 Documentation update
- [ ] ✅ Tests (adding or improving test coverage)
- [ ] ♻️ Refactor (no behavior change)
- [ ] 🎯 New skill (bundled or hub)
## Changes Made
- `tools/terminal_tool.py`: in `_rewrite_compound_background`, insert a `;`
after the rewritten `{ ... & }` brace group when the trailing suffix
begins with command text rather than a separator/terminator.
- `tests/tools/test_terminal_compound_background.py`: add
`TestTrailingStatementSeparator` for the string-level rewrites and
`TestRewriteIsValidBash`, which runs the rewriter output through
`bash -n` for parse validity plus one end-to-end execution check.
## How to Test
1. Run `pytest tests/tools/test_terminal_compound_background.py -q`.
2. Before the fix, `_rewrite_compound_background("echo hi && sleep 5 & echo done")`
returns `echo hi && { sleep 5 & } echo done`; `bash -n -c` on that string
exits 2 with `syntax error near unexpected token 'echo'`.
3. After the fix it returns `echo hi && { sleep 5 & } ; echo done`, which
parses and runs, and the trailing statement executes.
## Checklist
### Code
- [x] I've read the Contributing Guide
- [x] My commit messages follow Conventional Commits
- [x] I searched for existing PRs to make sure this isn't a duplicate
- [x] My PR contains only changes related to this fix
- [x] I've run `pytest tests/tools/test_terminal_compound_background.py -q` (50 passed)
- [x] I've added tests for my changes
- [x] I've tested on my platform: macOS 15.5
### Documentation & Housekeeping
- [x] I've updated relevant documentation — N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` — N/A
- [x] I've considered cross-platform impact: the rewrite is
platform-independent; the `bash -n` test is skipped when bash is absent
- [x] I've updated tool descriptions/schemas — N/A
Extract the signature-checking cleanup dispatch that cleanup_vm already had
into tools.terminal_tool._cleanup_env and call it from the prompt-time
backend probe instead of carrying a second copy. Comment on the probe
trimmed to the non-obvious part (why ssh is skipped).
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
The parent-chat suppression gate (afee35700e) keyed on evt task_id
starting with 'sa-'. But terminal_tool stamps ProcessSession.task_id
with the COLLAPSED container key from _resolve_container_task_id()
('default' or the session key — subagents intentionally share the
parent's container), so real child-spawned background processes carried
task_id='default' and their completion/watch notifications walked
straight past the gate into the parent conversation.
Fix: ProcessSession gains owner_task_id (the RAW spawning task id),
stamped by both spawn paths (spawn_local/spawn_via_env) from
terminal_tool's raw task_id, carried on every queued event
(completion, watch_match, watch_disabled, overflow), round-tripped
through the crash checkpoint, and used by both the drain suppression
gate and the attribution formatter (task_id remains the fallback so
synthetic/legacy events keep working).
Live repro: on origin/main a simulated subagent completion event with
the collapsed key was delivered to the parent drain (leak); on this
branch it is suppressed, parent-owned events still deliver, and
surface_child_process_notifications=true restores delivery with
attribution. 4 new regression tests fail on origin/main, pass here.
* fix(terminal): stop claiming a Linux environment — point at the env section; near-neutral tokens
* feat(terminal): default GIT_PAGER/PAGER=cat in session env; drop schema lines the runtime already enforces; fix pty backend claim
* refactor(terminal): unify notify_on_complete+watch_patterns into notify (bool|list); trim pipe-masking prose (runtime hint owns it)
* fix(terminal): background param referenced the unadvertised legacy arg name
* refactor(terminal): background-only modifiers (pty, notify) fail loud on foreground calls with corrected shape
* fix(execute_code): block the new notify arg in the sandbox terminal stub (foreground-only)
On macOS with TCC (Transparency, Consent, and Control), os.getcwd()
raises PermissionError: [Errno 1] Operation not permitted — not
FileNotFoundError — when the process CWD is under a protected location
(~/Documents, ~/Desktop, ~/Downloads) and the calling process lacks
Full Disk Access.
_safe_getcwd() only caught FileNotFoundError (deleted CWD), so the
terminal-tool cleanup thread, which calls _get_env_config() →
_safe_getcwd() every 60 s, logged a full stack trace on every tick.
This accumulated hundreds of MB of noise in mcp-stderr.log (observed
184 MB on a single-day session) without breaking functionality — the
cleanup thread's outer try/except swallowed the exception, but
exc_info=True kept emitting the traceback.
Fix: add PermissionError to the existing except clause so the fallback
chain (TERMINAL_CWD → $HOME) runs, matching the existing pattern for
deleted-CWD recovery (#17558). Complements #66306, which handles
PermissionError from subprocess.Popen(cwd=...) for an inaccessible
configured cwd on Linux; this handles the distinct case where the
live process CWD itself is TCC-blocked.
Tests cover: PermissionError fallback to $HOME, TERMINAL_CWD priority,
FileNotFoundError regression, happy path unchanged, and unrelated
OSError (NotADirectoryError) still propagating instead of being
swallowed.
Follow-up on @fangliquanflq's opt-in (#84775): after the profile-scoping fix
(#94560) the container cache key is resolved in _resolve_container_task_id,
so the shared key must unify profiles there too — 'shared:<key>' for every
session of every opted-in profile AND for CLI/no-session runs. Delivery adds
the shared sandbox layout as the first translation candidate. Empty key
keeps strict per-profile isolation; SSH ignores the key entirely.
Commit a270c4ade's session-key fallback in _resolve_container_task_id was
added to stop cross-profile SSH environment reuse, but it wasn't backend-
gated: persistent Docker silently fragmented into one container per gateway
session, breaking the product contract (one long-lived container per profile,
shared by CLI and every session of that profile). #93950's vanishing MEDIA
attachments were downstream damage.
- persistent Docker (container_persistent: true) now keys to the profile:
literal 'default' for the default profile (same container as CLI),
'profile:<name>' for named profiles
- SSH and non-persistent Docker keep session scoping (the original leak fix
and the #82731 isolation contract are untouched)
- gateway MEDIA translation follows the profile layout and keeps the legacy
bug-window per-session sandboxes as fallback candidates, trying each until
the file resolves — old sessions self-heal, no migration
- /root/.hermes credential-surface refusal preserved across all layouts
Follow-ups on top of the cherry-picked #93532 cap:
- Regression tests: suppressed (in-cooldown) matches must NOT consume the
lifetime budget; the cap trips exactly at the Nth DELIVERED match and
promotes to notify_on_complete with the watch_disabled summary queued
right after the final match.
- Extract _emit_lifetime_watch_disabled() and emit the summary even when
the global breaker drops the final match, so the user always learns why
watching went quiet (parity with the strike-limit path).
- Mention the lifetime cap in the terminal tool docstring (the schema text
was already updated by #93532).
Refs #93513
Per-session rate limiting only counts consecutive strike windows, so a
pattern that recurs at a cadence just above WATCH_MIN_INTERVAL_SECONDS
(e.g. a service restarted repeatedly over a day) never trips the
existing strike-limit disable — each match lands in its own clean
cooldown window. Every one of those matches still forces a full-context
agent turn, which stalls the event loop on large sessions (#93513).
Add WATCH_LIFETIME_MAX_HITS: once a session has delivered this many
watch_match notifications over its whole life, disable watch_patterns
and fall back to notify_on_complete, reusing the existing disable path.
delegate_task children run on worker threads of the parent process and
inherit the process-wide HERMES_INTERACTIVE=1 the CLI sets at startup.
_transform_sudo_command's interactive gate therefore fired inside
children with no sudo callback registered, falling through to the raw
/dev/tty password prompt: a password box printed mid-TUI from a
background thread, parallel children racing for the tty, and each child
blocked for the full 45s timeout.
Gate the prompt (and the sibling 'you will be prompted again' message
after an auth failure) on agent.delegation_context.is_delegated_child_context(),
the ContextVar set around every child run and propagated through
contextvars.copy_context onto the executor thread. Children now behave
as headless for sudo: configured SUDO_PASSWORD, the session cache, and
the NOPASSWD probe still work; otherwise the command fails gracefully
with a subagent-specific tip.
A/B verified: 3 regression tests fail on merge-base, 7/7 pass at head.
The terminal tool lifecycle guard and the gateway stop/restart CLI
guards keyed on the raw _HERMES_GATEWAY=1 env marker, which every
gateway descendant inherits (and importing gateway.run sets it too).
CLI/TUI agent sessions were falsely blocked from documented gateway
management commands. Gate on _is_supervised_gateway_process() instead,
which requires owning the live gateway PID file.
Salvaged from PR #92196 (guard half) by @nbxuhk. Fixes#92560.
Follow-up to the session-scoping fix: _get_sudo_password_cache_scope()
and _resolve_container_task_id() carried byte-identical copies of the
HERMES_SESSION_KEY lookup (contextvar + os.environ fallback). Collapse
both onto one helper adopting the bare-import convention approval.py
already uses — get_session_env() implements the fallback internally, so
the old try/except could only fire on import failure, where silently
degrading to process-global semantics would reintroduce exactly the
cross-session contamination the fix prevents.