A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).
- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
`served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
action env already funnel through (the dashboard site now calls it for its target).
Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
the profile is not the one the updater runs as; the module stays stdlib-only at
import time.
Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.
Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
The filename walk and the grep fallback each hand-built the same
`-type d -name .* ! -path <root> -prune` clause; the first fold drifted
between them, which is how #18473 stayed open. One `_hidden_prune_expr`
serves both; docstring says when the pruned branch is used and why.
Gate review on the previous commit: `find <root> -mindepth 1` yields nothing
when the root is a file, so `search_files(path="~/.hermes/config.yaml")` on
the grep fallback returned 0 matches (main returned 1). Reuse the
`! -path <root>` exemption the filename walk in `_search_files` already uses:
a dot-named root — directory or single file — is never pruned, hidden
descendants still are.
The root predicate now goes through `_normalized_filename_search_root`, which
already handles MSYS paths on Windows and posixpath for remote environments
instead of applying host `os.path` to a backend path.
The grep fallback passed --exclude-dir='.*' and the search root together. grep
applies --exclude-dir to the command-line root as well as to subdirectories
(GNU grep to every component of it), so any content search rooted at or under a
dot-directory such as ~/.hermes returned zero matches when ripgrep was absent,
while the same search returned results with rg. Reported in #18473 with a live
12-vs-0 comparison on GNU grep 3.7; BSD grep drops the root only when the root
itself is dot-named.
Route those roots through the existing find -prune pipeline, which already
mirrors rg's hidden-directory default for descendants, and give find -mindepth 1
so the prune cannot swallow a dot-named root. The macOS protected-folder branch
shares the pipeline, so the prune clause is now built only when there are
protected paths to prune.
Credit: @memosr (#34017) and @luxles (#48878) proposed dropping --exclude-dir for
hidden roots first; that variant leaks hidden descendants, which the find route
keeps excluded.
The shard fetch now routes through hub()._guarded_http_get (#114887), so a patch
on tools.skills_hub_skillssh._get_text intercepted nothing and the test reached
the live sitemap on CI. Patch tools.skills_hub._ssrf_safe_http_get like the
neighbouring guarded-GET test and model the dark shard as a 503.
A per-skill sitemap shard whose fetch failed (timeout/503 -> None) was read as
an empty shard and the resulting partial catalog was written to the shared
skills_sh_sitemap_v1 cache, so one slow shard silently dropped skills from the
published index for the whole TTL. Same class as the ClawHub page-walk fix in
this PR: retry the shard up to the shared CATALOG_PAGE_RETRIES bound, and if it
stays dark return the partial slice without caching it.
CATALOG_PAGE_RETRIES moves from ClawHubSource to the SkillSource base so both
catalog walks share one bound.
Trim the salvaged retry to the case the logs show: `_get_json` returning None
(timeout / non-200) is a hole in the cursor walk, not the end of the catalog.
A dict page with an empty item list stays terminal as before, and the
interactive browse path (12 s budget) no longer sleeps up to 8 s per retry.
Constant `CATALOG_PAGE_RETRIES` replaces the local literal.
Workflow: `skills-index.yml` `timeout-minutes` 50 -> 120. Since Sep 17 ClawHub
serves ~10 s per 200-item page (was ~2.5 s), so the ~400-page walk alone takes
60-70 min; the 06:28 UTC runs on Sep 17 and Sep 18 were cancelled at 50 min with
the clawhub crawl still running, and the 18:19 run stopped at 12,145 skills
(< 20,000 floor) after one failed page. Both halves are needed for the index to
ship again.
- HOST_INTERPRETER_KILL_REJECTION is one constant in cron/lifecycle_guard;
terminal_tool_guards and the execute_code lifecycle guard both use it,
so an image-name kill inside a cell now names the proc_* / explicit-PID
route instead of the generic "cannot restart or stop the gateway" text.
One invariant test with the generic path as control.
- gateway/restart.is_supervised_gateway_launch: the callee already maps
None to os.environ; pass environ straight through.
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.
The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).
Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.
Fixes#113667
A server that answers the legacy initialize with HTTP 200 but reports
2026-07-28 (regardless of the version offered) makes the SDK raise
'Unsupported protocol version from the server'. auto mode then falls back
to server/discover, which such a server rejects with a non-JSON 4xx, and
the connect died with 'both Streamable HTTP and SSE transports failed
(Streamable HTTP: Server returned an error response; SSE: 405)' even
though the same initialize/tools/list POSTs succeed from curl (#113359).
When both fail for that reason, re-run initialize ourselves, adopt the
result pinned to the version we offered (keeps later requests
legacy-shaped, which the handshake just proved the server accepts), send
notifications/initialized and proceed to tools/list.
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).
- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
ConnectionError; the no-fallback re-raise after a proven session, strict
redirect headers or a non-rejection) carry the detail. Debug-logs the
endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.
Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Review follow-up: swapping tools/mcp_oauth_manager.py back to origin/main kept every
PR test green, so the pre-flight PRM/ASM discovery stamp had no invariant. The existing
prefetch test now records the User-Agent of each mocked discovery request and asserts
it is DEFAULT_AUTH_REQUEST_USER_AGENT (red without the stamp, green with it).
hermes_cli/__init__.py is stdlib-only and defines __version__ at module top, so the
try/except around its import could never fire; import directly.
Widen the salvaged fixes to the whole class and add the pieces they missed:
- The default User-Agent for SDK-built OAuth requests moves from the manager's
bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
provider gets it too, and the manager's pre-flight metadata discovery (its own
client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
registration and token requests through client.send(), which never merges the
client's default headers; www.tradingview.com's WAF answers a header-less GET
with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
/register and /authorize on the MCP host, and the login died with
"Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
... != None" because no issuer was ever discovered).
- The callback listener now runs serve_forever() and is shut down before
server_close(). A thread parked in handle_request()'s select() keeps the
closed listening socket alive (the kernel holds the file for the duration
of the poll), so a flow cancelled mid-wait left the port bound and the
retry on the same pinned/cached port raised "OAuth callback port N is
already in use" with no external collider.
- When every authorization-server metadata fetch failed, a registration error
is re-raised leading with those statuses ("Could not read
authorization-server metadata (403 from ...); dynamic client registration
then fell back to a guessed endpoint on the MCP host and failed: ...").
humanize_oauth_registration_error leaves that message alone so the 403 in
it is not mistaken for a DCR allowlist refusal.
Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
The MCP SDK builds /.well-known discovery and dynamic client registration
requests as bare httpx.Request objects inside async_auth_flow, so the
client's default headers never apply and they leave with no User-Agent at
all. Some WAFs reject header-less requests outright: coda.io returns 403
on every discovery and registration call, which Hermes then misreports as
"'<server>' only allows pre-approved OAuth clients". The existing
oauth.user_agent option does not help because it is stamped only on
token-endpoint requests.
HermesMCPOAuthProvider's generator bridge now gives SDK-built requests a
default User-Agent ("Hermes-Agent") when the header is absent. The
caller's own MCP request is never touched, and an explicitly configured
oauth.user_agent on token requests still wins.
Reproduced against https://coda.io/apis/mcp: before, every well-known
and /register call returned 403; after, registration returns 200 and the
browser authorization flow proceeds.
Co-Authored-By: claude-flow <ruv@ruv.net>
A worker process that carries HERMES_DELEGATED_CHILD_CONTEXT next to its
HERMES_KANBAN_TASK is fenced by kanban_path_is_fenced's env-marker branch: both
auto-heartbeat writes raised PermissionError at DEBUG only, so the board showed a
worker that never beats while the process was alive (the salvaged commit already
made the return value honest, but _touch_activity ignores it). Log the refusal once
per process at WARNING with the cause, and skip the bridge for an in-process
delegate child BEFORE stamping the rate-limit window so a chatty child cannot starve
the worker's own heartbeat.
No self-write grant is added: the marker + TASK combination means "descendant, not
worker" by design (b578261584, 8a8c3634e8), and the only in-tree grant path,
kanban_db_dispatch._default_spawn, already pops the marker. The real-spawn grant
test now proves that from a dispatcher that itself carries the marker: the worker's
auto-heartbeat lands and its handoff completes.
Docs: worker guide notes the automatic claim extension and what the refusal warning
means.
Sibling of the previous commit: read_file_tool returned the unchanged-stub
before _record_read, and only _record_read calls
mark_background_review_skill_read. A skill file the parent had already read
left the fork with a stub and a refused skill_manage write. Same rule as
skill_view: no dedup stub when is_background_review().
Also: is_background_review hoisted to module level in skills_tool (no cycle —
skill_provenance imports only contextvars), comment reworded to the two real
reasons, test renamed to the shipped design, fixture seeds a fresh read-mark
store so the mark cannot leak between tests.
The review fork reuses the parent's session/task id for prefix-cache parity, so
its skill_view calls hit the parent's repeat-view dedup and got a
content_returned:false stub. The stub path never calls
mark_background_review_skill_read, so every skill_manage patch/write_file in
the fork was refused by the read-before-write guard: 408 refusals and zero
skill updates in one deployment, the same signature on three more (#95976).
A dedup stub is only valid while the referenced content is in the caller's own
context; the fork's context is not the parent's. Skip the dedup entirely when
is_background_review(): the fork's view is a real read (which marks the file
for the guard) and records nothing, so the parent's cache is untouched.
Salvaged from #95975 (@yaozhen), rebased onto the decomposed skills_tool and
narrowed from a namespaced bucket to no dedup in the fork.
Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).
- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
(rejections never reach the bridge; mentions ride the first payload only, stale bridge
fails closed); drop the fabricated legacy-sender branch that only existed to cover the
deleted probe. Mutation check: removing the first-payload gate turns the new test red.
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
`_try_dispatch_background_run` returned at the `async_delivery_supported()`
gate before reaching `_reap_stale_executions`, so on the one-shot
`hermes cron run` path (the CLI scopes `_SESSION_ASYNC_DELIVERY` to False)
the dead-owner reap never ran: an execution row left 'running' by a killed
prior run stayed that way until a scheduler tick — exactly the gap the
helper's own docstring says it exists to close.
Hoist the reap above the gate; the async dispatch path keeps its existing
reap-before-claim ordering.
Salvaged from #113938 (@kvnloo, source hunk verbatim; its test hunk is
replaced by an invariant test in a follow-up commit). #86981 (@pierrenode)
identified the same gap first and placed a duplicate reap in
`_execute_job_now`; superseded by this single hoist.
Part of #113923.
Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
The `--run-delivery` catch-all in tools/bot_mode_dm.py printed a structured
`{error, reason}` only for target_busy; any other exception went to stderr as
prose, so the sender's completion notification had no reason to branch on
(auth failure vs. transient) for the local lane while the relay lane had one.
The vocabulary-guarded helper the relay lane introduced moves beside the
vocabulary it guards (tools/bot_failure_reasons.delivery_failure_reason) and
both lanes call it: an exception's own `reason` is trusted only when it is
`target_busy` or in ALL_REASONS, otherwise the text is classified. The runner
now prints `{error, reason}` JSON on stdout for every exception (exit 1 kept).
Test: a classifiable, an unclassifiable and an out-of-vocabulary-reason
exception in --run-delivery each yield JSON with a vocabulary reason.
Review finding on #114941: the CLI-runner branch of `_spawn_delivery` still
answered `status: sent` with only a `process_id`, while the live-owner branch
answers `queued`/`claimed` + `delivery_id` and the relay branch `queued`. One
tool, three vocabularies; `sent` also reads as a delivery receipt, which is the
misreading this PR exists to remove.
`_spawn_delivery` now returns `status: queued`, a `delivery_id` and the
`process_id`. For CLI-runner deliveries the id is the same sha256 of the DM
file path the runner pins when it admits to a live owner (`_dm_delivery_id`,
replacing three inline copies), so the ack and any live receipt correlate; the
relay branch passes the envelope id (also on its notification_error ack).
`sent_at` becomes `queued_at`. Tool schema, Bot Mode user guide and the
assertions in tests/tools follow; two invariant tests pin the CLI-runner and
relay ids (red before, green after).
The CLI-runner fallback of `message_agent` returns `status: sent` synchronously
while delivery runs in a background process; when that process dies (e.g. the
`hermes` entrypoint missing from PATH in the Docker image, #114628) the failure
only arrives as the background-process completion notification. The old
`detail` ("Message dispatched ... when the delivery completes, its notification
carries the reply") and the schema description ("delivery acknowledgement")
read as a delivery receipt, which misled the calling agent into reporting a
successful hand-off and routing around the tool.
Reword `detail` and the tool description to say the ack is the hand-off to a
background delivery process, not a delivery receipt, and that the completion
notification carries the outcome — the reply or the delivery failure. Update
the Bot Mode user guide to match. `status: sent` is unchanged (renaming to
`queued` + `delivery_id` is a return-contract change left to the maintainer).
Fixes#114628
Supersedes #111562
`container_backend_for_task` already catches everything inside `_terminal_env_type_for_task`
and `_uses_container_paths` and returns defaults, and every CheckpointManager has
`unsupported_backend_reason`: call the method directly (no getattr fallback), drop the
try/except in `unsupported_backend_reason` and the suppress() around the import+call in
turn_explainers (the one around record_agent_write stays). The rollback.restore fakes in the
server tests gain the method they now must provide.
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.
Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).
Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
_title_match_result shaped its bookends and window with no max_content_len while
_bookend/_hydrate_hit cap FTS hits at 1200/4000, so a session found by title
could still return one 74K tool result verbatim. Same caps, same metadata.
#69334 capped discovery bookends (1200 chars) and scroll windows (4000)
via _shape_message(max_content_len=...) but _read_session still shaped
rows with no cap, so one archived tool result stored as a message came
back whole - a single read returned 74K chars and took a request from
~50K to ~89K tokens in one step.
Cap read-shape messages at 2000 chars with the same content_truncated /
original_content_chars metadata, and add a regression test.
Fixes#114344
The base64 guard covered only `eval` scripts. cmd.exe re-parses every
argument the .cmd shim forwards, so multi-line text sent through `fill`
(browser_type) was truncated at its first line and %VAR% expanded the same
way (#113838). Any non-eval command whose argv carries a newline or % now
runs as `agent-browser batch --json` with the command as a JSON array on
stdin (served from a temp file like stdout/stderr), and the single batch
entry is unwrapped to the usual {success, data, error} shape. The shim test
now drives _run_browser_command with _spawn_and_collect captured, so the
call-site wiring is guarded, not just the helper.
Follow-up to the cherry-picked one-line ``_GET_IMAGES_JS`` (#113844, @KoNit-K),
closing the class the issue asked to audit. Supersedes the earlier #82278 (@Clubheader),
which reached the same shim-truncation diagnosis via ``eval --stdin``; base64 needs no
stdin plumbing and also survives cmd.exe ``%VAR%`` expansion:
- ``browser_tool_session._shim_safe_eval_args``: when the resolved argv[0] is a
``.cmd``/``.bat`` shim (``npx.cmd``, npm's ``agent-browser.cmd`` on Windows)
the ``eval`` script is sent as ``--base64 <b64>`` (``agent-browser eval -b``,
present since the 0.26 floor). cmd.exe re-parses the child command line —
a newline ends the argument and ``%VAR%`` expands even inside quotes — so
this is the only lossless transport for model-authored ``browser_console``
expressions and the vault ``eval`` fallback, not just the bundled constant.
Every other spawn target (native binary, POSIX shim) keeps the raw argv.
- Tests: two invariants in ``tests/tools/test_browser_eval_shim_args.py``
(shim → base64 round trip with native/POSIX/non-eval controls; every
``*_JS`` constant across ``tools/browser_*`` is single-line). The
contributor's get_images regression test is dropped as subsumed by the
module-wide constant scan.
Host-specific (Windows): code-path proof. Live on this host: real
``agent-browser --json eval -b <b64>`` of the collapsed script returns the
image list (data: URIs filtered); ``eval "JSON.stringify("`` — the first line
the shim delivers — reproduces the reporter's exact
``SyntaxError: Unexpected end of input``.
Co-authored-by: Clubheader <Clubheader@users.noreply.github.com>
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:
- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
which logs the single final warning AND drops the supervisor from
``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
(``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
attached" instead of a dead entry, and the next browser call for the task
starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
(two invariants: bounded + unregistered after attach; first-dial failure
still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
main — that file is the opt-in real-Chrome E2E suite and its module-level
skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
reconnect.
Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.
Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
The scope install sat between `_connect_server_claim.set(None)` and the try, so
a raise from hydration or the scope build (lock timeout, .env decode error)
skipped the finally and left the claim cleared in this frame. Set scope_token
to None first and install as the first statement of the try; the finally now
resets both tokens on every exit path.
Invariant test (red before): a raising _install_owner_secret_scope leaves the
caller's claim in place after _connect_server propagates the error.
_load_mcp_config() interpolates ${VAR} header/URL refs through get_secret and
swallows the resulting UnscopedSecretError into {}. It runs at the top of
discover_mcp_tools() (and reconcile / get_mcp_status / probe) in the CALLER's
context, before the connect-site binding added by the previous commits could
run, so an unscoped discover for a routed profile whose config holds any ${VAR}
ref registered ZERO servers — stdio siblings included. The HTTP/SSE atom of
#113746 was therefore still open on that entry point.
Add _owner_secret_scope(), the sync twin of _install_owner_secret_scope()
(same owner rule: _mcp_registry_scope() -> hydrate -> build_profile_secret_scope;
an already-scoped caller or scope key None binds nothing), and wrap every
_load_mcp_config() read in tools/mcp_tool_discovery.py with it. The
connect-site binding stays: run_coroutine_threadsafe copies the loop thread's
context, so the run task's own binding still covers the stdio child env and
later revivals. No environ fallthrough; the secret scope is reset on exit.
Invariant test (red on the previous head): an unscoped discover for a routed
profile with a `${EXAMPLE_TOKEN}` Authorization header hands both servers to
registration with the OWNING profile's value, never the launch env's.
MCP discovery at gateway startup runs outside any per-turn secret scope. Under
multiplexing an unscoped get_secret fails closed, so both credential reads on the
connect path raised UnscopedSecretError:
- _build_safe_env (tools/mcp_tool_config.py) loops secret_source_names() for the
stdio child env. _SECRET_SOURCES is process-wide, so ONE profile configuring an
external secret source (secrets.command / bitwarden / 1password) broke every
stdio server for every profile in the gateway.
- _interpolate_env_vars resolves ${VAR} / ${env:VAR} refs, so HTTP/SSE servers
carrying a token in a header or URL failed the same way.
The server exhausted its initial-connect ladder and parked with zero tools
registered, reporting a credential unrelated to the server that failed.
_connect_server now binds the owning profile's secret scope before start().
start() ensure_futures the run task, which copies that context, so one binding
covers transport bring-up and every later revival inside the task.
The scope comes from _mcp_registry_scope() — the profile the connection is keyed
under — never the merely ambient one, so a served profile cannot be handed
another profile's token (#111151). An already-installed scope is left alone, and
a single-profile process (scope key None) binds nothing and is byte-identical.
Fixes#113746
Signed-off-by: moep90 <volleyballlive@googlemail.com>
Routing sitemap hops through _guarded_http_get dropped the explicit
Accept-Encoding: gzip header from 7050c052e3 (skills.sh serves sitemaps
brotli-compressed and the optional brotlicffi backend mis-decodes them when
streamed), leaving the explanatory comment orphaned. _guarded_http_get and
_ssrf_safe_http_get now accept plain request headers, sent on every hop, and
the sitemap fetches pass the pin again.
tools.skills_hub already owns _guarded_http_get (SSRF pre-check, guarded
client, bounded manual redirects that fail closed on a missing Location,
website-policy check). Route both the skills.sh sitemap index and each
<loc> sitemap through it instead of carrying a second inline guarded
client in the skills.sh adapter. The explicit Accept-Encoding: gzip
header goes away — httpx sends it by default.
Docs: the SSRF section now names the remote-party fetches the guard
covers and the LAN-provider opt-in.
Provider response URLs, model-supplied image refs, manifest-derived pet
URLs, and remote sitemap <loc> entries were fetched with raw
requests/httpx/urllib — bypassing tools/url_safety while every platform
media path already uses it. A hostile or compromised provider/manifest
endpoint could steer a server-side fetch at internal or metadata
addresses; several sites cache the body where it is deliverable back.
Apply the canonical is_safe_url + create_ssrf_safe_client pattern at
every site: per-hop revalidation at TCP connect (closing the
DNS-rebinding window), bounded redirect chains that fail closed on
missing Location, and caller headers scoped to the first hop only —
matching the openrouter provider's own documented contract that its
bearer key must never leave the operator-selected host.
Operator-configured endpoints and pinned release assets are out of
scope — those URLs are operator-selected, not remote-party-controlled.
Fixes#114468Closes#44728
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Ray <rayjun0412@gmail.com>
Co-authored-by: zapabob <1920071390@campus.ouj.ac.jp>
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).
#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.
- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
(RuntimeError subclass) so the dispatcher can tell a host refusal from a
card failure.
- _record_task_failure(infrastructure=True): run + event carry
`infrastructure: true`, consecutive_failures is left alone, the breaker
never trips; the card stays `ready` with the linger remedy in
last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
is such a refusal inside the rate-limit cooldown window, so a dead bus is
retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
managed gateway.
#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.
- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
is scope-wrapped when the user bus is reachable. Without a bus it degrades
with a loud once-per-process warning naming the consequence and both
remedies (enable-linger, KillMode=process) rather than refusing — the
unit's KillMode/lifetime is unknowable and a long-lived sequencer without
linger must keep working.
Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.
Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.
The error now names the resolved home's `.env` (and whether the platform's
token key is defined there), `config.yaml` (block absent / `enabled: false` /
no token) and the environment variable(s) checked, so a Windows or profile
home user can fix the file this process actually read. When a gateway started
from the same home already has the platform connected, the message says its
token lives only in that process's environment and which key to add to `.env`.
Docstrings in `gateway/channel_directory.py` and `send_cmd._load_hermes_env`
stop naming `~/.hermes`; the pipe-script-output guide documents the message.
`hermes send --help`, the `--list` empty-state hint and the "Platform 'x' is
not configured" error hardcoded `~/.hermes/...`, which does not exist on a
Windows install rooted at `%LOCALAPPDATA%\hermes` or under a profile home.
Build the strings from `get_hermes_home()` instead.
Salvaged from #114297 (the `_load_hermes_env` secret-scope rewrite and its
reload-based tests are dropped; see the PR body).
The wrapper (and its __all__ entry) had no production caller: the only
release path, tools/file_tools.py::clear_file_ops_cache, already goes
through file_state.get_registry().forget_task(). It existed solely for
the new registry test, which now calls get_registry().forget_task()
directly, the same path production takes.
Follow-up to the salvaged #114470 commit.
`AIAgent.close()` hands `_close_task_resources()` the agent's session_id, but the
file tools key `FileStateRegistry` by the per-turn task_id: cron runs use
`cron:<job>:<uuid>` while their session_id is `cron_<job>_<ts>`, and delegate
children use `subagent-N-xxxx` against a fresh uuid session. `cleanup_vm(session_id)`
therefore never reached `forget_task()` for the id that owns the read stamps and
writer claims, so the purge added to `forget_task()` had no production caller for
exactly the lifecycles the issue describes. `close()` now runs
`clear_file_ops_cache()` for every id in `_process_owner_task_ids` (the set the
turn context already maintains for process ownership) before dropping the session.
Dropped from #114470: the `HERMES_FILE_STATE_WRITER_TTL` env knob and the
time-based eviction in `check_stale()`. With the lifecycle end actually releasing
the finished task's claims, a TTL only weakens the concurrent case the guard exists
for (a live sibling's hour-old write is still a real conflict), and behavioural
env vars are not a config surface. The module-level `forget_task()` wrapper and
its `__all__` entry from the salvaged commit are kept.
Tests trimmed to two invariants in the mirroring file: `forget_task()` purges the
finished task's writer claims (a live sibling still fires), and `AIAgent.close()`
releases the file state of every task id it ran even though it receives the
session_id.
_scrubbed_env wrapped the scoped_passthrough_additions overlay in try/except Exception with a
debug log; a scope or config failure there would silently drop the declared secret again — the
exact failure mode #114209 complains about. The import cannot fail in-tree and _scrub_child_env
already calls it unguarded, so the guard is removed and one test pins that both local surfaces
propagate the error.