Commit Graph

4945 Commits

Author SHA1 Message Date
teknium1
b71359066c fix: /stop halts background subagents and returns their partial results as interrupted completions
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:

- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
  already did this in `_handle_reset_command`, the second request is
  idempotent) and the idle `_handle_stop_command` tail, which replied
  "No active task to stop." while a background child was running; it now
  stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
  spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.

The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.

An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.

Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.

Part of #114456
2026-09-18 20:34:42 -07:00
teknium1
3ed40556ce fix(profiles): a child spawned for another profile no longer inherits the spawner's authorization gates
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).

- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
  shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
  ...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
  `served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
  action env already funnel through (the dashboard site now calls it for its target).
  Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
  every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
  the profile is not the one the updater runs as; the module stays stdlib-only at
  import time.

Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.

Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
2026-09-18 15:11:47 -07:00
kshitijk4poor
1768a5566f refactor(tools): one find clause for the hidden-dir prune
The filename walk and the grep fallback each hand-built the same
`-type d -name .* ! -path <root> -prune` clause; the first fold drifted
between them, which is how #18473 stayed open. One `_hidden_prune_expr`
serves both; docstring says when the pruned branch is used and why.
2026-09-19 01:46:29 +05:30
kshitijk4poor
ab9a9503b7 fix(tools): a single-file search root under a hidden dir is searched too
Gate review on the previous commit: `find <root> -mindepth 1` yields nothing
when the root is a file, so `search_files(path="~/.hermes/config.yaml")` on
the grep fallback returned 0 matches (main returned 1). Reuse the
`! -path <root>` exemption the filename walk in `_search_files` already uses:
a dot-named root — directory or single file — is never pruned, hidden
descendants still are.

The root predicate now goes through `_normalized_filename_search_root`, which
already handles MSYS paths on Windows and posixpath for remote environments
instead of applying host `os.path` to a backend path.
2026-09-19 01:46:29 +05:30
kshitijk4poor
9106580025 fix(tools): search_files grep fallback finds content under hidden directories
The grep fallback passed --exclude-dir='.*' and the search root together. grep
applies --exclude-dir to the command-line root as well as to subdirectories
(GNU grep to every component of it), so any content search rooted at or under a
dot-directory such as ~/.hermes returned zero matches when ripgrep was absent,
while the same search returned results with rg. Reported in #18473 with a live
12-vs-0 comparison on GNU grep 3.7; BSD grep drops the root only when the root
itself is dot-named.

Route those roots through the existing find -prune pipeline, which already
mirrors rg's hidden-directory default for descendants, and give find -mindepth 1
so the prune cannot swallow a dot-named root. The macOS protected-folder branch
shares the pipeline, so the prune clause is now built only when there are
protected paths to prune.

Credit: @memosr (#34017) and @luxles (#48878) proposed dropping --exclude-dir for
hidden roots first; that variant leaks hidden descendants, which the find route
keeps excluded.
2026-09-19 01:46:29 +05:30
teknium1
d457e7f039 test(skills): seam the sitemap-shard retry test on the hub guarded GET
The shard fetch now routes through hub()._guarded_http_get (#114887), so a patch
on tools.skills_hub_skillssh._get_text intercepted nothing and the test reached
the live sitemap on CI. Patch tools.skills_hub._ssrf_safe_http_get like the
neighbouring guarded-GET test and model the dark shard as a 503.
2026-09-18 12:51:49 -07:00
teknium1
43eda09e2d fix(skills): retry a failed skills.sh sitemap shard; never cache a partial catalog
A per-skill sitemap shard whose fetch failed (timeout/503 -> None) was read as
an empty shard and the resulting partial catalog was written to the shared
skills_sh_sitemap_v1 cache, so one slow shard silently dropped skills from the
published index for the whole TTL. Same class as the ClawHub page-walk fix in
this PR: retry the shard up to the shared CATALOG_PAGE_RETRIES bound, and if it
stays dark return the partial slice without caching it.

CATALOG_PAGE_RETRIES moves from ClawHubSource to the SkillSource base so both
catalog walks share one bound.
2026-09-18 12:51:49 -07:00
teknium1
73f39086f6 fix(skills-index): retry only on a failed fetch; give the scheduled build time for slow ClawHub pages
Trim the salvaged retry to the case the logs show: `_get_json` returning None
(timeout / non-200) is a hole in the cursor walk, not the end of the catalog.
A dict page with an empty item list stays terminal as before, and the
interactive browse path (12 s budget) no longer sleeps up to 8 s per retry.
Constant `CATALOG_PAGE_RETRIES` replaces the local literal.

Workflow: `skills-index.yml` `timeout-minutes` 50 -> 120. Since Sep 17 ClawHub
serves ~10 s per 200-item page (was ~2.5 s), so the ~400-page walk alone takes
60-70 min; the 06:28 UTC runs on Sep 17 and Sep 18 were cancelled at 50 min with
the clawhub crawl still running, and the 18:19 run stopped at 12,145 skills
(< 20,000 floor) after one failed page. Both halves are needed for the index to
ship again.
2026-09-18 12:51:49 -07:00
joaomarcos
b26944ce3b fix(skills): retry ClawHub catalog pages on fetch failure
Treat timeout/non-200 catalog pages as a retryable hole instead of
cursor exhaustion so the skills index builder can finish a full walk.

Fixes #114525
2026-09-18 12:51:49 -07:00
teknium1
1f1fa9a87e fix(tools): execute_code shares the interpreter-kill rejection text; collapse redundant supervisor branch
- HOST_INTERPRETER_KILL_REJECTION is one constant in cron/lifecycle_guard;
  terminal_tool_guards and the execute_code lifecycle guard both use it,
  so an image-name kill inside a cell now names the proc_* / explicit-PID
  route instead of the generic "cannot restart or stop the gateway" text.
  One invariant test with the generic path as control.
- gateway/restart.is_supervised_gateway_launch: the callee already maps
  None to os.environ; pass environ straight through.
2026-09-18 12:49:13 -07:00
teknium1
726bccd28d fix(cron): lifecycle guard blocks kills aimed at the gateway's own interpreter image
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.

The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).

Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.

Fixes #113667
2026-09-18 12:49:13 -07:00
teknium1
9bd8257c02 fix(mcp): complete the handshake when a stateless server names a modern protocolVersion and lacks server/discover
A server that answers the legacy initialize with HTTP 200 but reports
2026-07-28 (regardless of the version offered) makes the SDK raise
'Unsupported protocol version from the server'. auto mode then falls back
to server/discover, which such a server rejects with a non-JSON 4xx, and
the connect died with 'both Streamable HTTP and SSE transports failed
(Streamable HTTP: Server returned an error response; SSE: 405)' even
though the same initialize/tools/list POSTs succeed from curl (#113359).

When both fail for that reason, re-run initialize ourselves, adopt the
result pinned to the version we offered (keeps later requests
legacy-shaped, which the handshake just proved the server accepts), send
notifications/initialized and proceed to tools/list.
2026-09-18 12:48:42 -07:00
teknium1
f0c26f0554 fix(mcp): name the HTTP status, URL and body behind "Server returned an error response"
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).

- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
  that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
  bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
  the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
  error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
  ConnectionError; the no-fallback re-raise after a proven session, strict
  redirect headers or a non-rejection) carry the detail. Debug-logs the
  endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.

Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-18 12:48:42 -07:00
teknium1
5a64a888da fix(mcp): pin the pre-flight discovery User-Agent in a test; drop the unfailing version import guard
Review follow-up: swapping tools/mcp_oauth_manager.py back to origin/main kept every
PR test green, so the pre-flight PRM/ASM discovery stamp had no invariant. The existing
prefetch test now records the User-Agent of each mocked discovery request and asserts
it is DEFAULT_AUTH_REQUEST_USER_AGENT (red without the stamp, green with it).

hermes_cli/__init__.py is stdlib-only and defines __version__ at module top, so the
try/except around its import could never fire; import directly.
2026-09-18 12:47:49 -07:00
teknium1
668e505ac7 fix(mcp): OAuth discovery/registration carry a User-Agent; cancelled login frees its callback port
Widen the salvaged fixes to the whole class and add the pieces they missed:

- The default User-Agent for SDK-built OAuth requests moves from the manager's
  bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
  provider gets it too, and the manager's pre-flight metadata discovery (its own
  client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
  registration and token requests through client.send(), which never merges the
  client's default headers; www.tradingview.com's WAF answers a header-less GET
  with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
  /register and /authorize on the MCP host, and the login died with
  "Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
  ... != None" because no issuer was ever discovered).

- The callback listener now runs serve_forever() and is shut down before
  server_close(). A thread parked in handle_request()'s select() keeps the
  closed listening socket alive (the kernel holds the file for the duration
  of the poll), so a flow cancelled mid-wait left the port bound and the
  retry on the same pinned/cached port raised "OAuth callback port N is
  already in use" with no external collider.

- When every authorization-server metadata fetch failed, a registration error
  is re-raised leading with those statuses ("Could not read
  authorization-server metadata (403 from ...); dynamic client registration
  then fell back to a guessed endpoint on the MCP host and failed: ...").
  humanize_oauth_registration_error leaves that message alone so the 403 in
  it is not mistaken for a DCR allowlist refusal.

Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
2026-09-18 12:47:49 -07:00
Ivyleaguelawyer
7eed615690 fix(mcp): send a User-Agent on SDK-built OAuth discovery/registration requests
The MCP SDK builds /.well-known discovery and dynamic client registration
requests as bare httpx.Request objects inside async_auth_flow, so the
client's default headers never apply and they leave with no User-Agent at
all. Some WAFs reject header-less requests outright: coda.io returns 403
on every discovery and registration call, which Hermes then misreports as
"'<server>' only allows pre-approved OAuth clients". The existing
oauth.user_agent option does not help because it is stamped only on
token-endpoint requests.

HermesMCPOAuthProvider's generator bridge now gives SDK-built requests a
default User-Agent ("Hermes-Agent") when the header is absent. The
caller's own MCP request is never touched, and an explicitly configured
oauth.user_agent on token requests still wins.

Reproduced against https://coda.io/apis/mcp: before, every well-known
and /register call returned 403; after, registration returns 200 and the
browser authorization flow proceeds.

Co-Authored-By: claude-flow <ruv@ruv.net>
2026-09-18 12:47:49 -07:00
teknium1
63fb1a7d36 fix(kanban): warn once when an inherited delegation fence rejects the worker's auto-heartbeat
A worker process that carries HERMES_DELEGATED_CHILD_CONTEXT next to its
HERMES_KANBAN_TASK is fenced by kanban_path_is_fenced's env-marker branch: both
auto-heartbeat writes raised PermissionError at DEBUG only, so the board showed a
worker that never beats while the process was alive (the salvaged commit already
made the return value honest, but _touch_activity ignores it). Log the refusal once
per process at WARNING with the cause, and skip the bridge for an in-process
delegate child BEFORE stamping the rate-limit window so a chatty child cannot starve
the worker's own heartbeat.

No self-write grant is added: the marker + TASK combination means "descendant, not
worker" by design (b578261584, 8a8c3634e8), and the only in-tree grant path,
kanban_db_dispatch._default_spawn, already pops the marker. The real-spawn grant
test now proves that from a dispatcher that itself carries the marker: the worker's
auto-heartbeat lands and its handoff completes.

Docs: worker guide notes the automatic claim extension and what the refusal warning
means.
2026-09-18 12:46:28 -07:00
KoNit-K
3aa76b2ce5 fix(kanban): report failed auto-heartbeats 2026-09-18 12:46:28 -07:00
kshitijk4poor
027d1a8a60 fix(tools): read_file's dedup stub skips the review fork too
Sibling of the previous commit: read_file_tool returned the unchanged-stub
before _record_read, and only _record_read calls
mark_background_review_skill_read. A skill file the parent had already read
left the fork with a stub and a refused skill_manage write. Same rule as
skill_view: no dedup stub when is_background_review().

Also: is_background_review hoisted to module level in skills_tool (no cycle —
skill_provenance imports only contextvars), comment reworded to the two real
reasons, test renamed to the shipped design, fixture seeds a fresh read-mark
store so the mark cannot leak between tests.
2026-09-19 00:33:58 +05:30
panyaozhen
bc471d68b8 fix(skills): the background-review fork never receives a skill_view dedup stub
The review fork reuses the parent's session/task id for prefix-cache parity, so
its skill_view calls hit the parent's repeat-view dedup and got a
content_returned:false stub. The stub path never calls
mark_background_review_skill_read, so every skill_manage patch/write_file in
the fork was refused by the read-before-write guard: 408 refusals and zero
skill updates in one deployment, the same signature on three more (#95976).

A dedup stub is only valid while the referenced content is in the caller's own
context; the fork's context is not the parent's. Skip the dedup entirely when
is_background_review(): the fork's view is a real read (which marks the file
for the guard) and records nothing, so the parent's cache is untouched.

Salvaged from #95975 (@yaozhen), rebased onto the decomposed skills_tool and
narrowed from a namespaced bucket to no dedup in the fork.
2026-09-19 00:33:58 +05:30
kshitijk4poor
8503ee4459 refactor(send): route WhatsApp mentions through the existing standalone chunker
Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).

- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
  copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
  payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
  is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
  sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
  already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
  feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
  changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
  (rejections never reach the bridge; mentions ride the first payload only, stale bridge
  fails closed); drop the fabricated legacy-sender branch that only existed to cover the
  deleted probe. Mutation check: removing the first-payload gate turns the new test red.
2026-09-19 00:24:50 +05:30
Andrey
b34ebc084a feat(send): add WhatsApp native mentions
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
2026-09-19 00:24:50 +05:30
Kevin Rajan
1131fd56c2 fix(cron): reap stale executions before the one-shot early return
`_try_dispatch_background_run` returned at the `async_delivery_supported()`
gate before reaching `_reap_stale_executions`, so on the one-shot
`hermes cron run` path (the CLI scopes `_SESSION_ASYNC_DELIVERY` to False)
the dead-owner reap never ran: an execution row left 'running' by a killed
prior run stayed that way until a scheduler tick — exactly the gap the
helper's own docstring says it exists to close.

Hoist the reap above the gate; the async dispatch path keeps its existing
reap-before-claim ordering.

Salvaged from #113938 (@kvnloo, source hunk verbatim; its test hunk is
replaced by an invariant test in a follow-up commit). #86981 (@pierrenode)
identified the same gap first and placed a duplicate reap in
`_execute_job_now`; superseded by this single hoist.

Part of #113923.

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-09-18 11:07:56 -07:00
teknium1
2d773ff091 fix: local delivery runner ships a typed reason for every failure
The `--run-delivery` catch-all in tools/bot_mode_dm.py printed a structured
`{error, reason}` only for target_busy; any other exception went to stderr as
prose, so the sender's completion notification had no reason to branch on
(auth failure vs. transient) for the local lane while the relay lane had one.

The vocabulary-guarded helper the relay lane introduced moves beside the
vocabulary it guards (tools/bot_failure_reasons.delivery_failure_reason) and
both lanes call it: an exception's own `reason` is trusted only when it is
`target_busy` or in ALL_REASONS, otherwise the text is classified. The runner
now prints `{error, reason}` JSON on stdout for every exception (exit 1 kept).

Test: a classifiable, an unclassifiable and an out-of-vocabulary-reason
exception in --run-delivery each yield JSON with a vocabulary reason.
2026-09-18 10:44:29 -07:00
teknium1
250bdcc311 fix(bot-mode): message_agent CLI-runner ack is queued + delivery_id like the other branches
Review finding on #114941: the CLI-runner branch of `_spawn_delivery` still
answered `status: sent` with only a `process_id`, while the live-owner branch
answers `queued`/`claimed` + `delivery_id` and the relay branch `queued`. One
tool, three vocabularies; `sent` also reads as a delivery receipt, which is the
misreading this PR exists to remove.

`_spawn_delivery` now returns `status: queued`, a `delivery_id` and the
`process_id`. For CLI-runner deliveries the id is the same sha256 of the DM
file path the runner pins when it admits to a live owner (`_dm_delivery_id`,
replacing three inline copies), so the ack and any live receipt correlate; the
relay branch passes the envelope id (also on its notification_error ack).
`sent_at` becomes `queued_at`. Tool schema, Bot Mode user guide and the
assertions in tests/tools follow; two invariant tests pin the CLI-runner and
relay ids (red before, green after).
2026-09-18 10:39:50 -07:00
teknium1
ef49567128 fix(bot-mode): message_agent ack says it is a dispatch hand-off, not a delivery receipt
The CLI-runner fallback of `message_agent` returns `status: sent` synchronously
while delivery runs in a background process; when that process dies (e.g. the
`hermes` entrypoint missing from PATH in the Docker image, #114628) the failure
only arrives as the background-process completion notification. The old
`detail` ("Message dispatched ... when the delivery completes, its notification
carries the reply") and the schema description ("delivery acknowledgement")
read as a delivery receipt, which misled the calling agent into reporting a
successful hand-off and routing around the tool.

Reword `detail` and the tool description to say the ack is the hand-off to a
background delivery process, not a delivery receipt, and that the completion
notification carries the outcome — the reply or the delivery failure. Update
the Bot Mode user guide to match. `status: sent` is unchanged (renaming to
`queued` + `delivery_id` is a return-contract change left to the maintainer).

Fixes #114628
Supersedes #111562
2026-09-18 10:39:50 -07:00
teknium1
1b083b85a3 fix(checkpoints): drop the defence layers around the never-raising backend predicate
`container_backend_for_task` already catches everything inside `_terminal_env_type_for_task`
and `_uses_container_paths` and returns defaults, and every CheckpointManager has
`unsupported_backend_reason`: call the method directly (no getattr fallback), drop the
try/except in `unsupported_backend_reason` and the suppress() around the import+call in
turn_explainers (the one around record_agent_write stays). The rollback.restore fakes in the
server tests gain the method they now must provide.
2026-09-18 10:38:08 -07:00
Sora-bluesky
3220b9ed2f fix(checkpoints): classify the backend on every /rollback instead of remembering it
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
2026-09-18 10:38:08 -07:00
Sora-bluesky
10ee9819f9 fix(checkpoints): do not feed container paths into host checkpoint storage
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.

Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).

Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
2026-09-18 10:38:08 -07:00
teknium1
0a41392489 fix(session_search): cap title-match discovery entries like FTS hits
_title_match_result shaped its bookends and window with no max_content_len while
_bookend/_hydrate_hit cap FTS hits at 1200/4000, so a session found by title
could still return one 74K tool result verbatim. Same caps, same metadata.
2026-09-18 10:30:57 -07:00
Simon Crouch
17869e8bd7 fix(session_search): cap per-message content on the read shape
#69334 capped discovery bookends (1200 chars) and scroll windows (4000)
via _shape_message(max_content_len=...) but _read_session still shaped
rows with no cap, so one archived tool result stored as a message came
back whole - a single read returned 74K chars and took a request from
~50K to ~89K tokens in one step.

Cap read-shape messages at 2000 chars with the same content_truncated /
original_content_chars metadata, and add a regression test.

Fixes #114344
2026-09-18 10:30:57 -07:00
teknium1
6c9e583860 fix(browser): route any newline/% argument past a Windows .cmd shim losslessly
The base64 guard covered only `eval` scripts. cmd.exe re-parses every
argument the .cmd shim forwards, so multi-line text sent through `fill`
(browser_type) was truncated at its first line and %VAR% expanded the same
way (#113838). Any non-eval command whose argv carries a newline or % now
runs as `agent-browser batch --json` with the command as a JSON array on
stdin (served from a temp file like stdout/stderr), and the single batch
entry is unwrapped to the usual {success, data, error} shape. The shim test
now drives _run_browser_command with _spawn_and_collect captured, so the
call-site wiring is guarded, not just the helper.
2026-09-18 10:29:47 -07:00
teknium1
35a03bce14 fix(browser): eval scripts reach a Windows .cmd shim base64-encoded; tests trimmed
Follow-up to the cherry-picked one-line ``_GET_IMAGES_JS`` (#113844, @KoNit-K),
closing the class the issue asked to audit. Supersedes the earlier #82278 (@Clubheader),
which reached the same shim-truncation diagnosis via ``eval --stdin``; base64 needs no
stdin plumbing and also survives cmd.exe ``%VAR%`` expansion:

- ``browser_tool_session._shim_safe_eval_args``: when the resolved argv[0] is a
  ``.cmd``/``.bat`` shim (``npx.cmd``, npm's ``agent-browser.cmd`` on Windows)
  the ``eval`` script is sent as ``--base64 <b64>`` (``agent-browser eval -b``,
  present since the 0.26 floor). cmd.exe re-parses the child command line —
  a newline ends the argument and ``%VAR%`` expands even inside quotes — so
  this is the only lossless transport for model-authored ``browser_console``
  expressions and the vault ``eval`` fallback, not just the bundled constant.
  Every other spawn target (native binary, POSIX shim) keeps the raw argv.
- Tests: two invariants in ``tests/tools/test_browser_eval_shim_args.py``
  (shim → base64 round trip with native/POSIX/non-eval controls; every
  ``*_JS`` constant across ``tools/browser_*`` is single-line). The
  contributor's get_images regression test is dropped as subsumed by the
  module-wide constant scan.

Host-specific (Windows): code-path proof. Live on this host: real
``agent-browser --json eval -b <b64>`` of the collapsed script returns the
image list (data: URIs filtered); ``eval "JSON.stringify("`` — the first line
the shim delivers — reproduces the reporter's exact
``SyntaxError: Unexpected end of input``.

Co-authored-by: Clubheader <Clubheader@users.noreply.github.com>
2026-09-18 10:29:47 -07:00
KoNit-K
1222172b2f fix(browser): preserve image eval payload on Windows 2026-09-18 10:29:47 -07:00
teknium1
aefa503479 fix(browser): gave-up supervisor unregisters itself; tests trimmed to two invariants
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:

- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
  which logs the single final warning AND drops the supervisor from
  ``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
  (``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
  attached" instead of a dead entry, and the next browser call for the task
  starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
  (two invariants: bounded + unregistered after attach; first-dial failure
  still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
  main — that file is the opt-in real-Chrome E2E suite and its module-level
  skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
  reconnect.

Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.

Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
2026-09-18 10:28:40 -07:00
KoNit-K
67b061a51c fix(browser): bound CDP supervisor reconnects 2026-09-18 10:28:40 -07:00
teknium1
92deb8e871 fix(mcp): _connect_server installs the owner scope inside the try so a failed install releases the claim
The scope install sat between `_connect_server_claim.set(None)` and the try, so
a raise from hydration or the scope build (lock timeout, .env decode error)
skipped the finally and left the claim cleared in this frame. Set scope_token
to None first and install as the first statement of the try; the finally now
resets both tokens on every exit path.

Invariant test (red before): a raising _install_owner_secret_scope leaves the
caller's claim in place after _connect_server propagates the error.
2026-09-18 10:25:46 -07:00
teknium1
7529ff0de1 fix(mcp): discover binds the owner's secret scope before ${VAR} interpolation at config load
_load_mcp_config() interpolates ${VAR} header/URL refs through get_secret and
swallows the resulting UnscopedSecretError into {}. It runs at the top of
discover_mcp_tools() (and reconcile / get_mcp_status / probe) in the CALLER's
context, before the connect-site binding added by the previous commits could
run, so an unscoped discover for a routed profile whose config holds any ${VAR}
ref registered ZERO servers — stdio siblings included. The HTTP/SSE atom of
#113746 was therefore still open on that entry point.

Add _owner_secret_scope(), the sync twin of _install_owner_secret_scope()
(same owner rule: _mcp_registry_scope() -> hydrate -> build_profile_secret_scope;
an already-scoped caller or scope key None binds nothing), and wrap every
_load_mcp_config() read in tools/mcp_tool_discovery.py with it. The
connect-site binding stays: run_coroutine_threadsafe copies the loop thread's
context, so the run task's own binding still covers the stdio child env and
later revivals. No environ fallthrough; the secret scope is reset on exit.

Invariant test (red on the previous head): an unscoped discover for a routed
profile with a `${EXAMPLE_TOKEN}` Authorization header hands both servers to
registration with the OWNING profile's value, never the launch env's.
2026-09-18 10:25:46 -07:00
moep90
d34423e4de fix(mcp): connect resolves credentials under the connection owner's profile scope
MCP discovery at gateway startup runs outside any per-turn secret scope. Under
multiplexing an unscoped get_secret fails closed, so both credential reads on the
connect path raised UnscopedSecretError:

- _build_safe_env (tools/mcp_tool_config.py) loops secret_source_names() for the
  stdio child env. _SECRET_SOURCES is process-wide, so ONE profile configuring an
  external secret source (secrets.command / bitwarden / 1password) broke every
  stdio server for every profile in the gateway.
- _interpolate_env_vars resolves ${VAR} / ${env:VAR} refs, so HTTP/SSE servers
  carrying a token in a header or URL failed the same way.

The server exhausted its initial-connect ladder and parked with zero tools
registered, reporting a credential unrelated to the server that failed.

_connect_server now binds the owning profile's secret scope before start().
start() ensure_futures the run task, which copies that context, so one binding
covers transport bring-up and every later revival inside the task.

The scope comes from _mcp_registry_scope() — the profile the connection is keyed
under — never the merely ambient one, so a served profile cannot be handed
another profile's token (#111151). An already-installed scope is left alone, and
a single-profile process (scope key None) binds nothing and is byte-identical.

Fixes #113746

Signed-off-by: moep90 <volleyballlive@googlemail.com>
2026-09-18 10:25:46 -07:00
teknium1
309023a00c fix(skills-hub): keep the gzip-only Accept-Encoding pin on sitemap fetches
Routing sitemap hops through _guarded_http_get dropped the explicit
Accept-Encoding: gzip header from 7050c052e3 (skills.sh serves sitemaps
brotli-compressed and the optional brotlicffi backend mis-decodes them when
streamed), leaving the explanatory comment orphaned. _guarded_http_get and
_ssrf_safe_http_get now accept plain request headers, sent on every hop, and
the sitemap fetches pass the pin again.
2026-09-18 10:22:14 -07:00
teknium1
5236a58005 refactor(skills-hub): sitemap fetches reuse the hub's guarded GET
tools.skills_hub already owns _guarded_http_get (SSRF pre-check, guarded
client, bounded manual redirects that fail closed on a missing Location,
website-policy check). Route both the skills.sh sitemap index and each
<loc> sitemap through it instead of carrying a second inline guarded
client in the skills.sh adapter. The explicit Accept-Encoding: gzip
header goes away — httpx sends it by default.

Docs: the SSRF section now names the remote-party fetches the guard
covers and the LAN-provider opt-in.
2026-09-18 10:22:14 -07:00
beardthelion
3933fdf63b fix(security): route remote-party-supplied URL fetches through the SSRF guard
Provider response URLs, model-supplied image refs, manifest-derived pet
URLs, and remote sitemap <loc> entries were fetched with raw
requests/httpx/urllib — bypassing tools/url_safety while every platform
media path already uses it. A hostile or compromised provider/manifest
endpoint could steer a server-side fetch at internal or metadata
addresses; several sites cache the body where it is deliverable back.

Apply the canonical is_safe_url + create_ssrf_safe_client pattern at
every site: per-hop revalidation at TCP connect (closing the
DNS-rebinding window), bounded redirect chains that fail closed on
missing Location, and caller headers scoped to the first hop only —
matching the openrouter provider's own documented contract that its
bearer key must never leave the operator-selected host.

Operator-configured endpoints and pinned release assets are out of
scope — those URLs are operator-selected, not remote-party-controlled.

Fixes #114468
Closes #44728

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Ray <rayjun0412@gmail.com>
Co-authored-by: zapabob <1920071390@campus.ouj.ac.jp>
2026-09-18 10:22:14 -07:00
teknium1
eb28bc1bf7 fix(kanban): host spawn refusals stop charging cards; oneshot-unit dispatch scope-wraps workers
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).

#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.

- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
  (RuntimeError subclass) so the dispatcher can tell a host refusal from a
  card failure.
- _record_task_failure(infrastructure=True): run + event carry
  `infrastructure: true`, consecutive_failures is left alone, the breaker
  never trips; the card stays `ready` with the linger remedy in
  last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
  is such a refusal inside the rate-limit cooldown window, so a dead bus is
  retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
  managed gateway.

#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.

- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
  old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
  is scope-wrapped when the user bus is reachable. Without a bus it degrades
  with a loud once-per-process warning naming the consequence and both
  remedies (enable-linger, KillMode=process) rather than refusing — the
  unit's KillMode/lifetime is unknowable and a long-lived sequencer without
  linger must keep working.

Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.

Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 10:21:42 -07:00
teknium1
2164221844 fix(send): name the default-root gateway and external secret sources in the 'not configured' error
Under `HERMES_HOME=<root>/profiles/<name>` the reporter's gateway ran from `<root>`, so
`_not_configured_error` also reads `<root>/gateway_state.json`; when the platform is
connected there under a live pid it appends "A gateway (pid N) running from <root> has
<platform> connected; this shell is scoped to profile home <home> whose .env has no
<VARS>." (#114272 step 5). The `--list` empty state likewise names the root's existing
`channel_directory.json`. The consulted-sources list gains "external secret sources
(<name>: enabled|disabled | none configured)" from agent.secret_sources.registry —
names only.
2026-09-18 10:11:14 -07:00
teknium1
7a99c4d13c fix(send): the 'not configured' error lists the home and sources it consulted
The error now names the resolved home's `.env` (and whether the platform's
token key is defined there), `config.yaml` (block absent / `enabled: false` /
no token) and the environment variable(s) checked, so a Windows or profile
home user can fix the file this process actually read. When a gateway started
from the same home already has the platform connected, the message says its
token lives only in that process's environment and which key to add to `.env`.

Docstrings in `gateway/channel_directory.py` and `send_cmd._load_hermes_env`
stop naming `~/.hermes`; the pipe-script-output guide documents the message.
2026-09-18 10:11:14 -07:00
Rook-CodeVolt
59e40220e2 fix(send): derive hermes send path hints from the resolved Hermes home
`hermes send --help`, the `--list` empty-state hint and the "Platform 'x' is
not configured" error hardcoded `~/.hermes/...`, which does not exist on a
Windows install rooted at `%LOCALAPPDATA%\hermes` or under a profile home.
Build the strings from `get_hermes_home()` instead.

Salvaged from #114297 (the `_load_hermes_env` secret-scope rewrite and its
reload-based tests are dropped; see the PR body).
2026-09-18 10:11:14 -07:00
teknium1
783f854b0f fix(file_state): drop the module-level forget_task wrapper
The wrapper (and its __all__ entry) had no production caller: the only
release path, tools/file_tools.py::clear_file_ops_cache, already goes
through file_state.get_registry().forget_task(). It existed solely for
the new registry test, which now calls get_registry().forget_task()
directly, the same path production takes.
2026-09-18 10:10:40 -07:00
teknium1
f7a422ee4a fix(file-state): close() releases file state for every task id the agent ran; drop the writer TTL knob
Follow-up to the salvaged #114470 commit.

`AIAgent.close()` hands `_close_task_resources()` the agent's session_id, but the
file tools key `FileStateRegistry` by the per-turn task_id: cron runs use
`cron:<job>:<uuid>` while their session_id is `cron_<job>_<ts>`, and delegate
children use `subagent-N-xxxx` against a fresh uuid session. `cleanup_vm(session_id)`
therefore never reached `forget_task()` for the id that owns the read stamps and
writer claims, so the purge added to `forget_task()` had no production caller for
exactly the lifecycles the issue describes. `close()` now runs
`clear_file_ops_cache()` for every id in `_process_owner_task_ids` (the set the
turn context already maintains for process ownership) before dropping the session.

Dropped from #114470: the `HERMES_FILE_STATE_WRITER_TTL` env knob and the
time-based eviction in `check_stale()`. With the lifecycle end actually releasing
the finished task's claims, a TTL only weakens the concurrent case the guard exists
for (a live sibling's hour-old write is still a real conflict), and behavioural
env vars are not a config surface. The module-level `forget_task()` wrapper and
its `__all__` entry from the salvaged commit are kept.

Tests trimmed to two invariants in the mirroring file: `forget_task()` purges the
finished task's writer claims (a live sibling still fires), and `AIAgent.close()`
releases the file state of every task id it ran even though it receives the
session_id.
2026-09-18 10:10:40 -07:00
Omid Zaferi
16d3b21edb fix(file_state): purge stale writer claims in forget_task and add writer TTL (#114446)
(cherry picked from commit b5c1ea32151e56d197e371dd7a5d682218903e8b)
2026-09-18 10:10:40 -07:00
teknium1
547fff7500 fix(tools): scope-only passthrough overlay raises instead of silently dropping the declared secret
_scrubbed_env wrapped the scoped_passthrough_additions overlay in try/except Exception with a
debug log; a scope or config failure there would silently drop the declared secret again — the
exact failure mode #114209 complains about. The import cannot fail in-tree and _scrub_child_env
already calls it unguarded, so the guard is removed and one test pins that both local surfaces
propagate the error.
2026-09-18 10:04:57 -07:00