Widen PR #87757 to cover the ZIP path: _update_via_zip() also calls
_finish_dashboard_update_cleanup() but never runs _reload_config_modules,
so the Windows git-broken fallback would still crash with the same
ImportError (cannot import name 'bounded_probe_run' from the stale cached
hermes_cli._subprocess_compat).
- new _reload_process_scan_modules() called inside
_finish_dashboard_update_cleanup itself, so every current and future
call site is covered; reloads dependency-first
(_subprocess_compat, then dashboard_procs)
- reload failures log at warning (a miss surfaces seconds later as an
ImportError in the same process)
- regression tests: reload-before-kill ordering, node-failure skip,
stale-module symbol restoration (the exact #87134 boundary state),
nonfatal reload failure, and the #87757 reload-list contract
hermes update runs in the PRE-pull Python process. After git pull updates
source files on disk, modules already in sys.modules still hold the OLD
code. The existing _reload_config_modules() reloaded only config modules,
but the post-update dashboard cleanup path (_finish_dashboard_update_cleanup
-> _scan_dashboard_processes) imports hermes_cli._subprocess_compat lazily;
a new symbol added there (e.g. bounded_probe_run) is invisible to the
cached module object, causing ImportError during the cleanup step.
Extend the reload list to include hermes_cli._subprocess_compat and
hermes_cli.dashboard_procs so the cleanup uses freshly-pulled code.
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
Secondary-gateway events were tagged with connectionId (store/gateway
fan-out) but no consumer read it: working/attention tracking, the
pruneSecondaryGateways keep-set, and the profile-scoped event gates
(skin.changed / change-watcher broadcasts / approval-mode reconcile)
all keyed by session id + bare profile name. Every registered source
exposes a 'default' profile (the roster force-unshifts it), so two
connected gateways collided — gateway B's 'default' activity was
attributed to gateway A's 'default', keeping the wrong socket alive
and applying the wrong source's config/skin/cron changes.
Thread connectionId through consumption using the existing composite
backendScopeKey helper:
- session-states records each registry-tagged event's (connectionId,
profile) scope per runtime session; liveSessionScopes() projects the
busy/needs-input ones as composite keys for the gateway keep-set.
- recomputeKeptGateways (use-gateway-boot) seeds the keep-set with
those scopes; pruneSecondaryGateways matches registry-scoped entries
ONLY on their composite key, while local entries keep matching bare
profile names (single-source path unchanged).
- gateway-event's 'from the active profile' gates now compare the
event's composite scope against the active gateway's connection via
the new activeGatewayConnectionId(); untagged local/primary events
behave byte-identically.
Display-only surfaces that already use roster handles are untouched.
ensureRegistryBackend delegated kind==='local' to ensureBackend(), which
follows the v1 connection.json routing table — under a v1 REMOTE global
mode (the migration keeps the mandatory 'local' entry AND makes that
remote the registry primary) the roster's 'This device' rows enumerated
and dialed the REMOTE primary: every profile appeared twice (forcing
-slug handles) and clicking a local agent talked to the remote box.
resolveRegistryLocalRoute() (pure, colocated with the registry helpers)
now decides the local entry's path: delegate to the legacy route only
when v1 is itself local (single-source behavior byte-identical);
otherwise spawn/reuse a forced-local pool child via spawnPoolBackend's
new forceLocal option, pooled under the composite conn:local::<profile>
key so it cannot collide with the v1 remote descriptor cached at the
bare profile key.
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.
Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
The Windows-only test asserted encoding/errors kwargs on a mocked
subprocess.run, but the scan now routes through bounded_probe_run
(#87134), so subprocess.run is never invoked. Assert the probe call's
contract instead (errors='ignore', finite timeout), verify the parsed
PIDs, and add a fail-open case for probe failure. The test no longer
needs a Windows host once the probe is mocked, so the windows_only
gate is dropped.
subprocess.run(capture_output=True, timeout=N) is not hang-safe on
Windows: after the timeout fires, run()'s cleanup kills the direct child
and then joins the pipe reader threads with an UNBOUNDED communicate().
A descendant (conhost.exe under wmic/powershell) holding duplicated pipe
handles keeps the pipes from EOF and the join never returns.
_scan_gateway_pids() runs its wmic / Get-CimInstance Win32_Process scans
exactly that way, and on machines where the full process scan genuinely
exceeds its 10/15s budget (cold WMI on first boot, ARM VMs, heavy
Update/AV activity) hermes update wedged forever inside
_pause_windows_gateways_for_update() before printing a single line —
observed live on a fresh Windows 11 ARM64 VM with a faulthandler stack
pinning the main thread in subprocess._communicate and only a conhost.exe
child surviving. The single-flight update lock then blocks retries until
the wedged process is killed by hand.
This is the same deadlock class bounded_git_probe already fixed for git
probes (#68609 / #66037). Generalize that proven pattern into a shared
bounded_probe_run() — explicit communicate(timeout), kill_process_tree on
failure, bounded 1s drain, then abandon the daemonic readers — and
migrate the whole call-site class onto it:
- hermes_cli/gateway.py _scan_gateway_pids (the site that hung; reached
from hermes update, cron, gateway restart/status, dashboard)
- hermes_cli/dashboard_procs.py wmic scan (same shape, reached on update)
- hermes_cli/claw.py tasklist + PowerShell probes (same shape; its
try/except cannot catch a hang because a hang raises nothing)
- bounded_git_probe now delegates to bounded_probe_run (identical
contract, one copy of the cleanup logic)
Unlike bounded_git_probe, bounded_probe_run returns the CompletedProcess
(or None) rather than collapsing to stdout, because the gateway scan
branches on returncode to trip its wmic -> powershell fallback.
Tests: tests/hermes_cli/test_bounded_probe_run.py covers success,
nonzero-exit passthrough, spawn failure, bounded timeout (fails against
the old unbounded semantics — verified by sabotage), errors= decoding,
DEVNULL stdin, POSIX process-group placement, and the bounded_git_probe
delegation contract. Existing test_git_probe_tree_kill.py passes
unchanged against the delegated implementation.
Closes#87134
Salvage hardening on top of #87308 (thanks @Dudeman456):
- Tri-state verdict: inconclusive probes (binary missing, --help
failed/timed out) return None and fall through to the normal spawn
path, preserving the established 'Could not start Copilot ACP
command' error instead of masking it. This also fixes the two
test_copilot_acp_client HOME-env regressions that went red on the
PR: their mocked-Popen path was intercepted by the new unmocked
subprocess.run probe.
- Cache definitive verdicts per binary path so CLIs that DO support
--acp pay the ~50ms --help cost once per process, not per prompt.
- Skip the probe entirely when custom ACP args don't include --acp.
- Fix the help-text regex: the old pattern never matched '[--acp]'
(leading '[' is neither start-of-string nor whitespace) and \b
after 'p' matched '--acpfoo'.
- Hermeticity: stub subprocess.run in the two HOME-env tests; add 6
probe-specific tests (fast-fail, fall-through, caching, skip).
CopilotACPClient unconditionally passes [self._acp_command] +
self._acp_args (default ['--acp', '--stdio']) to subprocess.Popen.
When the resolved CLI doesn't accept --acp (e.g. Claude Code
v2.1.233, where 'claude --acp --stdio' exits 1 with
'error: unknown option') the subprocess dies in ~250ms with the
error on stderr, but the parent ACP loop has no fast-fail for this
shape and waits the full child_timeout_seconds (default 600s,
observed 109s+ before user interruption) for stdout that never
arrives.
Add _acp_supported() that probes the CLI's --help output for the
--acp flag in ~50ms, then call it at the top of _run_prompt before
any spawn happens. When the probe fails, raise a RuntimeError that
names the unsupported flag, lists the expected fix (install
@github/copilot late 2025+, or set HERMES_COPILOT_ACP_*), and
returns control to the caller in ~280ms instead of hanging the
delegate_task parent for hundreds of seconds.
Measured locally against Claude Code v2.1.233:
- Before: delegate_task acp_command=claude hangs 109s+ then
returns tokens={input:0, output:0}.
- After: delegate_task acp_command=claude raises RuntimeError
in 280ms with a clear actionable message.
This does NOT change behavior for supported CLIs (the new
@github/copilot ships with --acp) — the probe returns True and
the spawn proceeds unchanged.
Refs the bundled claude-review-delegate skill which already
documents this class of transport-mismatch pitfall for users
who call 'claude -p' directly; this fix closes the same gap for
the delegate_task MCP path.
Response to AI review (Enough1122) on #86812:
1. `_create_titled_session`: log the underlying exception before returning
None so programmatic callers aren't left with an undebuggable "could
not be created" — failures (DB lock, I/O, import) now land in errors.log
via logger.exception.
2. Drop the source-reading `TestSourceGuard` — it violated the repo's
"never read source code in tests" rule and was a change-detector
(passed even if behavior regressed). The stderr routing is already
covered by the real-path `test_missing_session_fails_on_stderr`.
3. Bare `-c` + `--create-if-missing` now prints a stderr note explaining
the flag needs a session name, instead of silently ignoring it — makes
the no-op self-evident to programmatic callers.
Tests: 6 targeted + 8 adjacent, all passing.
`hermes chat -c "<title>" -q "<text>"` silently no-oped when no session
matched the title under quiet/programmatic use: the not-found message was
written to stdout (the channel quiet callers parse as the final response),
so a background send to a not-yet-existing named session vanished with no
error. Surfaces via Hermes-Bot-Mode bot-to-bot handoffs (#86794).
- not-found message now goes to stderr (exit 1 unchanged), so programmatic
callers always see it even with -Q/--quiet
- new --create-if-missing: with `-c <title>` and no matching session, create
a fresh session carrying the title and proceed — the deterministic
"send to this named thread, making it if needed" primitive plugins asked for
- extract the -c resolution block into _resolve_continue_arg for testability
Tests: flag parsing, titled-session creation, stderr routing, source guard.
Streamed replies return None so the adapter does not send twice.
The /loop hook then saw empty text and never ran, so
awaiting_response stayed true and later ticks never fired.
Stash the delivered text on the event and use it for the post-turn
hooks. /goal uses the same path.
Tests: tests/gateway/test_loop_command.py
The stale-lock removal tests asserted the sweep result while leaving
_git_proc_running() live: on CI the parallel per-file runner almost
always has a real git subprocess in flight, pgrep -x git hits, and
clear_stale_git_locks correctly refuses to sweep — failing the tests
for reasons unrelated to the code under test (surfaced on PR #86918,
which doesn't touch gitlock at all).
Monkeypatch the guard to False in the sweep tests and add an explicit
test pinning the guard's block-while-git-running behavior.
The previous except Exception silently fell back to os.getenv if the
_scoped_key_env import ever failed — under multiplex that is exactly the
fail-open this PR removes (secondary profiles would regress to 'No LLM
provider configured' with zero trace). Catch only ImportError, log a
WARNING naming the consequence, and let any other failure propagate.
Also replaces the lambda fallback with a named nested function.
Three regression tests: scoped DEEPSEEK_API_KEY is visible to auto
detection under multiplex (the #86917 failure); unscoped paths keep the
os.environ read; explicit config provider still wins.
resolve_provider's auto path read provider API keys with bare
os.getenv — under multiplex a secondary profile's keys live only in its
secret scope, so auto-detection found nothing and every secondary
profile with model.provider: auto failed with 'No LLM provider
configured' at agent init (reproduced on a live 7-profile gateway).
Route both env-key reads (the OPENAI/OPENROUTER tier and the
PROVIDER_REGISTRY loop) through _scoped_key_env, the scope-aware helper
auxiliary_client already uses: secret scope wins under multiplex,
UnscopedSecretError falls back to os.environ (default-profile/CLI
paths unchanged). Same bug class as #86905.
Verified in a gateway-accurate simulation (hermes_home_override +
profile scope): resolve_provider('auto') now returns the secondary
profile's own provider (deepseek) instead of erroring.
Lesson from the feishu DM multiplex investigation: under multiplex,
os.environ holds the default profile's values, so any profile-level env
config (credentials AND authorization) must be read scope-aware, and a
scoped miss with a scope installed must fail closed instead of borrowing
from os.environ. The _get_scoped_secret wrapper is copy-pasted across
~15 platform adapters — new adapters and edits to existing ones must
keep the fail-closed semantics.
"\x00" in script_path raises TypeError when a caller passes a non-str
(e.g. a pathlib.Path, which is not iterable) — the guard itself would
crash the scheduler. All current call sites pass plain str, but the
guard must be crash-proof: str() first, then check. Adds a regression
test running a real script through _run_job_script with a Path argument,
which fails with TypeError on the pre-fix guard.
_run_job_script wrapped only expanduser() in its ingestion try/except.
On Linux an unexpandable NUL-bearing value raises inside that call, so it
landed on the clean fail-with-report path; on Windows expanduser() never
expands '~user' (and so never raises), and the NUL surfaces later as an
uncaught ValueError from resolve()/exists() — crashing the scheduler.
Align with cron.lifecycle_guard._expand_candidate_path, which already
documents this as the whole-class fix (the per-syscall catching produced
#76762, #77703, #77780, #78256): reject '\\x00' eagerly at the ingestion
boundary so both platforms fail identically.
The Bedrock API-key flow stored the bearer token in OPENAI_API_KEY and set
a bare `provider: custom`. Since #28660 that variable is only honoured for
openai.com hosts, so for bedrock-mantle.*.api.aws the token was dropped and
requests went out with api_key="no-key-required", a 401 on every call.
Write a named `providers.bedrock-mantle` entry with
key_env: AWS_BEARER_TOKEN_BEDROCK instead. The named-provider branch in
runtime_provider.py resolves key_env; the bare-custom branch cannot.
Fixes authentication only. Per-model mantle route selection is separate.
Appending --daemonized after ORIGINAL_ARGS put it past the `--`
relaunch-args separator on Linux, so it was absorbed into
RELAUNCH_ARGS instead of being parsed as a flag. HANDOFF_DAEMONIZED
never got set, so the one-shot self-detach block re-fired on every
re-exec -- an unbounded self-exec loop (thousands of iterations/sec,
100%+ CPU, argv growing until execve fails with E2BIG) whenever
relaunch args were present, which is the normal invocation shape on
Linux.
Fixes#86957
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.
Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.
Fixes#86661
- WARN when the venv site-packages layout is unresolvable and the script
falls back to plain PYTHONPATH execution, so 'editable installs
invisible' failures are diagnosable.
- Docstring: note that runpy does not set __package__/__spec__ the way a
direct python script.py invocation does.
_windows_cron_python_invocation bypasses the uv venv launcher (to avoid
flashing a console window) and re-attaches the venv via PYTHONPATH — but
PYTHONPATH entries are plain sys.path additions and never get .pth
processing, so editable installs (pip install -e) were invisible to cron
script jobs (ModuleNotFoundError).
Bootstrap the script with site.addsitedir() on the venv site-packages,
then exec it as __main__ via runpy.run_path, preserving the script
directory on sys.path (python script.py semantics). Falls back to a
plain invocation when the venv layout is unresolvable.
Commit 4c34eeb416 fixed dead Ctrl+C by removing the Kitty protocol push
(CSI >1u) from _EXTENDED_ENTER_KEYS_SEQ, keeping only modifyOtherKeys
level 2. That regressed kitty-the-terminal completely: kitty removed
xterm modifyOtherKeys support (kovidgoyal/kitty#4075) and only speaks
its own protocol, so after the removal kitty users lost Shift+Enter and
every other extended key — the CSI >4;2m we still pushed is a no-op
there (kitty even logs a PARSE ERROR for it).
The original reason for removing the push is obsolete: #87511 mapped
CSI-u control sequences, so Ctrl+C as ESC[99;5u now parses to
Keys.ControlC and fires the existing c-c binding. (The kernel-INTR
concern in that commit was moot — prompt_toolkit's raw mode clears
ISIG, so Ctrl+C is always handled by the binding, never the kernel.)
Restore the dual push (CSI >1u + CSI >4;2m), exactly mirroring the Ink
TUI, and complete the alias table for what the kitty disambiguate flag
actually emits — #87511 left real gaps, some of which its PR body
wrongly claimed were covered:
- Esc key: ESC[27u (+ modifiers) — previously leaked '[27u' as text
- Ctrl+Backspace -> backward-kill-word (#78285 was closed on the wrong
claim that codepoint-127 mapping existed; it did not)
- Shift+Space -> space (#86866's second symptom; the Ctrl+Space
mapping never covered modifier 2)
- Alt+Enter -> newline tuple; Shift+Tab -> BackTab; Ctrl+Tab -> Tab;
Alt/Shift+Backspace
- Multi-modifier letters (Shift+Alt 4, Ctrl+Shift 6, Ctrl+Alt 7,
Ctrl+Alt+Shift 8) normalized onto their Ctrl/Escape-prefix targets,
both unshifted (kitty) and shifted (mok emitters) codepoints
- Kitty PUA functional keys: keypad -> non-keypad equivalents,
F13-F24, and Ignore for lock/media/modifier-event keys so they are
consumed instead of leaking (kitty emits these even in legacy mode)
Also: clear the VT100 parser's prefix cache after installing (stale
answers could misparse), and re-push extended keys after
_recover_terminal_input_modes' reset — the recovery previously popped
both modes mid-session and never re-enabled them, silently killing
Shift+Enter until restart.
Refs #87511, #87074, #56684, #56645, #78285, #86866, #87390.
saveRegistryConnection only rewrote the registry file: editing a
connection's URL/token/host left pooled backend descriptors under
'conn:<id>::*' and open renderer sockets pointing at the OLD endpoint —
the UI showed the new target while traffic kept flowing to the old one
until idle-reap.
When a save MATERIALLY changes an existing connection (endpoint / auth /
ssh routing fields, via the new connectionDialFieldsChanged helper),
main now stops that connection's pooled backends and tunnels
(stopRegistryConnectionBackends, same teardown as removal) and
broadcasts 'hermes:connections:changed' with reason 'updated' so
renderers dispose and re-dial their secondaries at the new target.
Label-only renames do not recycle.
Removing a connection stopped its pooled backends and ssh tunnels but
never told the renderer: for remote/cloud sources there is no local
process to die, so the removed connection's WebSocket stayed open and
kept streaming ghost events into the UI until page reload. If the
socket did drop, openSecondary -> getConnectionFor threw 'No connection
with id' and scheduleReconnect retried forever (backoff caps at 15s,
entry never evicted).
- main now broadcasts 'hermes:connections:changed' on removal (and on
material edits); preload exposes connections.onChanged.
- use-gateway-boot subscribes and calls the new
disposeSecondariesForConnection(), which disposes + evicts every
secondary scoped to the connection id (redialing on edits).
- reconnectSecondary fail-stops: when the Electron main reports the
connection no longer exists, the entry is disposed and evicted
instead of retrying forever; ordinary transport errors keep the
existing backoff behavior.
Remote/cloud registry descriptors (entry.process === null) shared the
POOL_MAX_BACKENDS cap with real spawned local backends, so a roster
refresh across N registered remote connections could LRU-evict a live
local backend idle past the keepalive window. Cap accounting and
cap-driven eviction now count only entries with a live child process;
descriptors remain subject to the idle reaper.
Independent second-pass review found three gaps:
- focusedUsage is null | UsageStats, not Partial — ClientSessionState.usage
is the full type (app/types.ts) and its only write site seeds the four
required fields before merging (gateway-event.ts). The earlier Partial
annotation traced the wrong type (SessionRuntimeInfo, an RPC payload).
Comment now names the genuinely optional fields instead.
- The tile-focus contract test never seeded $sessionTiles/$sessionStates,
so it proved focusedStoredSessionId follows a tile but could not
distinguish focusedSessionId/focusedUsage working from broken. Seed a
bound runtime with distinct usage and assert both readout atoms move.
- The two public host.state references (website docs + bundled skill
reference) enumerated the old six atoms; plugin authors would never
discover the new ones. Both lists updated.
tsc/eslint/vitest green (4/4).
ClientSessionState.usage is Partial<UsageStats> (app/types.ts) — the
backend streams whichever fields changed — so the computed produces
ReadableAtom<Partial<UsageStats> | null>. Annotate the entry honestly
instead of claiming full UsageStats, and document the fallback rule for
plugin authors. Also collapse the three-argument expect() calls in the
contract test (vitest takes one message arg). Addresses triage review on
PR #80461.
Locks the plugin-facing contract: the focused atoms exist as readonly
nanostores, mirror the primary session while no tile is focused, project
the focused session's usage, and — the behavior this PR exists for —
follow the interacted tile while the primary-only $activeSessionId
stays put.
Disk plugins read app state exclusively through host.state, which only
exposed the primary workspace tab ($activeSessionId). In the multi-tile
layout, clicking a tile never touches that atom — and tile focus is a
pure renderer concern, invisible to both gateway RPC and the event
stream — so a plugin cannot follow the session the user is actually
looking at.
The core statusbar solves this same problem by reading the focused-
session atoms (use-statusbar-items.tsx). Widen the generic plugin
surface with the same signals, per the contribution rubric:
- host.state.focusedSessionId — runtime id of the focused session
(interacted tile, else the primary), the key for session.* RPC
- host.state.focusedStoredSessionId — durable id for navigation and
session-list matching
- host.state.focusedUsage — live streamed UsageStats projection
(context_used/max/percent, tokens, cost_usd), no RPC needed
Additive only; no existing behavior changes. tsc --noEmit clean.
Verified end-to-end with a disk plugin that now tracks the focused
session across tiles.
Gate composer submit and plugin host busy on the target session slice, not a leftover foreground busyRef. Staff can keep typing while a worker session is running.
Includes the follow-up test that submit uses the target session busy flag.
Add a compact Import popover to the MCP Capabilities page that accepts
anything a user might copy from an MCP server README and infers the
server config:
- mcp.json snippets (mcpServers-wrapped, bare name->config maps, single
unnamed server objects, Cursor/Claude `type` normalized to `transport`)
- bare npx/bunx/uvx/node/docker command lines (name inferred from the
package basename, e.g. server-filesystem -> filesystem)
- `claude mcp add NAME [--transport http|sse] [-e K=V] [-H ...] [--] CMD
ARGS...` and `claude mcp add NAME URL`
- bare http(s) URLs (name inferred from the hostname)
- Cursor deeplinks (cursor://anysphere.cursor-deeplink/mcp/install with
a base64-encoded JSON config payload)
The parser is a pure module (src/lib/mcp-import.ts) with unit tests for
every format plus garbage input. The popover previews the inferred
name + config and, on confirm, merges the entries into the editor draft
exactly like addServer's starter entry: unique keys, dirty (unsaved)
draft, first new block focused. Placeholder env values (YOUR_KEY,
TOKEN_HERE, ...) are kept verbatim for the user to edit in the editor
before saving.
i18n keys added under settings.mcp for en, zh, zh-hant, ja (ar falls
back through defineLocale).
Follow-up to the salvaged #87558 commit: the PR's docs promised the flags
follow "the focused chat", but PRIMARY_SESSION_VIEW is the primary
workspace tab only — a focused session TILE would read the wrong chat.
Wire host.state.busy / host.state.awaitingResponse through the focused
slice ($focusedStoredSessionId / $focusedSessionState), same semantics
as the statusbar busy pulse, with the primary view (and its draft
fallback) while the workspace holds focus.
Adds a tile-focus vitest case and corrects the docs wording.