`hermes gateway status`, `hermes gateway list`, `hermes profile list/show`
and the dashboard profiles payload keyed liveness off the profile's own
gateway.pid / gateway_state.json, so a satellite profile served by the
default multiplexer (gateway.multiplex_profiles) showed "not running"
even though the multiplexer is its live inbound process.
Reuse the single lookup the start guard and cron liveness already share —
named_profile_served_by_running_multiplexer() — with an optional
profile_name so list surfaces can ask about any profile, and OR it into
gateway_running for named profiles. Default profile and unserved named
profiles are unchanged.
Salvage of #69118 rebased onto the shared helper (which post-dates it).
Co-authored-by: Isaac Dobson <isaac@dobsonheadlights.com>
Co-authored-by: Mushisushi28 <133449918+Mushisushi28@users.noreply.github.com>
`asyncio.start_unix_server` does not exist on Windows (no AF_UNIX event-loop
support in asyncio), so arming the loop-tick witness in
`loop_heartbeat_forever` raised AttributeError on every native-Windows
gateway start. The broad except swallowed it and recorded
`loop_tick_socket=False`, so every stale-heartbeat probe classified the
gateway as UNKNOWN — never WEDGED, never ALIVE-with-stalled-write. The
two-witness interlock from a1c83ef9 (issue #90502 follow-up) has been
effectively disabled on Windows since it landed: a wedged native-Windows
gateway could never be detected, and an alive one could never be
distinguished from a stalled heartbeat write.
On non-POSIX platforms the witness now arms over a TCP loopback server on
127.0.0.1 (OS-assigned dynamic port) instead:
- same protocol — connect, read one byte "1"
- same semantics — pure in-memory, zero disk I/O, answered only while the
loop is dispatching, armed by the loop task itself (an awaited
`asyncio.start_server` is structurally loop-owned exactly like the Unix
variant, so a wedged loop cannot keep answering pings)
- the assigned port is published in the heartbeat payload as
`loop_tick_tcp_port`, and `probe_gateway_loop_liveness` prefers the TCP
witness when the producer published a port, falling back to the AF_UNIX
socket for POSIX/legacy producers
POSIX behavior is unchanged: the AF_UNIX arm (including the stale-node
sweep) stays gated behind `os.name == "posix"` so the missing attribute can
never raise on Windows again. Legacy heartbeats without `loop_tick_tcp_port`
keep the existing socket-node contract untouched.
Tested end-to-end on native Windows: witness arms, port is published,
`_probe_loop_tick_tcp` answers from an external thread while the loop
dispatches, and the existing loop-liveness suite passes unchanged (the
AF_UNIX structural test still passes — the Unix arm text is preserved
inside the POSIX branch).
Adds two tests pinning the new behavior: an E2E test that arms the TCP
witness and probes it (skipped on POSIX, where the Unix arm is the real
witness), and a structural test that the TCP arm stays awaited on the loop
task and the AF_UNIX arm stays POSIX-gated.
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.
Three surgical changes:
1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
The inlined watcher now routes the respawned gateway's stray
stdout/stderr to the same sidecar log gateway_windows._spawn_detached
uses (DEVNULL only as fallback), so a gateway killed moments after
respawn leaves a trace. Direct implementation of the 4th repro's
hardening suggestion (1).
2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
canonical _spawn_detached, so the respawned gateway's exit-diag /
lifecycle records show whether it escaped the parent Job Object — a
job-teardown kill is no longer indistinguishable from any other silent
death.
3. Post-update resume verifies liveness before vouching
(hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
runs the same provisional-hit + 2s-confirmation liveness poll every
other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
with all_profiles= for the fleet) before printing ✓, writes the #91675
start attestation for the verified PIDs, and fails the resume with a
"restart could not be verified" warning + recovery hint when no stable
gateway appears. Suggestion (2) of the 4th repro; closes the last
silent-success hole in the family (#84185 fixed the cold-start leg,
#91675 the direct-start leg; this is the relaunch leg).
Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.
Fixes the Bug-1 relaunch-trust leg of #48820.
/simplify-code quality+efficiency reviewers (converged, verified): on
the standard 'hermes gateway run' path the argv fast-path arms BEFORE
run_gateway's config bridge executes, and arm_startup_watchdog() is
idempotent — so gateway.startup_watchdog: false and
startup_watchdog_timeout_seconds were dead knobs (env bridged, live
handle untouched). run_gateway now applies the config to the live
handle: disarm on disable; disarm+re-arm on a bridged config timeout so
the fresh handle covers the remaining pre-loop startup with the
configured deadline. config_defaults comment updated to match reality.
E2E (real module, fast-path armed first): disable path disarms the live
handle; timeout path re-arms a fresh handle at 123s.
Review follow-ups on the salvaged #89750:
- gateway.startup_watchdog / gateway.startup_watchdog_timeout_seconds in
config_defaults, bridged to the internal HERMES_STARTUP_WATCHDOG env
vars in run_gateway() (the argv fast-path arms before config can load,
so env remains the mechanism; config.yaml is the user-facing surface
per policy — explicit env values still win as operator override).
- hermes_cli/main.py argv sniff now requires the ADJACENT token pair
'gateway run' instead of independent membership, so unrelated commands
mentioning both words can't arm a 300s hard-exit timer; profile-flagged
invocations (-p work gateway run) still arm.
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.
P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.
P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.
P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.
P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.
P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.
Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.
Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
A hosted gateway (hermes-doubleam-2568) deadlocked at startup with every
thread parked in futex_wait_queue before the asyncio loop came alive:
zero log lines, /health unreachable — but s6 saw a live PID so it never
respawned the process, and a stale gateway_state.json from the previous
life told every status surface "draining" for ~30 hours.
Every existing liveness backstop assumes startup succeeded: the
loop-liveness watchdog is armed inside the running loop's startup path,
the shutdown watchdog arms at stop(), and the heartbeat file is written
by an asyncio task. None can fire when the process wedges before the
loop exists.
New gateway/startup_watchdog.py: a plain daemon OS thread armed at
process entry (both gateway.run.main() and the `hermes gateway run` CLI
wrapper), disarmed the moment GatewayRunner confirms a live event loop —
the point where the existing loop-liveness watchdog takes over. If
startup neither reaches that milestone nor exits within the deadline
(default 300s; slowest legitimate pre-loop work is the 120s-bounded MCP
discovery wait), the watchdog:
* dumps all-thread stacks via faulthandler,
* appends a JSON record to logs/gateway-startup-watchdog.log,
* records the exit in the NS-608 lifecycle ledger
(reason=startup_liveness_watchdog) so the next boot classifies it
instead of reporting an unclean SIGKILL/OOM death,
* os._exit(75) so s6/systemd respawn the process.
Config is env-only (HERMES_STARTUP_WATCHDOG=0 to disable,
HERMES_STARTUP_WATCHDOG_TIMEOUT_S to tune, floor-clamped to 30s):
the watchdog must be armed before config.yaml is loaded — a wedge
during config parsing is exactly in scope — so it cannot depend on
config for its own enablement. Everything is best-effort; a watchdog
failure never affects the startup it observes.
Arm sites are placed after the PID-file/--replace conflict guards so a
--replace loser exiting early never arms a watchdog. Disarm happens even
when the loop guards are config-disabled (gateway.loop_watchdog: false)
— the startup watchdog only covers the pre-loop window, never adapter
connects or steady-state, so WhatsApp pairing / npm cold installs are
unaffected.
Tests: tests/gateway/test_startup_watchdog.py (29 tests — config
resolution, arm/disarm idempotency, fire path with captured exit,
lifecycle-ledger marking, dump record, disable knob).
Fixes OOF-298.
Generalize the HERMES_S6_SUPERVISED_CHILD supervisor-marker mechanism so
ANY supervised gateway launch (systemd, launchd, Windows Scheduled Task,
external supervisor) skips the active_profile redirect in
_apply_profile_override(). Previously only the s6 container marker was
honored, so a systemd-launched default-profile gateway with
HERMES_HOME=<root> followed the sticky active_profile file and silently
assumed another profile's identity — logging under that profile's tree
and connecting with its Telegram bot token (double-polling a token owned
by that profile's own live gateway).
- hermes_cli/main.py: honor HERMES_SUPERVISED_CHILD (new generalized
marker), HERMES_S6_SUPERVISED_CHILD (back-compat), INVOCATION_ID
(systemd; gateway commands only, since it leaks into every descendant
of systemd-launched processes), and HERMES_GATEWAY_EXTERNAL_SUPERVISOR.
- hermes_cli/gateway.py: export HERMES_SUPERVISED_CHILD=1 in generated
systemd units (user + system) and the launchd plist.
- hermes_cli/gateway_windows.py: export it from the Scheduled-Task cmd/vbs
launchers and the windowless respawn env overlay.
- hermes_cli/service_manager.py: export it alongside the s6 sentinel.
- tests: regression coverage for all markers + non-gateway INVOCATION_ID
neutrality + generated-unit marker presence.
Fixes#74872
Add _pid_record_belongs_to_current_profile() helper that verifies a
PID record's persisted hermes_home matches the current process. Use
it in get_running_pid() and get_runtime_status_running_pid() so the
default-profile gateway never mistakes another profile's gateway PID
as its own.
In _apply_profile_override(), clear HERMES_HOME instead of returning
early when it points to a profile directory but no --profile flag was
given, letting the sticky active_profile logic resolve the right one.
In _guard_existing_gateway_process_conflict(), detect stale PID files
from other profiles and emit a warning.
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):
- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
false-positives), and an explicit start-time expectation is now honored
on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
ProcessRegistry._terminate_host_pid (previously unverified), and the
session-close path now runs the same daemon identity verification as
the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
probes (real spawned processes, real psutil ancestry) wired into the
on-demand windows-latest wine2e lane.
Fixes#98814Fixes#89614
Re-applied onto 3aee29089 after `hermes update` reset main to origin/main.
1. web_routers/sessions.py: _resolve_session_id() classifies malformed-DB
errors via the existing is_malformed_db_error() and raises 503 at all five
call sites. delete_session_endpoint was the worst — an unresolvable id
counted as idempotent success, so DELETE reported it had removed a session
that was still on disk.
2. gateway/lifecycle_ledger.py: check_state_db_integrity() runs PRAGMA
quick_check(1) on the unclean-exit path only (~2s on 500MB) and records the
verdict into gateway-exit-diag.log. The 2026-08-31 corruption sat undetected
for 3.5 days because nothing ever looked.
3. hermes_cli/gateway.py: `gateway run --replace` gave the outgoing gateway 5s
before SIGKILL; SessionDB.close() runs a PASSIVE WAL checkpoint that does
not finish in 5s on a WAL 4x past the autocheckpoint threshold, and a kill
mid-checkpoint tears b-tree pages. Grace raised to 30s via a testable
_await_gateway_exit() that also re-checks after the final sleep (a PID
exiting in the last interval must not be SIGKILLed — PID-reuse hazard).
NOT added: wal_checkpoint(TRUNCATE) at shutdown — removed upstream in #45383
because a TRUNCATE reset races the live writer and tears b-tree pages.
Adversarial review: Codex gpt-5.6-sol, 9.0/10 across three groups, no must-fix.
/simplify-code efficiency reviewer (verified): cleanup_stale=False does
NOT deliver the exclusion the salvaged fix intended — get_running_pid
returns None whenever a record fails liveness VALIDATION (start-time
mismatch, argv drift, lock hiccup) regardless of the flag, which only
controls unlinking. In exactly the at-risk scenario the recorded PID
still never joined the exclusion set.
Exclusion evidence now comes from the RAW pidfile + lock records (no
validation, no unlink side effects); the validated non-destructive
probe is kept for the runtime-status fallback PID. For a KILL exclusion
list this is strictly safer: a stale recorded PID at worst spares one
process for one sweep, while a validation false-negative would
TerminateProcess a live gateway. Regression test reworked to drive the
real function semantics (raw record present, validation rejects);
mutation-checked: removing the raw-record read fails the test.
A named profile has no local gateway.pid, so cron warned that jobs
would not fire and recommended hermes gateway install — which the
start guard then refuses with exit 78. Share the multiplexer-serving
probe with the start guard and count it as liveness.
Repairs #98790 where
✓ Gateway is running — cron jobs will fire automatically
PID: 4165
Ticker heartbeat: 39s ago
4 active job(s)
Next run: 2026-08-30T22:50:18.762041+03:00 in profile B incorrectly reports
that jobs will fire based on profile A's gateway process.
Root causes:
1. executed ,
which enumerated the entire systemd fleet regardless of ,
violating the docstring "only PIDs belonging to the current profile".
2. checked when
(no heartbeat file) should trigger a warning — instead, it fell through
to the "✓ Gateway is running" green branch.
Changes:
- hermes_cli/gateway.py::_get_service_pids: pattern = get_service_name()
when all_profiles=False, filtering to the current profile's systemd unit.
- hermes_cli/cron.py::cron_status: guard hb_age is None first with an
explicit yellow warning: "ticker has not reported a heartbeat".
Regression test suite guards both systemd scoping (default + all_profiles)
and heartbeat branching (None vs fresh vs stale).
/simplify-code follow-ups on the 90502 salvage:
- _probe_loop_tick_socket_sustained: the two 'result is None' arms were
byte-identical — saw_node was effectively write-only. Collapsed to one
arm with one honest comment.
- loop_heartbeat_forever: sweep sibling gateway.loop-tick.*.sock nodes
from dead PIDs at arm time (POSIX-only liveness probe; Windows never
creates AF_UNIX nodes) so state/ does not accumulate nodes across
os._exit(75)/SIGKILL restarts. The reviewer's EADDRINUSE re-bind claim
was DISPROVED for this call site — asyncio's create_unix_server
os.remove()s an existing node before binding — but the contract is now
pinned by test_producer_rebinds_over_stale_socket_node (a live
producer arms and answers over a dead process's leftover node).
- test tmp_path fixture: yield + rmtree so the short-path mkdtemp no
longer leaks a directory per test run.
The two-witness contract from the first review round still granted
destructive authority on ONE silent 1s socket probe: stale heartbeat +
armed tick socket + a single miss returned WEDGED immediately, and the
#86860 consumers take the bounded SIGTERM/SIGKILL path on that verdict.
A short transient synchronous stall (reconnect storm, heavy synchronous
callback, scheduler delay) can outlast one recv timeout, so a lone miss
is exactly the false-wedge class this change exists to prevent.
WEDGED now requires the loop to stay silent across a sustained window:
tick_strikes consecutive misses (default 3, tick_gap_s apart). Any
answer inside the window proves the loop is dispatching and returns
ALIVE; a single miss returns UNKNOWN and keeps the graceful drain path
(which also preserves #86684's cron drain floor). A witness that
vanishes mid-window is ambiguity, never a wedge.
New regression coverage:
- unit: single silent probe recovers to ALIVE; sustained silence is
required for WEDGED; vanishing witness stays UNKNOWN.
- composed (real producer + consumer): heartbeat write stalled while
the loop is frozen for longer than one tick timeout but shorter than
the wedge window -> probe is ALIVE and launchd_restart drains, never
escalates; loop frozen for longer than the window -> WEDGED.
The default probe window is ~3.4s worst case, still far inside the 10s
subprocess query tier.
The off-loop heartbeat write broke the producer->consumer invariant #86860
depends on: file freshness no longer equals loop schedulability, yet the
probe still classified a stale file as WEDGED — and WEDGED is destructive
authority (SIGTERM -> SIGKILL, bypassing the #86684 cron drain floor). The
measured motivating stall (112.6s max) exceeds the 90s stale budget, so a
healthy loop blocked inside the watchdog's own write could be killed, and
executor saturation produces the same false positive. The inverse edge
also existed: an off-loop write landing after the loop froze refreshes the
file mtime, manufacturing a false-fresh liveness proof.
The gateway loop now also arms a loop-scheduling witness: a UNIX socket
(state/gateway.loop-tick.<pid>.sock) answered by the loop itself via
await asyncio.start_unix_server — socket-buffer writes, no fsync, no disk
I/O, so it keeps working on the filesystem that stalls the heartbeat
write. The heartbeat payload records whether the witness is armed
(loop_tick_socket).
The classifier is now two-witness:
- socket answers -> ALIVE (file age irrelevant: a stalled write
or saturated executor can no longer produce a wedge verdict)
- file fresh, socket silent -> UNKNOWN (a late off-loop write can no
longer manufacture a liveness proof)
- file stale, socket silent, producer armed -> WEDGED (both witnesses
agree the loop stopped scheduling)
- legacy payload (no flag) -> unchanged single-witness contract: the
legacy producer wrote on-loop, so staleness is still proof
- any conflict/ambiguity -> UNKNOWN, never escalate
Tests are a producer->consumer composition: a real heartbeat loop with a
stalled write probes ALIVE while the file is past the stale budget, and
launchd_restart fed by the real probe drains instead of escalating; a
silent socket with a fresh file denies ALIVE; WEDGED requires the armed
socket to agree; a bind-failed producer disables stale escalation; legacy
payloads keep the old contract; a source-inspection test pins that the
witness is awaited on the loop. Mutation-checked: reverting either source
file fails the new tests. 45 tests pass across the watchdog suites; ruff
clean.
The generated unit only counted restart_drain_timeout, so a default
cron drain (30s + 10s cleanup) could still be inside budget when
systemd SIGKILLed the cgroup. Size TimeoutStopSec from max(drain,
cron floor + reserve) plus headroom so an in-budget stop is not killed.
- launchd_restart resolves _launchd_domain() once (live launchctl probe,
up to 2x5s per call; two calls could also disagree)
- wedged-integration tests mock _wait_for_launchd_service_pid so the
observation poll doesn't burn 15s of real sleep per test (39s -> 16s)
- PEP8 blank lines in test_platform_base.py
A graceful SIGUSR1 exit alone doesn't prove supervision: detached-fallback
gateways (macOS 26 unsupported-domain marker) and unloaded jobs also exit
cleanly with nobody to revive them, and _graceful_restart_via_sigusr1
returns True for an already-gone PID — the CLI would print success while
the gateway stayed down. Poll _wait_for_launchd_service_pid (15s) after a
graceful exit and fall through to kickstart -k when no replacement
appears, mirroring systemd_restart's replacement observation. Adds the
no-replacement regression test and strengthens the budget assertion.
`hermes gateway restart` on macOS never took the graceful path, so every
restart — including deliberate ones — was reported to chat as an unplanned
shutdown.
`launchd_restart()` diverged from `systemd_restart()` in two ways, each
sufficient to break it on its own:
1. Wrong helper. It called `_request_gateway_self_restart()`, which is gated
on `_is_pid_ancestor_of_current_process()`. That holds only when the CLI
was spawned *by* the gateway (in-chat `/restart`). Invoked from a shell the
gateway is a sibling, so the guard returns False and SIGUSR1 is never sent.
`_graceful_restart_via_sigusr1()` — same job, no ancestry gate, already
used by `systemd_restart()` and the updater — had no launchd call site.
2. Wrong budget. It waited `_get_restart_drain_timeout()`, which defaults to
0, so `_wait_for_gateway_exit(timeout=0.0)` could never succeed. The
systemd branch uses `_get_restart_exit_wait_budget()`
(drain + after_turn + 15s headroom); `resolve_restart_exit_wait_budget()`
documents that callers falling back to a hard kill must cover both phases
or they reintroduce #77184.
The result was a bare SIGTERM followed immediately by `kickstart -k`. Since
SIGTERM leaves `restart_requested` False, the gateway exited 1 instead of 75
and announced "⚠️ Gateway shutting down — Your current task will be
interrupted." instead of "restarting", dropping the resume_pending handoff
that lets a session resume after the bounce.
Observed on macOS 27.0 / Hermes 0.20.4:
→ Stopping gateway (PID 49787) — draining in-flight runs (up to 0s)...
⚠ Gateway PID 49787 still running after 0.0s — restart may fail
⚠ Gateway drain timed out after 0s — forcing launchd restart
Send SIGUSR1 with the exit-wait budget and return on success, leaving
launchd's unconditional KeepAlive to revive the process. `kickstart -k` stays
as the fallback for a genuine drain timeout, but must not run after a
successful graceful exit or it would kill the replacement instance.
The wedged-loop escalation (#81642) still short-circuits ahead of this, so a
provably dead event loop is not handed a signal it cannot process.
Tests: adds a launchd counterpart to the existing systemd graceful-restart
test, asserting SIGUSR1 with the exit-wait budget and no bare SIGTERM or
kickstart on success. Updates the three wedged-gateway tests, which asserted
the old SIGTERM-plus-drain shape; they also now stub
`_graceful_restart_via_sigusr1` so no real signal escapes to the fake PID.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt branch from upstream/main f751a8c546 and re-applied the PR
changes. Resolved one conflict in tests/gateway/test_platform_base.py:
main had added TestDockerProfileSandboxMediaTranslation in the same
region — kept both main's new tests and the PR's
TestPlatformLockTakeoverGovernance regression suite.
Local: tests/gateway/test_platform_base.py +
tests/hermes_cli/test_gateway_service.py — 185 passed, 2 skipped;
ruff clean.
Refs: #79096
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.
This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:
- gateway/status.py: expose service-ownership discovery for gateway
runtimes (find_windows_gateway_services maps validated gateway PIDs
through process ancestry to running SCM service PIDs, with
create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
venv mutation and restart them afterward. Stops wait for a stable
SCM 'stopped' state AND for the original descendant processes to
exit (service 'Stopped' is not proof the child released its
handles). Failure to prove ownership, stop a service, or restart it
fails closed; rollback restores attempted services, and rollback
failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
or a service that will not reach a stable state abort the update
before any file mutation.
Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.
Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The terminal tool lifecycle guard and the gateway stop/restart CLI
guards keyed on the raw _HERMES_GATEWAY=1 env marker, which every
gateway descendant inherits (and importing gateway.run sets it too).
CLI/TUI agent sessions were falsely blocked from documented gateway
management commands. Gate on _is_supervised_gateway_process() instead,
which requires owning the live gateway PID file.
Salvaged from PR #92196 (guard half) by @nbxuhk. Fixes#92560.
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).
_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.
That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.
Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.
A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.
Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.
tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
Fixes#88654.
After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:
restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
if restart_mode is None:
continue
A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:
cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'
with no operator signal of any kind.
Two changes.
_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.
When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
hermes_cli/gateway.py's restart-wait sizing (from #92175) was the only
cross-module import of an underscore-private shutdown_forensics helper.
Promote it (private alias retained for existing patchers).
`_guard_named_profile_under_multiplexer` correctly refuses a named-profile
gateway while the default gateway is multiplexing — starting a second one would
double-bind that profile's platforms. The refusal is right; its exit code was
not.
The refusal is decided entirely by configuration (`multiplex_profiles` plus the
allowlist), so it is permanent: no number of retries can change the answer.
Exiting 1 made it look transient to a service manager.
That matters because this module generates the systemd unit, and the template
pairs `Restart=always` / `RestartSec=5` with `StartLimitIntervalSec=0` — it
deliberately trades systemd's generic start-rate limiter for the specific
`RestartPreventExitStatus=GATEWAY_FATAL_CONFIG_EXIT_CODE` backstop declared
three lines below it. Returning 1 left that backstop unarmed with the limiter
already disabled, so a correct, permanent refusal became an unbounded restart
loop. Observed on a host running `multiplex_profiles: true` with a leftover
per-profile unit: 136 refusals in ~13 minutes, stopped only by hand.
`GATEWAY_FATAL_CONFIG_EXIT_CODE` (78, EX_CONFIG) is this codebase's existing
answer for exactly this case — `gateway/restart.py` documents it as the fatal
configuration error that the s6 finish script translates into 125 "permanent
failure" (#51228). This adopts that contract rather than inventing one, so the
fix also works on s6 hosts, not just systemd.
After: one refusal, `status=78/CONFIG`, `NRestarts=0`, unit settles in `failed`.
Also strengthens the two guard tests. They asserted
`pytest.raises(SystemExit, match="1")`, but `match=` is a regex search over
`str(exc)`, so it passed for 1, 21, 100 and 111 alike — it read like an exit-code
assertion while pinning nothing. They now assert
`excinfo.value.code == GATEWAY_FATAL_CONFIG_EXIT_CODE`. The exit code is the
contract here: it is the only thing that tells a supervisor the failure is
permanent.
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:
- _restart_macos_launchd_gateways(): the invoking profile keeps the
existing launchd_restart() path; every other gateway of this install
is drained via SIGUSR1 (same as systemd siblings), then hard-
kickstarted unless KeepAlive already respawned it, then verified on a
fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
counts toward failed_or_stale_units — including timeouts during
liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
derives labels from THIS install's profiles (get_default_hermes_root),
not by globbing the shared per-user ~/Library/LaunchAgents — a
sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
installs) must never enumerate, let alone restart, another install's
fleet. This also keeps the hermetic test suite blind to a dev
machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
liveness, kickstart, and fresh-PID verification all use the domain the
service was actually located in (gui/<uid> vs user/<uid> probed per
label via `launchctl print`). This addresses the #41403 review defect:
the process-wide _launchd_domain() cache resolves the current profile's
domain and must never be reused for a sibling. _launchd_domain() itself
becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
sweep excludes every gateway service PID (mirror of the systemd
hermes-gateway* pattern) so it cannot mistake a freshly respawned
sibling service for a stale manual gateway. Default-scope callers
(gateway status, cron checks, stop_profile_gateway's orphan reaper —
which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
hints for launchd labels alongside the systemctl ones.
Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).
Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to the #74075 salvage: _reap-path _get_service_pids() call now
passes all_profiles=True. With the ps scan fixed, the reaper's process
scan surfaces sibling-profile launchd gateways on macOS; excluding only
the current profile's label would misclassify them as unsupervised
orphans and reap them (same class as the update-sweep sites the
contributor fixed). Also refresh the stale 'ps -A eww' comment.
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
on macOS/BSD ps, making the fallback silently return [] on every macOS
machine. The matcher only needs argv (not env vars), so e is
unnecessary. -ww keeps unlimited-width output on both BSD and
procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
on macOS, enumerate every ai.hermes.gateway* launchd agent across
profiles via bare launchctl list instead of only the current
profile's label. This prevents the update sweep from misclassifying
sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
_get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
pass all_profiles=True so the update fleet sweep excludes every
service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
launchctl list <label>, all_profiles uses bare launchctl list
with prefix filtering, handles empty/broken output gracefully, and
preserves systemd behavior.
Tranquil-Flow
Widens #90327's line_input() to the whole bug class: all 46 bare input()
free-text prompt sites across the setup wizards (model_setup_flows, setup,
config, gateway, auth, auth_commands, plugins_cmd, skills_hub, bundles,
setup_whatsapp_cloud, main) now route through line_input(), and the shared
cli_output.prompt() / setup.prompt() helpers do too — so every CLI wizard
gets cursor editing, not just the custom model prompt.
Redirected stdin and missing prompt_toolkit keep builtin input() behavior.
E2E: real PTY with raw escape bytes through line_input, cli_output.prompt,
and setup.prompt (arrows + Ctrl+A/E edit correctly); redirected-stdin
fallback verified.
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.
This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:
- the options
- the renderers for config.yaml, .env and the documents
- the activation body
- the command lines of the processes
nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.
`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.
The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.
The change also makes four corrections that apply to both modules:
- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
had only the gateway. But Hermes Desktop and the web dashboard connect
to a different process, so six of the configurations in public repos
add a second unit by hand. serve and dashboard are one entry point with
one flag of difference, and you can run only one of them. Thus the
option is an enum. The NixOS module asserts against container mode with
a backend, and does not make a unit that cannot start.
- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
installs into the working directory, which is correct for AGENTS.md but
wrong for SOUL.md and memories/. Hermes reads those files from
HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
made a workspace file that Hermes never loaded as the identity. The
documentation said this in prose, but two directory diagrams showed the
opposite. This change corrects both. A key in either option can now
contain subdirectories.
- `documents` needs an explicit `workingDirectory`. The default of that
option is bad on both modules. It is the home directory of the user on
Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
workspace files without a directory therefore gets a place that the
user did not select. The place is also different on each module. The
modules now refuse that combination.
The test is on the priority of the option and not on its value. An
option that nothing sets keeps the priority of its own default, and
each definition from a user is stronger. Thus a directory with the same
text as the default still counts as a selection, and so does a
mkDefault. A comparison of values detects neither case.
- Each activation writes .env again from a base in the Nix store, and
does not add to the file that exists. Thus a second activation cannot
put the same secret in the file two times, and a removed
environmentFile goes away. environmentFiles keeps the type `listOf
str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
the Nix store, which all users can read.
- HERMES_MANAGED and the .managed marker now hold the name of the system
that manages the install. Thus a refusal says "managed by home-manager"
and not "managed by NixOS", and `hermes update` gives the Nix guidance
for both shapes. The CLI does not print a rebuild command for each
system. It names the owner, and the user knows their own tool. A bare
`true` and an empty marker still mean NixOS, so this does not change an
existing install.
Verification. Six new checks, all built:
nixos-module evaluates the module with evalModules and the
NixOS module list. It asserts both units, one
HERMES_HOME, and that the module refuses
container mode with a backend.
home-manager-module evaluates the module with the
homeManagerConfiguration function of
home-manager. The process assertions run against
systemd units on Linux and launchd agents on
Darwin.
module-option-parity asserts that each shared option is on both
modules, and that the two exclusion lists name
only options that exist.
env-file-assembly runs the real .env script and checks the
contents, the mode, that a second run gives the
same bytes, and that a removed file goes away.
workspace-files-need-a-directory
checks that the module refuses `documents`
without a directory, and accepts a directory
that has the same text as the default.
service-argv runs each command line that the modules build
through the real parser of the CLI, with one
sentinel flag added, and requires that argparse
refuses only the sentinel.
`nix flake check` passes, with 21 checks in total.
The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.
Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:
- a lost --no-open
- a backend that runs the gateway
- an overwritten config.yaml
- documents in the wrong directory
- a different HERMES_HOME on the two processes
- a lost HERMES_HOME export
- a missing backend unit
- a removed assertion
- an .env file that grows at each activation
- an install that reports NixOS
- an empty .managed marker
- an option on the NixOS module only
- a stale entry in an exclusion list
- a renamed subcommand
- an unknown flag
- the workspace-files assertion always passes
- the assertion compares values instead of priorities
- an off-by-one that lets an untouched default through
- the assertion also fires for hermesHomeFiles
- a mkDefault no longer counts as a selection
- the Home Manager module stops wiring the assertion
- the NixOS module stops wiring the assertion
The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.
These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:
- test_git_probe_tree_kill.py (2 tests)
- test_update_import_guard.py (1 test)
- test_telegram_media_read_timeout.py (2 tests)
- test_teams.py (a collection error)
Closes#9056
# Conflicts:
# hermes_cli/main.py
# hermes_cli/update_cmd.py
# hermes_cli/web_server.py
subprocess.run(capture_output=True, timeout=N) is not hang-safe on
Windows: after the timeout fires, run()'s cleanup kills the direct child
and then joins the pipe reader threads with an UNBOUNDED communicate().
A descendant (conhost.exe under wmic/powershell) holding duplicated pipe
handles keeps the pipes from EOF and the join never returns.
_scan_gateway_pids() runs its wmic / Get-CimInstance Win32_Process scans
exactly that way, and on machines where the full process scan genuinely
exceeds its 10/15s budget (cold WMI on first boot, ARM VMs, heavy
Update/AV activity) hermes update wedged forever inside
_pause_windows_gateways_for_update() before printing a single line —
observed live on a fresh Windows 11 ARM64 VM with a faulthandler stack
pinning the main thread in subprocess._communicate and only a conhost.exe
child surviving. The single-flight update lock then blocks retries until
the wedged process is killed by hand.
This is the same deadlock class bounded_git_probe already fixed for git
probes (#68609 / #66037). Generalize that proven pattern into a shared
bounded_probe_run() — explicit communicate(timeout), kill_process_tree on
failure, bounded 1s drain, then abandon the daemonic readers — and
migrate the whole call-site class onto it:
- hermes_cli/gateway.py _scan_gateway_pids (the site that hung; reached
from hermes update, cron, gateway restart/status, dashboard)
- hermes_cli/dashboard_procs.py wmic scan (same shape, reached on update)
- hermes_cli/claw.py tasklist + PowerShell probes (same shape; its
try/except cannot catch a hang because a hang raises nothing)
- bounded_git_probe now delegates to bounded_probe_run (identical
contract, one copy of the cleanup logic)
Unlike bounded_git_probe, bounded_probe_run returns the CompletedProcess
(or None) rather than collapsing to stdout, because the gateway scan
branches on returncode to trip its wmic -> powershell fallback.
Tests: tests/hermes_cli/test_bounded_probe_run.py covers success,
nonzero-exit passthrough, spawn failure, bounded timeout (fails against
the old unbounded semantics — verified by sabotage), errors= decoding,
DEVNULL stdin, POSIX process-group placement, and the bounded_git_probe
delegation contract. Existing test_git_probe_tree_kill.py passes
unchanged against the delegated implementation.
Closes#87134
runuser/su/sudo -u from a root shell leaks XDG_RUNTIME_DIR=/run/user/0 into
the child. _user_systemd_socket_ready() stat-ed sockets under it with a bare
Path.exists(), which only suppresses ENOENT/ENOTDIR/EBADF/ELOOP — EACCES on
the 0700 root-owned dir escaped as a raw PermissionError traceback instead of
the documented UserSystemdUnavailableError remediation path.
- _path_exists_safe(): Path.exists() that treats EACCES as absent; used at
both the readiness and DBUS-detection call sites.
- _ensure_user_systemd_env(): drop an XDG_RUNTIME_DIR that is unset or owned
by another user in favour of our own /run/user/{uid}, so the restart
actually succeeds after su/sudo -u instead of only failing cleanly.
Regression tests cover the EACCES readiness probe, foreign-dir replacement,
and that preflight raises UserSystemdUnavailableError (not PermissionError).