The api_server platform wrote runtime status exactly once at bind (the
_connected mark) and never again: last_heartbeat stayed at boot time and
metrics_today froze at zero, so the dashboard showed stale API Server
activity until a full app restart (#52323).
The adapter now keeps daily request/message/token counters and a bounded
latency sample, publishes a metrics-bearing snapshot at bind, records
metrics after each completed _run_agent turn and /v1/runs run, and a
30-second heartbeat loop re-publishes while connected. gateway.status
gains a platform_metrics field on the platform payload, and
/health/detailed serves the live adapter metrics alongside the
persisted platform map.
Salvaged from #52345 by @itsflownium (Flownium) — reworked onto current
main (run-worker submission, bind-retry loop, readiness work counts).
Fixes#52323
Co-authored-by: Flownium <157689911+itsflownium@users.noreply.github.com>
* fix(update): stop reading gateway identity off the restart watcher's argv (#107002)
The detached restart watcher is spawned as
`python -c <watcher source> <old_pid> <python> -m hermes_cli.main gateway run`.
Its trailing argv is the command it will spawn LATER, but the canonical matchers
read identity straight off the joined command line, so the watcher itself was
classified as a live `gateway run` process — the documented "never infer process
identity from argv substrings" bug class, on the exact surface `hermes update`
uses to verify a post-update relaunch.
Also budget the post-relaunch liveness poll against the watcher's own deadline:
the watcher respawns the gateway only after the PID it was handed exits, so a
30 s window can expire before the relaunch it verifies was scheduled to start.
* test(windows-live): run the gateway-ancestor harness parent from a script file
A `python -c <src>` parent is an interpreter running inline source and carries
no readable Hermes identity, so it is no longer a gateway to any classifier —
the harness's own comment already said a realistic gateway argv is not a -c blob.
* test(windows-live): share one sleeper SCRIPT across the live process-topology fixtures
Four live Windows E2E files stood processes up as `python -c "sleep" <hermes argv tail>`.
That shape no longer carries a readable Hermes identity, so the fixtures stopped standing
in for the gateways they simulate. One shared sleeper script replaces the -c spelling.
* fix(gateway): drop the duplicated _INLINE_SOURCE_FLAG_RE definition
The constant was emitted twice around command_line_runs_inline_source. Same
pattern both times, so behaviour is unchanged — but one definition is enough.
* test(stderr-timestamp): run the gateway-lookalike children from script files
Both lookalikes stood a gateway child up as `python -c <src> <gateway tail>`.
That shape no longer carries a readable Hermes identity (#107002), so the
wrapper correctly stopped treating them as gateway spawns and the tests failed.
A `-c` tail is data for a program the inline source may spawn LATER, never the
child's own identity — the real wrapper child is `python -m hermes_cli.main
gateway run`, which has no `-c`. Running the stand-ins from a real script file
restores what the tests mean to assert without depending on the misread.
* test(windows-live): restore the tempfile import dropped with the local sleeper helper
* test(windows-live): wait on the sleeper SCRIPT name, not its source text
The live fixtures proved argv visibility by waiting for `time.sleep(120)` in the
spawned process's command line. That string only ever appeared there because the
sleeper was spelled `python -c "import time; time.sleep(120)"`; now that it runs
from a file the source is in the file, so the probe timed out ("sleeper argv
never visible") even though the argv was perfectly visible.
Wait on the script name instead, exported as SLEEPER_MARKER next to the script
so the probe and the spelling cannot drift apart again.
* fix(tests,gateway): keep the live-system guard blocking -c-wrapped gateway spawns
The #107002 identity fix made _gateway_command_subcommand return None for
'python -c <src> … -m hermes_cli.main gateway run'. tests/_fixtures/live_system_guard.py
shares that matcher, so the autouse guard stopped blocking the detached restart
watcher: real gateways leaked out of the e2e run and squatted the webhook port.
Add gateway.status.gateway_spawn_intent_subcommand — the spawn-intent mirror of the
identity matcher, peeling the inline-source wrapper token-wise and re-running the same
canonical matcher on each suffix (still no substring matching) — and point the guard at
it. Read-only subcommands stay spawnable.
* fix(gateway): make the inline-source option walk value-aware so -X utf8 -c is not read as a gateway
The walk that decides whether a command line is an interpreter running inline
source (`python -c <src> ...`) treated every token starting with `-` as a flag
and the first non-flag token as the end of the option block. CPython options
that take a SEPARATE operand (-X/-W/-Q, --check-hash-based-pycs, --jit) break
that model: the operand was mistaken for the end of the block, so the walk
never reached the -c behind it and the watcher was read as a live gateway
again -- exactly the #107002 misclassification, one shape further out.
- Reuse the canonical operand sets from hermes_state_holders rather than
hand-rolling a second copy (AGENTS.md: parser-derived flag sets).
- Walk case-preserving tokens: operand-taking -Q/-W/-X must not be conflated
with operand-less -q/-b, so callers no longer lowercase before the walk.
- Handle clustered short options precisely (-uc is inline source, -Xc is -X c).
- Replace the ad-hoc _INLINE_SOURCE_FLAG_RE rescan in
gateway_spawn_intent_subcommand with the index the same walk returns; the
regex could not find spellings the walk accepts and would raise
StopIteration.
Reported by an automated review on PR #121635 and reproduced here.
Refs #107002
---------
Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
Atomic Hermes' bundled desktop runner (desktop-gateway.py) shares
HERMES_HOME with the CLI but was not recognized by the gateway
command-line matcher, so `gateway run --replace` skipped the
terminate-and-scoped-lock-handoff path and collided with the desktop
runner's still-held scoped locks (e.g. the Discord bot-token lock),
leaving Discord responses down until the desktop gateway was killed by
hand. Recognize desktop-gateway.py as a gateway entrypoint in
_gateway_command_subcommand, covering both live process detection and
the PID-record metadata fallback.
Co-authored-by: namhyuk kim <happykimnh@icloud.com>
utf-8-sig exists to tolerate BOMs that Windows tooling adds to files
users edit. /proc and /sys files are generated by the Linux kernel, never
BOM'd and absent on Windows, so -sig there only muddies the read/write
policy. Switch every literal /proc/ and /sys/ read to utf-8 and teach
the footgun read rule that string literals starting with /proc/ or
/sys/ are exempt (user-edited files keep utf-8-sig).
Slim redo of the copied-profile-dir guard (#119772) at the seams that actually leaked on main:
- get_runtime_status_running_pid(expected_home=...) now rejects a record whose hermes_home
stamp names another home (rung 3 of resolve_gateway_liveness and every direct caller:
live_gateway_pid_for_home, the dashboard messaging/status readers, the update inventory).
Rung 1 already applied recorded_gateway_home_conflicts inside get_running_pid.
- A scoped read with no gateway_state.json hands rung 3 an empty record, never None -- None
re-read the PROCESS home's record and lent its live PID to the copied directory.
- multiplexer_liveness_for_profile only answers for <default root>/profiles/<name>: the roster
is matched by name and a same-named directory under another root is not the served home.
Two invariant tests replace the PR's four (same contract, real records instead of probe stubs).
uv writes no __pycache__ (pip does), so every module of a fresh generation
was compiled in the foreground of the first user request that imported it
(#100461; main's f298911467 / d380651a9f were lost with lazy_deps).
--compile-bytecode on the sync covers the whole install, transitive deps
included. Also drops the stray copy of that rationale that a merge left
inside gateway/status.py::_posix_is_zombie.
The cron guidance, the eager-activation guard and the launch-env freeze test all encoded
the pre-ruling behaviour: a `false` that disarmed the credential guard, and a "LEGACY
(pre-multiplex topology)" per-profile install the status output no longer offers.
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:
* `hermes -p served cron status` printed "Gateway is not running" and told the
user to run `hermes gateway install` / `gateway run` for that profile — i.e.
to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
checked systemd linger against the CURRENT profile's unit, so a doctor run
under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
which returns None for a served profile — the token-conflict warning was
SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
profile.
New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.
Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:
- `gateway run` for a profile the host gateway already serves ATTACHES: print
its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
to re-scan `profiles/` (control socket) and attach once the answer includes
it. Refuse only when the host gateway cannot be made to serve it. Under a
service supervisor the attach exits 78 instead of 0 so a redundant unit is
parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
the owner's control socket, so it no longer depends on the claim ordering in
start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
process: they restart the host multiplexer and preserve its served set. A
secondary still running its own gateway is reported with the
`gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
argv: a host singleton runs bare/default argv and can never prove it serves
profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
multiplexer is whichever profile launched the one host process.
Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
Review of the first cut:
- The Home Assistant errno hint matched a bare 65 and two error strings on
every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
genuinely unreachable HA host was told macOS was blocking launchd. Gate on
darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
`gateway run`, so `hermes gateway stop` would have signalled osascript
alongside the gateway. The canonical matcher now bails on an exact
argv[0] basename of osascript (the gateway is its child and is matched on
its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
carries (osascript's own output) instead of implying it is redundant.
main folded the auto-archive housekeeping into per-profile state.db maintenance; the plugin
update-check chore stays. The consolidated-away TestWorkdirParallelPool stays removed. Two new
utf-8 reads in gateway/host_rendezvous.py read utf-8-sig.
Reviewer findings on the host rendezvous record + serve attach.
- Attach now PROVES the owner before exiting 0: a bounded TCP connect to the
recorded endpoint plus a token-authenticated GET /api/host/identity that must
answer as the recorded pid+role. A supervisor or `hermes update` relaunch
landing in the old process's graceful-shutdown window (socket closed, atexit
not yet run) previously exited 0 with NOTHING listening, so the service
reported success for a dead backend. Anything short of a proven owner falls
through to the bind.
- Unprovable liveness (no psutil, an unexpected psutil error) is a CANDIDATE,
not an owner: it goes to the same probe instead of exiting 0, which is what
turned a record for a long-dead pid into a permanent silent outage.
- An explicitly typed --port/--host the owner cannot serve is a non-zero refusal
naming the owner, never a silent loopback redirect; and a `hermes dashboard`
user is never routed to a headless `serve` backend (servesSpa in the identity
answer).
- SIGTERM — the NORMAL stop (systemd stop, docker stop, the update relaunch) —
now clears the record and the 0600 token and releases the host lock. atexit
does not run on it: uvicorn's capture_signals re-raises into the default
disposition, so a live session token outlived its process indefinitely. The
handler only prepends cleanup and hands off to the previous handler, leaving
the shutdown sequence unchanged.
- Host lock claim is tri-state (acquired / held-by-other / could-not-open) and
logs the OSError: an unwritable lock dir used to be reported as "another
gateway owns this host", sending operators hunting a process that never
existed. Its handle cache is keyed by (role, resolved lock path), so a changed
lock dir can no longer make owns_host_lock() lie.
- Windows: the token is written through the SSH runtime's protected
owner+SYSTEM DACL writer (os.open(0o600) sets no ACLs there, and os.replace
fails against an open reader); read_token's docstring no longer claims the
mode bits prove same-OS-user.
- gateway/status: a RELATIVE $XDG_STATE_HOME is ignored per the XDG spec — it
made the host lock dir CWD-relative, so two serves started from different
directories shared no singleton.
Live A/B (real processes): kill -TERM of a fully started `hermes serve` left
host-serve.json + host-serve.token on disk on base, removes both on head (exit
status -15 unchanged). A record with a LIVE pid and a closed port made base exit
0 with no bind; head binds. `serve --port 8899` against an owner on another port
exits 1 naming the owner.
Multiplex-only (Teknium ruling): exactly ONE `hermes serve` and ONE
`hermes gateway run` per host, each multiplexing every profile. The existing
gateway lock/PID files are anchored to HERMES_HOME, so N profiles were N
independently acquirable locks and `hermes serve` had no singleton at all.
- `gateway/host_rendezvous.py`: a host lock + a published record (pid,
creation time, port, protocolVersion, tokenFingerprint, served-profile set)
under the one cross-profile lock root already in the tree. A record whose PID
is dead, or alive with a different creation time (PID reuse), is stale and is
never attached to. Removed on clean exit.
- Gateway: takes the host lock ALONGSIDE the per-home lock and publishes the
record. Observe-only — a second gateway still starts, and logs the owner.
- Serve: publishes the record plus a 0600 token file on bind, and a second
`hermes serve` for ANY profile discovers it, prints the live pid/port and
exits 0 instead of binding a second port. The bare-TCP `_dashboard_listening`
probe (which proved only that *something* answers) is replaced by
record-based discovery with creation-time proof; the unified re-exec stays as
the no-record fallback. `--isolated` and HERMES_DESKTOP=1 opt out.
`spawn-ledger.json` keeps being written unchanged, so the merged Desktop attach
ladder (#117983) keeps working.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
`_read_process_cmdline` tried `/proc`, then `ps`, then psutil. macOS and BSD have no `/proc`, so
every call there forked `ps -p <pid> -o command=` while psutil — a pinned core dependency, which
pyproject calls "the canonical answer" for PID questions — sat unused behind it, returning the same
string in-process.
That call is the gateway-identity check guarding against PID reuse, reached for every profile with
a live gateway, and `list_profiles()` runs it per profile. `list_profiles` is the shared body of
`GET /api/profiles` and `profiles.list`, which the Bots roster polls every 5s per connection, so
the forks repeat forever on an idle machine.
Measured on macOS: one call 4.17ms via `ps` against 0.022ms via psutil. End to end against a home
whose profiles each hold a live gateway lock, A/B in one process state, two rounds: 4 gateways
23.9ms -> 4.0ms, 8 gateways 49.0ms -> 9.3ms.
`ps` stays as the fallback rather than being replaced: on macOS psutil raises AccessDenied for a
process owned by another user, which `ps` still reports — real when a dashboard probes root-owned
LaunchDaemon gateways. That path is unchanged and simply pays a cheap failed lookup first. For a
readable process both sources return byte-identical strings, so the command-line matchers see no
difference, and Linux still answers from /proc without reaching either.
Regressions: psutil answers without any fork (a raising `subprocess.run` proves it); the `ps`
fallback still answers when psutil refuses, as it does across users; a dead PID reports nothing
from either; and both sources agree on a live gateway-shaped process. Restoring the old order
fails the first.
Fixes#117270
_pid_exists() ran psutil.Process(pid).status() (~7 ms) unconditionally before the
Windows branch — once per registry entry inside the unfair exclusive session file
lock, starving every 2 Hz poller at >=7 leases (#115578). Windows has no zombie
state, so the probe is pure cost there; pid_exists()/ctypes decide instead.
active_session_registry_snapshot() read AND pruned under _FileLock. It now snapshots
raw entries under the lock, probes liveness after release, and re-locks only to drop
the lease ids proven dead, so a lease acquired in between survives.
Tests trimmed from the contributor branch to two invariants + a POSIX control; the
Windows half is proven on the real runner via the wine2e red/green receipt.
Salvages #115591
hermes_home_assignments parses a space-joined command line, so an unquoted
`HERMES_HOME=C:\Users\John Doe\.hermes` token was cut to `c:/users/john` and the
default-profile predicate (assignments non-empty, own home absent) judged the
gateway NOT to belong to its own home -- a regression from the substring test
this PR replaced. Add command_line_names_hermes_home: the token-bounded parse
first, then a token-bounded literal match of the whole home, used by both
gateway.status and the hermes_cli PID scan so the two predicates stay mirrors.
Review follow-up (#115039): a supervisor that exports HERMES_HOME with a
trailing separator (systemd Environment=, sh -c wrappers) put a spelling on
the command line the token-bounded comparison rejected, so gateway status
reported that gateway not-running for its own home. Strip one trailing
separator on both the extracted values and the profile home, keeping the
longer-sibling rejection (ops/2 is still a different home).
Also bound the assignment NAME: FOO=hermes_home=/x embeds the name inside
another token and is not a HERMES_HOME assignment (the value side already
was token-bounded; the substring test this replaces matched it).
Drive _scan_gateway_pids in a test so the mirrored predicates in
gateway.status and hermes_cli.gateway cannot drift apart again.
The named-profile branch of _command_line_belongs_to_profile tested
`hermes_home={home}` as a substring of the command line, so a process
declaring HERMES_HOME=/root/profiles/ops2 satisfied the ops profile's
predicate (the -p side already uses token equality for exactly this
reason). The default-home branch and the CLI mirror
hermes_cli.gateway._matches_current_profile had the same hole.
Extract every HERMES_HOME=<value> assignment token-bounded with quotes
stripped (hermes_home_assignments) and compare full values in all three
places. Quoted assignments (ps/wmic re-quoting paths with spaces) now
match too.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
_prepare_runtime_status_update deep-copied the module snapshot into the new
payload, deep-copied that again as previous_payload, and deep-copied the
result back into the module snapshot; submit() then copies once more. The
module snapshot is only ever reassigned (never mutated in place), and the
transition emitter only reads previous_payload, so hand out the old snapshot
as previous_payload and store the new payload directly.
Both callers of _get_runtime_status_writer() already hold
_runtime_status_state_lock (an RLock); the flush paths read the module
attribute directly and never initialise. A second lock plus double-checked
init protected nothing, so the getter serialises on the state lock instead.
The 18-keyword signature was copy-pasted onto write_runtime_status and
publish_runtime_status and forwarded one by one. Keep the explicit
keyword-only signature on _prepare_runtime_status_update (typo safety at the
boundary) and let the two public wrappers forward **fields. No positional
callers exist (all call sites use keywords).
flush_runtime_status_async re-implemented _RuntimeStatusWriter.flush() with a
hand-rolled relay thread parked on the writer Condition plus a future and
call_soon_threadsafe/shield/wait_for. flush() already returns the same
True/False outcome at its deadline, so run it under asyncio.to_thread.
With the relay gone, wait() loops on settled() instead of re-inlining the
same predicate.
The background writer persisted the in-memory canonical snapshot verbatim. origin/main re-read gateway_state.json on every write, so keys stamped by other processes survived; with the snapshot they were clobbered on the next publish. Out-of-process writers exist: hermes gateway migrate --standalone (_reconcile_standalone_runtime) rewrites the live multiplexer's file, and container_boot seeds desired_state/migrated_from. Re-read the file in the writer thread immediately before writing and lay the gateway's fields over it: the gateway wins for every key it owns, foreign keys survive.
write_runtime_status(wait_timeout=, reload_existing=) are part of the
call contract, not private state: drop the leading underscores at the
definition and every call site. flush_runtime_status_async no longer
polls the writer every 20 ms; a daemon waiter parked on the writer's
Condition relays the settled generation into a loop future, so the loop
wakes on the event and a wedged write past the timeout cannot hold
interpreter exit.
Gate review: `degraded` became a whole-life serving state but four
sibling predicates only accepted `running`: derive_gateway_busy /
derive_gateway_drainable (the NAS drain gate reported a degraded gateway
with in-flight turns as idle and undrainable), the stale-heartbeat
detectors in `hermes gateway status` and the dashboard, and the Windows
doctor probe. All four accept `degraded`; a dead watchdog-stamped
`degraded` is still excluded by `gateway_running=False`.
The parked-platform ERROR pointed at `/platform resume`, which only
resumes platforms in the retry queue - a non-retryable failure never
enters it. The remedy is `hermes gateway restart`.
Test: a degraded gateway with active agents is busy; not live => not
drainable.
python -m gateway.run never passes through the CLI's
normalize_hermes_home_env(), so the raw Path(os.environ['HERMES_HOME'])
readers for PID/lock/status, lifecycle ledger and heartbeat files kept a
literal ~ relative. Fallbacks are unchanged; the forensics marker probe
expands inline to stay import-light.
The reporter's case in #113372 is "not a crash": the process stays alive,
`gateway_state.json` keeps saying `running`, and housekeeping, cron and the
kanban dispatcher are frozen. Nothing wrote `updated_at` periodically, so the
file was a stored status, not a heartbeat, and both `hermes gateway status`
and `/api/status` rendered a wedged gateway as healthy (the existing stale arm
only fires when the recorded PID is gone).
Make the housekeeping tick re-stamp `gateway_state.json` first thing every
tick (60 s), so `updated_at` is a heartbeat that stops when the thread — or a
chore blocked on the loop — wedges. Readers then warn on `running`/`starting`
+ stale stamp + live PID: `hermes gateway status` prints
`⚠ Gateway heartbeat stale: housekeeping has not refreshed gateway_state.json
for N s (event loop or housekeeping wedged; pid X alive)`, `/api/status`
carries `gateway_heartbeat_stale_s` (null when healthy) and the sidebar strip
shows "Heartbeat stale". Liveness (`gateway_running`, busy/drainable) still
keys off the PID, never the stamp. Draining is excluded: shutdown drains the
housekeeping thread before the process exits, so its stamp legitimately ages.
Invariant tests: the CLI line names the age and the live PID and is silent on
a fresh stamp; /api/status sets/clears `gateway_heartbeat_stale_s` the same way.
`/api/status` mapped a not-running gateway's retained record through
`retained_gateway_state`, which only kept `startup_failed`; a watchdog-stamped
`degraded` + `exit_reason` of a dead PID became a bare `stopped` with the
reason nulled, so the sidebar strip and System page read "Stopped" while
`hermes gateway status` said "exited degraded: event loop stopped dispatching".
Retain `degraded` under the same rule as `startup_failed` (only while
`desired_state` still wants the gateway running, and only for the watchdog
reasons in the new `gateway.status.WATCHDOG_EXIT_REASONS`); the resolver
already keeps `gateway_exit_reason` for any non-`stopped` verdict. The sidebar
strip gains the `degraded` label (warning while live, destructive when the
process is gone) and the System page describes a dead `degraded` record as a
watchdog exit. `/api/messaging/platforms` keeps yielding `gateway_stopped` for
it — the channels really are down.
Invariant tests: retained_gateway_state keeps/drops the verdict by
desired_state and exit_reason; /api/status carries degraded + exit_reason for
the dead PID and stopped/null after `hermes gateway stop`.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
`/api/messaging/platforms` still read a not-running gateway's retained
`gateway_state == "startup_failed"` verbatim, so a profile the operator
stopped with `hermes gateway stop` (which keeps the last failure and its
`exit_reason` on disk for diagnostics under `desired_state: stopped`) showed
every configured platform as `startup_failed` with the stale reason ("Port
8642 already in use") — the same symptom #112517 describes, on the sibling
surface `/api/status` had just been fixed for. The two docstrings promise
the sidebar strip and the Channels page "can never disagree on one page
load"; they did.
Lift the verdict into one shared `gateway.status.retained_gateway_state()`
used by both routers: `startup_failed` only while `desired_state` still
wants the gateway running, else `stopped`. Behaviour of `/api/status` is
unchanged (same condition, one call site); messaging now yields
`gateway_stopped` with no error for the operator-stopped case and keeps
`startup_failed` + `exit_reason` when the profile is meant to be running.
Invariant test `test_operator_stopped_gateway_does_not_report_retained_startup_failure`
(red on the previous head: state='startup_failed',
error_message='Port 8642 already in use').
Fixes#112517
The boot guard logged its refusal once and nothing else knew. A user whose default says multiplex
but whose gateway serves one profile would read `hermes gateway status` and see a healthy gateway.
The verdict now lands in gateway_state.json (multiplex_standalone_reason) and status prints it with
both remedies; a boot that does multiplex clears it.
`_flag_reconnect_needs_attention` stamps needs_attention=True +
retrying_since when a reconnect loop passes the attention threshold, but
only `_install_reconnected_adapter` (the watcher's own success path)
cleared them. Every OTHER writer of `connected` — the startup stamp in
`_start_connect_pending`, `BasePlatformAdapter._mark_connected`, Telegram's
in-place polling recovery (`send_path_degraded` -> healthy) — left the
flags in place. A gateway restarted after an escalation therefore reported
telegram as `connected, needs_attention: true, retrying_since: 2026-08-30`
for two weeks on a healthy bot.
Fix at the single seam: `write_runtime_status(platform_state="connected")`
now defaults needs_attention=False / retrying_since=None unless the caller
passed them explicitly, so all four writers agree without each growing a
copy of the clear. Discord's connect() (cherry-picked from #102557 by
@sudhirpatil) calls `_mark_connected()` instead of a bare `_running = True`
so its stale `fatal` stamp clears through the same seam (#102554).
Both default-profile process matchers (`gateway.status._command_line_belongs_to_profile`
and `hermes_cli.gateway._scan_gateway_pids._matches_current_profile`) rejected a named
gateway with a substring test for `--profile ` / ` -p `, which the equals spelling the
CLI pre-parser accepts (`--profile=ops`) slipped past. The default home's identity check
then adopted that gateway's PID, and a default-profile `gateway stop` with no pid file
scanned the process table and could SIGTERM the named gateway (review of #108352,
finding E). Both sites now ask `profile_flag_value()`, the same tokenizer the named
branch already uses.
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.
One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
Under gateway.multiplex_profiles a secondary's api_server and webhook are never built as
adapters (run_adapters skips SHARED_LISTENER_MIRROR_PLATFORMS: the default's listener answers
/p/<profile>/...). The multiplexer record therefore has no `<profile>:api_server` entry,
profile_platforms_from_multiplexer() returned {} for them and both /api/messaging/platforms
and /api/status?profile= fell through to `pending_restart`: the Desktop Messaging card and
Command Center said "Restart needed" forever for a platform that was answering.
- gateway.status.shared_listener_mirror_platforms projects the default's LIVE api_server /
webhook entry onto every served secondary with `ingress_url` = `<listener>/p/<profile>/v1`
(`.../webhooks/<route>`); a dead default listener is not mirrored. The api_server / webhook
adapters stamp the listener they actually bound (`listener_base`) on connect so the URL is
the real one, not a config guess. `hermes status` lists those URLs beside the other
shared-ingress platforms.
- /api/status?profile= reports `gateway_shared_with` (every profile the multiplexer carries)
when the served rung answered; null for a standalone gateway.
- Desktop: the messaging card shows the URL line; "Restart gateway" from a served profile
(statusbar menu, Cmd+K, messaging/webhooks banners, Command Center) confirms "Restart the
shared gateway? All bots on this device reconnect: default, alpha, beta" (Restart all /
Cancel) and toasts "Shared gateway restarted (3 bots)". Standalone keeps the silent path.
- Dashboard: same confirm + toast on the System page and the sidebar restart; the 409 from
start/stop on a served profile renders as an inline notice instead of a raw error toast.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.