The hubs are navbar-only. On a phone Docusaurus opens the drawer on the
doc sidebar and puts the navbar behind 'Back to main menu', so mobile
readers cannot find Skills or Plugins at all. Two top-level sidebar
links make them one tap away on every doc page (desktop sidebar too).
The startup reclaim caps the log a previous run left behind, but a shell
that stays up for days writing Chromium ERRORs is unbounded until it
restarts — the case @ehz0ah flagged in review.
Chromium owns that descriptor in append mode for the life of the
process, so rotation is the wrong primitive: renaming leaves the writer
on the renamed inode and the cap silently stops applying. Poll the live
size and truncate in place instead; an O_APPEND writer resumes at offset
0, so a noisy process stays bounded at ~cap plus one interval's output.
The timer is unref'd and every tick is best-effort.
Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>
Refs #100573
Two follow-ups to #117851, both raised in review by @ehz0ah:
1. The diagnostics block created HERMES_HOME/logs with an unguarded
module-level mkdirSync, before app readiness. A read-only or invalid
logs path terminated the Linux desktop at startup — optional
diagnostics killing the app they exist to diagnose. Wiring now runs
through enableLinuxCrashDiagnostics(), where every step is
best-effort: no writable logs dir degrades to "no Chromium log" (and
keeps the crash reporter, which writes elsewhere), and a Crashpad
handler that refuses to start is not a startup failure either.
2. Electron opens an explicit --log-file with APPEND_TO_OLD_LOG_FILE, so
desktop-chromium.log accumulated ERROR/FATAL output across launches
with no bound. desktop.log is capped at 10 MiB x 3 backups precisely
because it once reached ~326 GB and exhausted the disk; the new file
now goes through the same planner before Chromium appends to it.
Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>
Refs #100573
desktop.log's cap (10 MiB x 3 backups, plus the discard ceiling that
reclaims a boot-loop log outright instead of renaming it to .1) lived
inline in main.ts, keyed to one hardcoded path, and main.ts cannot be
imported by a test.
Move the planner to its own module, parameterized by base path, so the
Chromium diagnostic log added for #100573 gets the same bound and the
behaviour is provable without booting Electron.
Refs #100573
Rebasing onto main turned 21 tests red in four files whenever the operator supplies
HERMES_GATEWAY_LOCK_DIR. Not a behaviour collision: the record this PR introduces lives in
the per-OS-USER host state dir, and the root conftest pins that dir per test only when the
caller supplied nothing (#118097 deliberately keeps the documented override working). With
one supplied, every test in a file shares the dir, so a case that arms the obligation makes
the next case read a restart it never owed.
Clearing the record around each hermes_cli test — rather than re-pinning the dir — fixes the
leak without touching #118097's resolution rule or any expectation.
Review findings on the host-scoped update→restart obligation.
- update_cmd_fleet: an unwritable host state dir (read-only HERMES_GATEWAY_LOCK_DIR,
container UID that does not own $HOME) made write_host_obligation return False and the
caller ignored it, so an interrupted update left ZERO obligation — stale code, no
warning, no catch-up restart. The return is now propagated: the legacy per-home marker
(still read by every reader) carries it, and a host that can write neither says so.
- update_cmd_fleet::_obligation_fields: a PRESENT but unparseable/foreign-versioned host
record no longer falls through to the legacy marker; terms nobody can read cannot be
discharged by another record's terms.
- update_cmd_fleet::_restart_identity_sha: zip/pip/Docker installs resolve no checkout
SHA, so the restart-once stamp was "" and could never match — every profile's update
re-killed the one shared multiplexer. Falls back to the record's expected_sha, then the
receipt's post-update identity.
- update_host_obligation: any main_pid probe error is unproven identity (keep its own
restart), never an aborted restart pass.
- run_notifications: the online notice dedupes per home CHAT, so two served profiles
sharing one chat get one message (accounting stays per profile); transport resolution is
isolated per profile, so one broken adapter no longer starves the rest of the fan-out.
- run_adapters / run_profile_reconcile: _profile_configs is pruned with the served set, so
a failed or removed profile no longer owes a notice nothing can deliver and
.restart_pending.json is unlinked.
Tests cover the new format's own hazards: unwritable record dir, foreign-version record
with a legacy marker present, non-git install, shared home chat, broken adapter, pruned
config, plus a parity test for the duplicated host-state-dir resolver.
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
The session-bus socket was created lazily and never closed. main.ts
already tears down its pooled keep-alive sockets in will-quit so nothing
holds the event loop open or leaks an fd past app teardown; the new
long-lived socket joins that policy. Disposal goes through the existing
close path, so live notifications are failed the same way a daemon crash
fails them.
The transport was picked from process.platform, so the ipc test mocked
the Linux module and the only vitest lane (ubuntu) never exercised the
Electron Notification branch that Windows and macOS run, while the Linux
tests skipped everywhere else. Make the platform an injectable host
parameter: the ipc test drives the Electron branch as darwin, the Linux
tests pass linux explicitly and lose their skipIf.
Every failure in show() armed the global 10 s cooldown, including the
ones caused by the notification daemon being replaced mid-call: the
owner-changed fences, and the bus's NameHasNoOwner error for a call
already addressed to the vanished unique name. Linux has no Electron
fallback, so every notification in that window was silently dropped
exactly when the new daemon was healthy.
Classify by generation rather than by error: a failure against an owner
that is still current is the daemon's fault and cools down; one against
an owner that has since been replaced does not. The fixture now answers
calls to a stale unique name with NameHasNoOwner like the real bus, and
the race test asserts the next notification reaches the new owner
immediately.
`finished` cannot be observed inside show(): close() aborts the signal
first, and dbus-native settles an aborted invoke synchronously, so the
awaiting code throws instead of resuming past the check. `connection !==
state` is implied by the generation check (disconnect bumps the
generation before clearing the connection), `connection !== target.state`
in close() is implied by `delivered` (disconnect fails every live entry,
which clears it), and receive() can never run after release() deleted its
map entry.
A malformed Notify id now gets its own error instead of being reported as
an owner change.
Every Chromium CHECK/LOG(FATAL) traps at the same instruction (the
ImmediateCrash tail of logging::LogMessage::HandleFatal), so the cores
collected for #100573 all share one address and none say which check
fired. The message goes to stderr, which the .desktop entry and the
Omarchy wrapper both discard, and the app never started Crashpad.
On Linux, route Chromium's log (ERROR and above) to
HERMES_HOME/logs/desktop-chromium.log and keep local minidumps; nothing
is uploaded. Child processes inherit the switches, so zygote/GPU checks
land in the same file.
Refs #100573
Review findings on this PR.
- `_from_host_record` accepted a record with `createTime: null`: `_same_incarnation` treats a
missing create_time as "matches", so `liveness_is_proven()` blessed whatever process happens to
own that PID today. A record pointing at an unrelated `sleep 60` made `cron status` print
"the host gateway (PID ...) serving profiles default, served". Require a recorded create_time.
- `cron status`'s host-record rung printed `hermes gateway restart`, which exits 78 for exactly
its audience (a served NAMED profile); use the `--profile default` form both rungs now share.
- doctor's WORST state (supervision slots exist, nothing owns the gateway role) emitted a warning
but appended no issue, while the strictly less severe LEGACY case did.
- Drop three defensive wrappers in `host_topology`: a try/except around an always-importable
`normalize_profile_name` that silently degraded to a DIFFERENT normalisation, a blanket except
per rung, and a swallowed `get_active_profile_name`.
- `get_systemd_unit_path`, `_legacy_unit_search_paths`, the doctor host-unit probe and the
profile-delete service cleanup hardcoded `~/.config/systemd/user`, so on a host that moves
$XDG_CONFIG_HOME the unit probe false-negatived -- the very bug this PR fixes survived there.
One `user_systemd_unit_dir()` helper honours the spec default.
Tests: the two "Scheduler host: the host gateway" assertions were prefixes of BOTH rungs, so the
rungs were indistinguishable and the host-record rung was never exercised (which is how the
restart-command defect shipped green). Both now assert the full line, plus a served-host-record
case. Rendezvous isolation uses `tmp_path` instead of a leaked `tempfile.mkdtemp()` and also
covers `test_gateway_multiplex_served_record.py`, which read the real host record; the
`tests/conftest.py` hook in #118097 supersedes all of them once it lands.
``host_rendezvous``'s contract is that a record whose (pid, createTime) cannot
be POSITIVELY matched is a candidate, never an owner. Accepting an unprovable
record turned a leftover record into a permanent "gateway running" answer on
every reporting surface. Require ``liveness_is_proven`` before treating the
record as the host gateway, with an invariant test.
Two existing tests that simulate "no gateway anywhere" now point
HERMES_GATEWAY_LOCK_DIR at an empty dir, so an unrelated host record on the
machine cannot make the profile look served.
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:
* `hermes -p served cron status` printed "Gateway is not running" and told the
user to run `hermes gateway install` / `gateway run` for that profile — i.e.
to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
checked systemd linger against the CURRENT profile's unit, so a doctor run
under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
which returns None for a served profile — the token-conflict warning was
SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
profile.
New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.
Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:
- `gateway run` for a profile the host gateway already serves ATTACHES: print
its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
to re-scan `profiles/` (control socket) and attach once the answer includes
it. Refuse only when the host gateway cannot be made to serve it. Under a
service supervisor the attach exits 78 instead of 0 so a redundant unit is
parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
the owner's control socket, so it no longer depends on the claim ordering in
start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
process: they restart the host multiplexer and preserve its served set. A
secondary still running its own gateway is reported with the
`gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
argv: a host singleton runs bare/default argv and can never prove it serves
profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
multiplexer is whichever profile launched the one host process.
Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
Review of the first cut:
- The Home Assistant errno hint matched a bare 65 and two error strings on
every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
genuinely unreachable HA host was told macOS was blocking launchd. Gate on
darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
`gateway run`, so `hermes gateway stop` would have signalled osascript
alongside the gateway. The canonical matcher now bails on an exact
argv[0] basename of osascript (the gateway is its child and is matched on
its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
carries (osascript's own output) instead of implying it is redundant.
macOS Local Network Privacy attributes a socket to the executable launchd
spawned for the job. A bare venv Python has no application identity and is
not platform-entitled, so every LAN connect from the supervised gateway
failed with errno 65 (No route to host) while the same URL worked from
Terminal — and nehelper never showed a prompt that could grant it, so a
headless Mac had no way out.
Run the job through /usr/bin/osascript: `do shell script "exec …"` spawns
its child as osascript-responsible — an Apple platform binary — and the
child is exempt. Verified live on macOS 26.3.1 with a fresh ad-hoc binary
under the gui launchd domain: bare → 65, `/bin/sh -c exec` → 65,
`/usr/bin/time` → 65, osascript → reachable; an ad-hoc-signed helper .app
with NSLocalNetworkUsageDescription (the #115196 approach) stays denied with
no prompt even after lsregister, matching the #57812 dead-end table.
`do shell script` buffers the child's output until exit, so the shell
command redirects stdout/stderr to the log files the plist already routes;
`exec` keeps the gateway in the job's process group, so `launchctl bootout`
still delivers SIGTERM to it (live: stop → no orphan, restart → exit 75 →
KeepAlive respawn, stop → parked). The existing plist-staleness refresh
picks the new definition up on `hermes gateway install`/`start`.
Mechanism proposed by @leewaiho in #57812.
Co-authored-by: leewaiho <18321182+leewaiho@users.noreply.github.com>
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".
Salvaged from #115196 (remedy text points at the regenerated launchd job).
A hydrated page can start mid-turn: the first row is the tool-only
assistant, the second its tool result, and the answer arrives third with
no active bubble to append to. The bubble that fold creates stands for
all three backend rows, but only the two pending rows were counted
(serverRowSpan: 2), so the release rewound the older-page offset short
of what it gave up and 'Show earlier' skipped rows the reader could no
longer reach.
Count the current answer row too, and pin it with a regression whose
page starts on the tool-only assistant/result pair.
Found in review of 72a5aa3fd2 by @ehz0ah.
The hydration fold merges a turn's tool rows into the bubble they belong
to, so a ChatMessage is not one backend row (a ten-tool-call turn is one
message over eleven rows). The older-page offset is counted in backend
rows, so rewinding it by the store's message count skipped history: the
next 'Show earlier' page started past rows the reader could never reach.
Carry the count out of the fold as ChatMessage.serverRowSpan and rewind
by it, which makes the rewind exact instead of approximate. Drop the
'has a rowId' release guard for the same reason it was wrong before: a
folded bubble can have no single durable id while every row behind it is
persisted and re-fetchable; what must not be released is a row still in
flight (pending).
The backend pages display history (include_compacted reads group by display
order; inactive rows are not counted), so a locally counted row total is not
guaranteed to be the offset's unit. Decrement the offset the backend reported by
the rows the store gave up, clamped at zero: a mismatch can only overlap a page
already in memory (which the merge dedupes), where an over-counted absolute
offset would skip rows the reader could then never reach.
Three copies of the tail-entry resolution (record the page, rewind, read the
state) now go through one resolveTailEntry; the retention planner reports
released:false without walking the weights or counting persisted rows, and the
hook checks for a fetch route before planning instead of discarding the plan
afterwards.
The transcript window bounds what reaches assistant-ui, but the session store
underneath keeps every message it ever materialized: the tail hydration, every
older page "Show earlier" fetched, and every turn the window streamed. Renderer
retention therefore grows with session content and never comes back down — the
remaining half of #77311 after #117681.
`boundRetainedTranscript` keeps the live window plus one window page of slack
(so the first "Show earlier" still pages through memory) and releases the rows
older than that, which are off-screen and already persisted. Two invariants
keep the release safe: nothing is released without a durable `rowId` (a row the
backend cannot serve again would be content lost), and a branch group is never
split at the cut, for the same reason the window cut does not.
`rewindTranscriptTail` points the session's tail bookkeeping at the retained
prefix, so the released rows stay reachable: the existing older-page backfill
fetches them back on the next "Show earlier", and the rewind is refused outright
when the session has no recorded page route to fetch from.
`useTranscriptRetention` runs on the window's own cut moves — streaming grows
the tail, "Show earlier" prepends, a re-cut moves the anchor — so the weight walk
never happens per token. It computes the cut inside the store updater, from the
array it is about to replace: a view snapshot can be a flush behind the store,
and trimming a stale copy back into it would drop the newest rows of a live turn.
The first cut appended the recovery only in `_tcc_row`, which is built by the
0.10 fallback probes. Drivers that serve health_report (0.22+, the ones a
stale row actually bites) render their own tcc_* rows untouched, so `doctor`
never showed it. Apply the hint at the report seam like the display-count
guard so both paths get it.
Also: one CUA_DRIVER_BUNDLE_ID (permissions.py is the leaf; the daemon
imports it), the CLI derives the field set from the same table, and the
unverified "--upgrade repairs stale rows" claim is dropped from hint + docs
(the user hitting this is already on a current driver).
`hermes computer-use permissions status` and the doctor's fallback `tcc_*`
rows told a user whose System Settings toggle already showed CuaDriver ON to
"grant it in System Settings" — the one step that cannot help. macOS keys the
TCC row to the app's code-signing requirement; a row written for an earlier
CuaDriver build stops matching after a driver update and flipping the toggle
does not rewrite it (trycua/cua#3170, repaired by the cua-driver >= 0.22
installer on update). The daemon then reports Accessibility / Screen Recording
false while the pane shows ON, and every click silently no-ops (#99732).
`permissions.stale_tcc_grant_hint(*missing)` renders the reset for exactly
the missing services (`tccutil reset Accessibility|ScreenCapture
com.trycua.driver`, then `hermes computer-use permissions grant`); both
surfaces append it only when a grant is actually reported False. Docs: the
computer-use page still said bounded/unrestricted daemons run "under the
Hermes host identity" — they launch through CuaDriver.app since #95381 — and
gains the stale-row troubleshooting entry.
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
Each per-profile pass could spend 15s against a 5s caller budget, so on a host with
several served profiles the trailing wildcard pass — the only one that stops the shared
MCP loop and its connections — might never run. shutdown_mcp_servers takes a timeout and
the caller gives each pass budget/(N+1).
The worker also ran under copy_context(), carrying a caller's HERMES_HOME override into
a pass that documents it has none: from inside profiles/poison's scope the wildcard pass
resolved live_home to that profile. It now starts from a fresh Context().
Both call sites replace `with suppress(Exception)` with a logged WARNING: the suppress
swallowed a real TypeError from this signature change, which surfaced only as a missing
teardown in an unrelated test.
_default_spawn gated the scrub on is_multiplex_active(), so on a single-profile host —
the common Kanban deployment — build_subprocess_env never reached _sanitize_subprocess_env
and profile B's worker env was byte-identical to the dispatcher's, OPENAI_API_KEY and
systemd-injected tokens included. The authority test is "is this worker acting for a
ROUTED home", exactly as served_profile_child_env decides it, not the gateway-wide flag.
Drops install_profile_terminal_scope from the bind_home=False branch: that path has no
terminal-scope consumer and it cost a config.yaml parse per spawn. The test now asserts
the worker ENV rather than that a scope object was bound.
Three defects in the eager activation gate:
- profiles_to_serve() takes the multiplex flag as an ARGUMENT and never reads config, so
any host with >=2 profile dirs armed the fail-closed guard at boot — including hosts
that deliberately pinned gateway.multiplex_profiles: false. The gate now reads
GATEWAY_MULTIPLEX_PROFILES then config.yaml and skips activation when it is explicitly
false; the lazy backstop is unchanged.
- The snapshot was taken at the TOP of start_server, before keepalive / auth gate /
uvicorn build. It is the ONLY source for launch keys with no .env to rebuild from
(systemd Environment=, op run, Compose), so a credential injected or rotated by a later
boot step was invisible for the process lifetime. Activation moves to the last boot step.
- A failing probe returned False silently: the guard never armed and the caller's
diagnostic was dead code. It now logs at WARNING with exc_info and fails CLOSED — an
unreadable profiles/ cannot prove the host is single-profile.
Also: a profile dir holding only an EMPTY .env is what a crashed `hermes profile create`
leaves behind, and counting it flipped a single-profile host fail-closed at its next
boot. The gate counts named_profile_has_servable_identity instead.
_session_profile_runtime_scope hydrates a profile's EXTERNAL secret sources at entry —
it shells out to the operator's secret command (op run / bws / sh -c) with a 30s CLI
budget while holding the process-global source-cache lock. The exit-flush worker has a
5s TOTAL budget, so with a slow source it flushed 0 transcripts where base flushed 1,
and the loss was silent: both callers discard the return value. The periodic reaper
tick paid the same cost per session per tick.
Persisting a transcript needs no external credential: pass hydrate_secrets=False (the
flag this branch already uses in run_profile_reconcile), and WARN when the exit flush
persists fewer transcripts than it found.
Docstrings corrected while here: the transcript ROWS were never at risk (a profile
session owns a SessionDB whose db_path is frozen at construction); the real defect is
the config / token-ledger / provider resolution that _persist_session does at call time.
_profile_home_or_none answered None both for "this body is the launch profile's own
work" and for "this NAMED profile's home no longer resolves" (deleted or renamed
mid-run, invalid name, transient OSError), and the handler factories captured that
answer once. With the launch-profile fallback added earlier on this branch, an inbound
message on secondary B's own bot was then handled with the LAUNCH profile's .env and
frozen env — a silent cross-tenant credential fallback where origin/main raised.
_routed_profile_home now returns an UNRESOLVED_PROFILE_HOME sentinel and logs a
WARNING; that case binds nothing, so an unscoped get_secret raises UnscopedSecretError
under multiplexing. None keeps exactly one meaning: launch-owned.
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:
- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
now enter the SESSION's profile scope, the same chokepoint _finalize_session
already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
profile's scope (mirrors startup discovery and the reconcile chore), with a
trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
eviction and state/memory handle release now run inside the deleted
profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
profile no longer means "no scope". launch_profile_scope_if_multiplexed()
binds the launch profile once the process multiplexes; before activation it
is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
assignee's secret AND terminal scope for toolset resolution and the spawn-env
build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
when the host has more than one servable profile home, instead of lazily on
the first ?profile= request after earlier work already ran unscoped.
Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
Independent review found the previous commit could make the unbounded-growth
symptom it fixes PERMANENT, and that its argv narrowing re-opened #92401 inside
a single install.
- other_generations_for_path() counts only LIVE generations. A retired one is
already write-fenced (StateDbReplacedError, close-time checkpoint disabled)
and leaves the registry only when its holder releases -- which a gateway
handle does not do before shutdown. One inode replacement (repair swap,
backup restore, snapshot) therefore skipped auto-VACUUM for that path for the
whole process lifetime.
- _argv_scoped_to_other_home is ranked evidence now. argv[0] is the SHARED
install binary for every profile on a host, so it is neutral, never proof of a
hold; the process's own --hermes-home / HERMES_HOME= / --profile / -p
selection decides which home it serves, and a token naming ANOTHER profile's
store is tested before any own-prefix check. argv[0] also stops dismissing a
holder of a store whose home is not part of an install layout (a custom
HERMES_HOME is served BY the binary under ~/.hermes).
- install_root comes from hermes_constants.named_profile_home, not
basename(parent) == "profiles": an arbitrary <X>/profiles/<n>/ tree no longer
promotes all of <X> to "ours".
- The launch profile keeps its gateway.sessions_dir override when its own store
is pruned; every other served profile prunes under its own <home>/sessions.
Pruning under the wrong dir orphaned transcripts forever.
- glob.escape on the request_dump_<id>_* sweep (pre-existing).
Tests: the retired-generation test asserted the starvation mechanism; it is
replaced by the invariant (a live sibling defers VACUUM) plus a red-on-base test
that a retired, write-fenced generation does not. New argv cases cover the
shared binary with -p other, another profile's store token, and a non-Hermes
<X>/profiles/ tree; the housekeeping fixture now asserts the unpinned store
still resolves inside the sandbox before yielding.
Three silent data-correctness bugs on a host where ONE process serves every
profile:
- The gateway constructor ran auto-archive and auto-prune/VACUUM once, on a
handle pinned to the construction-time launch home, so a served secondary
profile's state.db was never pruned or vacuumed by anybody. Both now run per
SERVED profile, inside that profile's runtime scope, against its own
SessionDB, its own `sessions:` config and its own transcript dir.
- The dashboard/`hermes serve` auto-archive sweep archived an arbitrary
profile's store but read the config through the PROCESS HERMES_HOME, so one
profile's sessions.auto_archive_days governed every other profile's
retention. It now loads the config of the home whose store it sweeps.
- Auto-VACUUM admission read "no FOREIGN holder" as "the store is quiet", but
the /proc scan skips our own pid by design. VACUUM plus its TRUNCATE
checkpoint therefore retired a WAL generation another live SessionDB in THIS
process still held. The path-keyed registry now answers the in-process half.
Also: argv naming only the install root no longer dismisses a multiplexer as
"another instance" when the store is a profile store beneath that root — under
one-process-per-host that argv is exactly what the holder looks like.
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.
`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.
Also in this pass:
- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
the gate the serve/Desktop ticker already passes). On a host pinned to
per-profile gateways both processes raced that profile's tick lock, and when
the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
instead of the profile's live adapters. The gate compares the liveness PID
against `os.getpid()`: this process holds the launch `gateway.pid` and
publishes every served profile in `served_profiles`, so a bare liveness answer
would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
host-wide bare-id union, which let profile A's running `daily-brief` report
profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
ticked set; pools lived until `atexit`, so every home ever ticked kept a
ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
`_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
fails behaviourally on base instead of as a collection error.
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.
- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
(`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
covered every profile ticked. It now asks `scheduler_ownership`:
`owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
home) and `live_gateway_ticking(home)` (another live host gateway whose
published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
`_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
`_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
`_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
`daily-brief` no longer read as one job. The public accessors still report
the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
per-profile key, and the single global pool was sized by whichever profile
ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
`gateway.multiplex_profiles`: that flag gates adapters, and with it off every
non-launch profile's jobs sat in a store no ticker visited.
Reviewer findings on the host rendezvous record + serve attach.
- Attach now PROVES the owner before exiting 0: a bounded TCP connect to the
recorded endpoint plus a token-authenticated GET /api/host/identity that must
answer as the recorded pid+role. A supervisor or `hermes update` relaunch
landing in the old process's graceful-shutdown window (socket closed, atexit
not yet run) previously exited 0 with NOTHING listening, so the service
reported success for a dead backend. Anything short of a proven owner falls
through to the bind.
- Unprovable liveness (no psutil, an unexpected psutil error) is a CANDIDATE,
not an owner: it goes to the same probe instead of exiting 0, which is what
turned a record for a long-dead pid into a permanent silent outage.
- An explicitly typed --port/--host the owner cannot serve is a non-zero refusal
naming the owner, never a silent loopback redirect; and a `hermes dashboard`
user is never routed to a headless `serve` backend (servesSpa in the identity
answer).
- SIGTERM — the NORMAL stop (systemd stop, docker stop, the update relaunch) —
now clears the record and the 0600 token and releases the host lock. atexit
does not run on it: uvicorn's capture_signals re-raises into the default
disposition, so a live session token outlived its process indefinitely. The
handler only prepends cleanup and hands off to the previous handler, leaving
the shutdown sequence unchanged.
- Host lock claim is tri-state (acquired / held-by-other / could-not-open) and
logs the OSError: an unwritable lock dir used to be reported as "another
gateway owns this host", sending operators hunting a process that never
existed. Its handle cache is keyed by (role, resolved lock path), so a changed
lock dir can no longer make owns_host_lock() lie.
- Windows: the token is written through the SSH runtime's protected
owner+SYSTEM DACL writer (os.open(0o600) sets no ACLs there, and os.replace
fails against an open reader); read_token's docstring no longer claims the
mode bits prove same-OS-user.
- gateway/status: a RELATIVE $XDG_STATE_HOME is ignored per the XDG spec — it
made the host lock dir CWD-relative, so two serves started from different
directories shared no singleton.
Live A/B (real processes): kill -TERM of a fully started `hermes serve` left
host-serve.json + host-serve.token on disk on base, removes both on head (exit
status -15 unchanged). A record with a LIVE pid and a closed port made base exit
0 with no bind; head binds. `serve --port 8899` against an owner on another port
exits 1 naming the owner.
Red on base: `gateway.host_rendezvous` does not exist there, and the real-process
A/B shows base binding two ports for two profiles where head binds one.
- two HERMES_HOMEs each take their OWN per-home gateway lock (unchanged) but
only the first takes the host lock;
- a live record ends a second `hermes serve` at exit 0 before the bind;
- a record whose creation time does not match its live PID is a recycled PID
and is ignored; `--isolated` never attaches.
Multiplex-only (Teknium ruling): exactly ONE `hermes serve` and ONE
`hermes gateway run` per host, each multiplexing every profile. The existing
gateway lock/PID files are anchored to HERMES_HOME, so N profiles were N
independently acquirable locks and `hermes serve` had no singleton at all.
- `gateway/host_rendezvous.py`: a host lock + a published record (pid,
creation time, port, protocolVersion, tokenFingerprint, served-profile set)
under the one cross-profile lock root already in the tree. A record whose PID
is dead, or alive with a different creation time (PID reuse), is stale and is
never attached to. Removed on clean exit.
- Gateway: takes the host lock ALONGSIDE the per-home lock and publishes the
record. Observe-only — a second gateway still starts, and logs the owner.
- Serve: publishes the record plus a 0600 token file on bind, and a second
`hermes serve` for ANY profile discovers it, prints the live pid/port and
exits 0 instead of binding a second port. The bare-TCP `_dashboard_listening`
probe (which proved only that *something* answers) is replaced by
record-based discovery with creation-time proof; the unified re-exec stays as
the no-record fallback. `--isolated` and HERMES_DESKTOP=1 opt out.
`spawn-ledger.json` keeps being written unchanged, so the merged Desktop attach
ladder (#117983) keeps working.
`catalog_mapping` took an optional `target_profile` and fell back to the launch
process's `HERMES_PROFILE` (then "default"), and `execution_policy_mapping`
loaded whatever gateway home was active while labelling the result with the
requested `target_profile`. On a multiplexed / Desktop-spawned backend a remote
Bot invited for profile B could receive a catalog signed for the launch
profile, or B's name over A's approvals/turn-limit/toolsets.
- `target_profile` is a required keyword on `catalog_mapping`; an empty or
mismatched profile fails loudly naming the field, no env fallback.
- `execution_policy_mapping(config=None)` resolves the SERVED profile's own
config: a no-op when it is the active home (callers that already scope are
unchanged), else `_profile_runtime_scope(get_profile_dir(profile))`; a profile
that does not exist is refused instead of silently reading the launch config.
- `issue_room_grant`'s default digest inherits the served-profile resolution.
- Tests: 27 call sites now name the profile; two invariants red on base.
- Docs: multi-profile resolution table row for the RoomLink catalog/policy.
Fixes#116900
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:
* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
loop) and the success marker; with no success marker on disk it printed
"✓ Gateway is running — cron jobs will fire automatically", and with a stale one
it pointed at "Check the gateway log" instead of naming the cause. The persisted
`CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
and reported as "Gateway is running STALE code — fires NOTHING", with both
revisions and the restart command. A fresh heartbeat with a recorded error and no
success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
to the existing drain-first `request_restart` path (SIGUSR1) via
`hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
contract the restart phase already uses for unmapped manual gateways. The drain
budget computation is shared (`_gateway_drain_budget`).
The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).
Fixes#117275