Follow-up to the salvaged restart-wait commits:
- A drain or cron timeout of .inf now means "wait indefinitely" instead of
an OverflowError from the integer stop envelope, which crashed
`hermes gateway restart` and made `hermes update` silently fall back to
its 45s floor. The fleet "draining (up to Ns)" lines and the drain
progress report format the budget instead of int()-ing it, so an
unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
envelope and that configured cron reaches the CLI wait, instead of
re-deriving the formula; the negative wording assertion on the pending
footer is dropped (change-detector).
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.
Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
The fleet-restart-pending discharge parsed the marker inventory and
probed a gateway-less host inline (CC 36). _marker_owed_gateways now
turns the inventory into the owed set (raising ValueError, which the
caller's existing handler already maps to "keep the marker"), and
_discharge_gatewayless_marker owns the #118742 host probe. Every
verdict is unchanged; the caller is at CC 21.
- tools/browser_tool_install.py: keep pm-clean's frozen old-updater stub; main's
UTF-8 decode fix touched only the npx prefetch body it replaces.
- tests/hermes_cli/test_update_scoped_reconciliation.py: keep pm-clean's test
subset (catch-up rides the PM completion owner) and take main's gateway-less
host evidence (#120740): the updated seed that holds the host at a running
gateway, and the two gateway-less matrices for the source change that merged
cleanly into update_cmd_fleet.py.
* refactor(update): one predicate for serve rows outside the gateway matrix
The inventory branch of `_marker_only_restart_obsolete` inlined the rule for which
serve/dashboard rows the gateway matrix neither covers nor needs to (supervisor-owned,
or a manual serve handed to its own reminder). The inventory-less branch needs the same
rule for #118742, so it moves to `update_cmd_fleet_gatewayless.runtime_outside_gateway_evidence`
and both branches will read one definition. No behaviour change.
* fix(update): a Desktop-only host settles an inventory-less restart obligation
A host that runs no gateway (the Desktop app alone) can be left with an inventory-less
fleet-restart obligation: an updater that died before recording its inventory, or the
pre-inventory writer. With no owed set, the live gateway matrix is its only evidence, and
on that host the matrix is empty forever, so `_marker_only_restart_obsolete` never settled
and every later `hermes update` exited 1 with "gateways are still off the checkout code"
(#118742).
An empty fleet alone cannot tell that host from one whose gateway the dying update stopped,
so the inventory-less branch now asks the live host, never a historical receipt:
`host_owes_no_gateway_restart` settles only when no profile's `gateway_state.json` claims a
state other than stopped/startup_failed (a gateway that went away without a clean stop keeps
the obligation) and every live runtime sits outside the gateway matrix
(`runtime_outside_gateway_evidence`, shared with the inventory branch). HEAD must still
contain the pulled SHA (`checkout_contains`, same rule as #119367). Probe failures keep it.
Tests: the scoped-reconciliation matrix now holds its host at "the update stopped a gateway"
so it keeps pinning receipt independence; a new host-evidence matrix covers Desktop-only,
clean stop, carried commit, stopped gateway in a named profile, unclassified and
unidentified serves, a gateway row without fleet identity, and a diverged checkout. Two
manual-serve tests that assumed an empty fleet always stays pending now stub the live host
and assert the manual reminder survives the gateway obligation settling.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
---------
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
`_marker_only_restart_obsolete` bailed out with "a newer pull moved HEAD" whenever
`checkout_sha != expected_sha`, and otherwise held the fleet to `expected_sha` exactly.
A cherry-picked hotfix on top of the pulled SHA moves HEAD without any pull arming a
fresh obligation, so the recorded one could never discharge: every CLI start warned
and every no-op `hermes update` exited 1 through the armed veto in
`_pending_fleet_restart_needed`, while the gateway verifiably served HEAD.
The bail-out now fires only when HEAD no longer CONTAINS `expected_sha`
(`checkout_contains`, `git merge-base --is-ancestor`, fail-closed on any probe
failure), and the live fleet is held to the checkout SHA — the code it actually runs.
A gateway on genuinely stale code still keeps the obligation armed.
Closes#119367
`hermes update --yes` from a cron job inside the gateway takes the #100179
ancestor branch (`_drain_or_signal_gateway_for_update` -> self-restart request,
fire-and-forget) — the gateway can only restart after the updater exits. The
post-restart fleet matrix then saw that same gateway on the pre-update code_sha,
printed STALE + "Update not complete" and exited 1 on every nightly run (#119597).
The restart phase now records the pids that accepted a self-restart
(`_GatewayRestartOutcome.self_restart_pending_pids`, threaded from the systemd,
manual and launchd paths) and `collect_fleet_versions(self_restart_pending=)`
turns exactly those rows into `restart_pending` — rendered as "restart pending
(deferred until this process exits)" and excluded from the stale_or_down
verdict. Every other stale/down gateway keeps its verdict; the same pid without
the recorded acceptance is still STALE. The receipt keeps the row's old sha, so
the next `hermes update` / CLI startup hint still verifies the restart landed
via the live fleet (`_receipt_reports_stale_runtime` -> `_live_fleet_covers_receipt`).
Closes#119597
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.
Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
Update stale test assumptions for PM install hints and terminal heartbeat,
keep Unix socket fixtures below AF_UNIX path limits, and remove timing and
host-process races from gateway tests. Read systemd units with BOM-tolerant
UTF-8 so the Windows footgun scan remains clean.
A pre-2026.7 `Restart=on-failure` system unit has no
RestartPreventExitStatus=78, so a permanent refusal restarted it ~180
times while the regenerated user units parked (#118282). The gateway
refreshes its USER unit at boot (`gateway run` -> refresh_systemd_unit_if_needed
(system=False)) but nothing ever rewrote an installed SYSTEM unit unless
the operator ran `sudo hermes gateway install|start|restart --system`.
The update-time fleet restart is the only routine contact with that unit:
as root it now regenerates the stale unit before the drain; without root
it names the exact repair command. A unit that already parks on 78 is
left alone. HERMES_HOME is restored after the refresh so the rest of the
update keeps running for the invoking profile.
Closes#118282
Every update route now finishes through update_completion._complete_selected,
which restarted the whole fleet unconditionally -- including the "Already up to
date" route that main sent through the pending-restart catch-up. Net effect:
each cron tick and each profile's `hermes update` drained and re-killed the one
multiplexed gateway.
Port the catch-up path's two live guards into the completion tail:
- host_restart_already_completed(checkout sha): a sibling profile attaches to
the restart this host already stamped (#95294); the restart phase now stamps
it via mark_host_restart_completed.
- every planned runtime AND every live fleet row current at the checkout sha
(#117051, d6b0d37ece). Both are required: the live matrix lists gateways
only, so a planned serve still on pre-update code keeps the restart.
The already-current route also arms the host obligation with the checkout sha;
an SHA-less arm replaced the standing record and wiped the restarted proof.
Host-scoped update-restart obligation (061195fac1 / 953b6f6f08 / 3be255eca6)
lands on the PM model: the obligation record, its readers and the legacy
per-home marker compat come in as-is. The catch-up restart path
(`_apply_pending_fleet_restart_catchup` / `_run_pending_fleet_restart`) is
retired here (the fleet restart rides the completion owner), so main's
per-host restart-once guard on that path is not carried; its unit→live
MainPID collapse IS ported into the live post-update systemd pass
(`_restart_systemd_gateway_units`), with the two collapse tests rewritten
against that function (red on the pre-port tree: `_unit_main_pid` absent).
Tests that exercised only the retired catch-up path are dropped.
Desktop: main's shared log-rotation planner replaces the inline constants
in main.ts; the merge keeps our machine-profile import beside its import.
utf-8 → utf-8-sig on the three new BOM-intolerant reads (footguns lint).
Review findings on the host-scoped update→restart obligation.
- update_cmd_fleet: an unwritable host state dir (read-only HERMES_GATEWAY_LOCK_DIR,
container UID that does not own $HOME) made write_host_obligation return False and the
caller ignored it, so an interrupted update left ZERO obligation — stale code, no
warning, no catch-up restart. The return is now propagated: the legacy per-home marker
(still read by every reader) carries it, and a host that can write neither says so.
- update_cmd_fleet::_obligation_fields: a PRESENT but unparseable/foreign-versioned host
record no longer falls through to the legacy marker; terms nobody can read cannot be
discharged by another record's terms.
- update_cmd_fleet::_restart_identity_sha: zip/pip/Docker installs resolve no checkout
SHA, so the restart-once stamp was "" and could never match — every profile's update
re-killed the one shared multiplexer. Falls back to the record's expected_sha, then the
receipt's post-update identity.
- update_host_obligation: any main_pid probe error is unproven identity (keep its own
restart), never an aborted restart pass.
- run_notifications: the online notice dedupes per home CHAT, so two served profiles
sharing one chat get one message (accounting stays per profile); transport resolution is
isolated per profile, so one broken adapter no longer starves the rest of the fan-out.
- run_adapters / run_profile_reconcile: _profile_configs is pruned with the served set, so
a failed or removed profile no longer owes a notice nothing can deliver and
.restart_pending.json is unlinked.
Tests cover the new format's own hazards: unwritable record dir, foreign-version record
with a legacy marker present, non-git install, shared home chat, broken adapter, pruned
config, plus a parity test for the duplicated host-state-dir resolver.
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:
* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
loop) and the success marker; with no success marker on disk it printed
"✓ Gateway is running — cron jobs will fire automatically", and with a stale one
it pointed at "Check the gateway log" instead of naming the cause. The persisted
`CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
and reported as "Gateway is running STALE code — fires NOTHING", with both
revisions and the restart command. A fresh heartbeat with a recorded error and no
success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
to the existing drain-first `request_restart` path (SIGUSR1) via
`hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
contract the restart phase already uses for unmapped manual gateways. The drain
budget computation is shared (`_gateway_drain_budget`).
The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).
Fixes#117275
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
`hermes update` on macOS printed "✓ Restarted ai.hermes.gateway" for the
profile running the update whenever launchd reported *any* pid for the label —
including the pre-update process the restart was supposed to replace.
f29ee96dd3 closed this for SIBLING labels only
(`_wait_for_launchd_service_pid(label, old_pid=old_pid, ...)`); the invoking
profile kept `wait_for_launchd_gateway_supervision` →
`_launchctl_label_supervising_process`, a predicate that is true for "launchctl
list exits 0 and some positive pid" and never compares against the pid seen
before the restart.
- `_launchctl_supervised_pid()` exposes the pid `_launchctl_label_supervising_process`
already parsed; the boolean is now a thin wrapper over it.
- `wait_for_launchd_gateway_supervision(..., old_pid=...)` accepts a fresh pid
only, mirroring the sibling loop's contract. `old_pid=None` keeps the old
"any supervised pid" meaning for callers with no pre-restart observation.
- `_restart_launchd_gateway_after_update()` snapshots the pid before
`launchd_restart()` and passes it in; the failure line now says launchd is
not supervising a NEW process. The snapshot is verification-only, so the
no-`launchctl list`-gating invariant of #74973 is untouched.
Also from round-2 review of this PR:
- the supervised-serve discharge test parametrized 5 supervisors, 3 of which
the inventory writer can never put on a serve/dashboard row; cut to the two
it can emit and documented the parity-only members.
- `_marker_only_restart_obsolete` now names who still owns a discharged
launchd-supervised serve (update-time `report_unaccounted_runtimes`, exit 1).
A `fleet_restart_pending` marker records the pre-update plan's runtimes as its inventory,
including `serve`/`dashboard` rows. `_marker_only_restart_obsolete` rejected every non-gateway
inventory row outright, so on any host that runs a dashboard or the Desktop `serve` backend the
marker could never discharge: `hermes version` kept printing "a previous `hermes update` ... did
not restart running gateways" after every gateway was already current on the pulled SHA.
Receipts already draw this boundary (`_receipt_owed_gateways`, #115090) and so does the restart
phase for the Desktop backend (#111494); the marker path was the one place left counting a
supervisor-owned serve row as evidence against the gateways it does not cover. A manual-serve row
still needs its durable handoff and an unclassified backend still fails closed.
Also corrects `_owed_stale_serve_rows`'s docstring, which justified the Desktop exclusion by the
"armed forever" consequence this change removes.
The catch-up rewrote the receipt's fleet matrix before checking whether that
actually cleared the debt, so an exit-1 run (marker still owes a gateway,
blocked serve_restart_pending storage) left latest.json mutated —
tests/hermes_cli/test_update_scoped_reconciliation.py pins the receipt as
byte-identical on the failure path. Decide on the in-memory settled copy and
write only when the pending restart is discharged (#117051).
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).
In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.
Fixes#117051
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").
Part of #117051
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
A KeepAlive LaunchAgent backend's recorded ledger spawner (the bootstrap
shell) is long dead, so _collect_ledger_runtimes read the row as
manual-serve: the plan proposed a respawn-argv restart that fights the
job's own KeepAlive, and a stale survivor was reported as "manual" with
no kickstart hint.
Classify via the existing loaded-job matcher
(_loaded_launchd_backend_jobs / _launchd_job_owning_backend) before the
spawner probe, record the job's domain/label in the row detail, keep the
recovery-partition skip reason and the receipt's gateway-coverage matrix
consistent with the new supervisor value, and name the launchctl
kickstart command in the stale-survivor warning on macOS.
Units labelled before the profile-name suffix scheme (ai.hermes.gateway-<hash>
from the historical _profile_suffix) are invisible to the label derivation, so
the macOS update restart pass left them on pre-update code with no warning.
legacy_launchd_labels_for_install() credits such a plist only when attribution
to THIS install is provable — its ProgramArguments run this install's venv
python, or its pinned HERMES_HOME is one of this install's homes — keeping the
#41403 boundary (a sandboxed HERMES_HOME never touches another install's
fleet). Unreadable or unattributable plists are skipped fail-closed.
Fixes#115254
Two further dead ends in the fleet_restart_pending class that the salvaged
commits left open:
* An inventory-less marker whose expected_sha is no longer HEAD. Every pull
that does not come from `hermes update` (deploy scripts, `git pull` by hand,
a patch-preservation wrapper) skips both the arm and the clear, so a marker
left by an old update survives for weeks with an expected_sha the checkout
will never be at again; `checkout_sha != expected_sha` kept it pending
forever (#115638). Such a marker records no owed set, so the fleet being
current on the checkout is the whole of the evidence its warning can be
about: settle it against checkout HEAD instead. Inventoried markers keep the
moved-HEAD gate — their owed set belongs to the SHA they recorded.
* `_receipt_owed_gateways` returned None for any serve/dashboard row that was
not a manual-serve row already handed to its reminder — so a host running a
Desktop-hosted, systemd, or launchd `hermes dashboard`/`serve` carried a veto
in every receipt and the gateway warning could never be discharged by
evidence (#115090). Those backends are their supervisor's to restart and say
nothing about gateway coverage; only an unclassified backend or a failed
manual transfer still makes coverage unverified.
Tests: the empty-inventory shape moves out of the stays-pending list
(#115311), the identity test's "marker" case gets the inventory its comment
describes (an owed identity at a SHA the fleet does not serve), and the
plan-order test's unsupported row becomes a genuinely unclassified serve
row rather than a Desktop-hosted one.
A pre-update plan that records zero runtimes (Desktop-hosted `serve` with no
gateway services, or a fleet of only manually-deferred serve/dashboard
processes) armed `fleet_restart_pending` with an empty inventory. Nothing
navigates the marker's fail-closed discharge path for an empty `owed` set —
`_marker_only_restart_obsolete` returned False before ever probing, so every
later `hermes update` on the host reported "Fleet restart incomplete" and
exited 1 with no gateway to restart.
- `_write_fleet_restart_pending_marker`: never arm a marker for an explicit
empty inventory (`runtimes=[]`); `None` still arms the conservative
"unknown obligation" breadcrumb.
- `_marker_only_restart_obsolete`: build the owed set before the SHA/probe
gates; an inventory owing nothing is discharged and cleared immediately,
without a fleet probe. Inventories that DO owe restarts keep the existing
fail-closed probe-until-proven-current path; legacy/malformed inventories
stay fail-closed.
Regression tests: marker not armed for `runtimes=[]`; pre-fix empty-inventory
marker discharged without touching the fleet probe; an already-up-to-date
`hermes update` carrying the legacy marker exits 0 with no restart run.
Fixes#115311
_marker_only_restart_obsolete() stayed fail-closed forever when the marker
carried no inventory= line (legacy markers, or a pull that raced a missing
pre-update plan), so every later CLI call printed the interrupted-update
warning even after the fleet provably served expected_sha (#115638).
An inventory-less marker records no owed set, so it now settles on the same
live-fleet bar as an inventoried one minus the owed-coverage check: at least
one row, every row a current gateway under a known profile at expected_sha,
checkout not moved past the marker. Malformed or unsupported inventories stay
fail-closed, as do empty probes and stale/down rows.
Tests: discharge + stale-keep for inventory-less markers; legacy-marker arms
of the scoped-reconciliation matrix now pin live-evidence discharge with the
receipt left intact; explicit inventory=null settles the same way.
`_update_owes_fleet_restart` held a receipt whose restart phase completed to exactly
`post_update.sha`. A gateway restarted afterwards onto a checkout moved by hand runs code
NEWER than the update pulled, so the remedy the warning itself names (`hermes gateway
restart`) could not clear it until the next `hermes update` wrote a fresh receipt
(#113350 steps 3-4). A completed update is discharged when every owed gateway serves the
code it pulled OR today's checkout; the catch-up predicate `hermes update` runs is
unchanged.
`_live_fleet_covers_receipt` and `_marker_only_restart_obsolete` each looped the fleet
twice to fold `served_profiles` into the covered identities and re-validated a shape
`_fleet_row` already enforces. One `_fleet_covered_gateways(fleet)` answers the
`(kind, profile)` set (or None for an unidentified row) for both.
Keep the pending marker until a fresh fleet observation covers its owned gateway inventory. Transfer identified manual serve obligations to durable reminders so an already-verified gateway restart does not keep asking for another restart. Failed transfers retain the marker.
Pin the non-deferred alpha/beta gap and verified-restart-survival cases, including failed probes, storage recovery, and current-successor controls. Both regressions fail without the production change.
Reported-by: ehz0ah
Reported-by: cadamec
Manual serve/dashboard runtimes with no restart mechanism kept re-arming fleet_restart_pending: every CLI start warned about a completed update. This keeps the false alarm out while never hiding a real obligation: reminders are keyed by (pid, create_time), survive receipt rotation and gateway recovery, and clear only when the incarnation is provably gone or explicitly handed off. The pending marker owns its runtime inventory from the moment it is written; settlement derives coverage only from that inventory, so an older receipt can never discharge a newer update's obligation. Markers without an inventory stay pending. Startup keeps the completed-restart evidence path: a receipt whose restart phase finished vouches for the fleet at its post-update SHA.
Reported-by: adamkrawczyk
Diagnosis credit: KoNit-K (#107237)
Mechanism findings: g3org3yo, kokhlo, fmercurio
Reported-by: andrexibiza (P1 generation ownership)
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
`hermes` printed "A previous `hermes update` pulled new code but did not restart
running gateways" on every launch after an update whose receipt was `failed` for a
post-restart crash: `fleet` was empty (the matrix never ran), so the check fell back to
the pre-pull `plan.runtimes[].code_sha` against today's checkout and the live gateway,
correctly restarted onto the pulled commit, still read as "not restarted" once a manual
`git pull` had moved the checkout again. Reproduced on the maintainer's box against the
#112604 receipt: gateway on the update's `post_update.sha`, checkout two pulls ahead.
The startup hint now holds a receipt whose restart phase completed (`gateway_restart`
present, `incomplete` false, no `phase_error`) to the code it actually pulled: every owed
gateway serving `post_update.sha` discharges the update. Later drift is
`gateway/code_skew.py`'s job, not a warning that misreports the update. The catch-up
predicate `hermes update` itself runs keeps its checkout semantics (the fleet must reach
HEAD), so the same box still gets its gateway restarted on the next update.
`_live_fleet_covers_receipt` compares stamped `code_sha` against the expected sha for
`current` AND `stale` rows (both labels are relative to the checkout, not to what the
update owed); `unknown`/`down` rows still never cover.
`hermes update` started in an interpreter that had imported the PRE-pull tree, then
kept running every post-swap phase (dependency sync, Node/web/Desktop builds,
maintenance, config migration, fleet restart, verification, receipt) in that same
process, lazily importing NEW source into an OLD `sys.modules` graph. Any rename
between the two commits surfaced as an ImportError/AttributeError inside the updater
after the code swap had already succeeded (#87134, #111271, #112465, #112558, #112604).
Each incident added another purge, reload list or per-step isolation, and each moved
the crash to the next module nobody had listed.
The pre-pull process now stops at the swap: it writes the open receipt, the pre-update
fleet plan, the pre-update version/active features and the Windows pause token to a
hand-off file and re-executes `hermes update <same flags> --post-swap <file>` under the
venv interpreter. The child imports exclusively from the pulled tree, resumes the
receipt and owns the rest of the run; the parent relays its exit code. Git and ZIP paths
both hand off. On Windows, when the updater runs from `hermes.exe`, the child is spawned
detached exactly as the shim hand-off already did (the shim cannot be awaited while the
sync must replace it).
With no pulled code ever executing in a pre-pull interpreter, the stale-module layer is
dead and removed: `_purge_stale_hermes_modules`, `_stale_purge_prefixes`,
`_evict_module`, `_STALE_PURGE_*`, `_reload_updated_runtime_modules`,
`_reload_process_scan_modules`, `_reload_config_modules`, `_UPDATE_RUNTIME_RELOAD_MODULES`
and their tests. `_run_config_check_fresh` / `_run_migrate_config_fresh` keep their names
and simply call the config API.
Tests: the hand-off boundary (child argv/env, detached receipt + plan in the payload,
exit-code relay) and the child side (receipt resumed with its history, plan rebuilt,
pre-update snapshots taken from the payload). Mocked updater flows run the tail
in-process through the same payload round-trip (autouse fixture; opt out with
`@pytest.mark.real_post_swap_handoff`).
Two ways the 120 s fleet settle window in `_collect_fleet_snapshot` closed before
the relaunched gateway could publish its identity, both turning a clean update
into a wrong verdict:
* `_restarted_units_gone()` asked `systemctl is-active` in the scope named by the
bookkeeping string. A user unit asked in the system scope answers `inactive`
(rc 4) exactly like a dead unit, so one scope-mismatched entry aborted the poll
on its first iteration -> "Fleet version check returned no rows", exit 1,
`fleet_restart_pending` left behind (#112466). Death is now proven positively:
`show -p LoadState,ActiveState` in that scope, and only `LoadState=loaded` with
a non-active state counts; `not-found` (wrong scope, bare/legacy name) is
inconclusive and keeps waiting like the existing systemctl-missing/slow cases.
* An `unknown` row counted as "row-capable -> done", so the one-shot probe
returned ~2 s after a relaunched gateway's process appeared but ~10 s before its
first runtime-status write, and the matrix printed "gateway predates version
stamping; restart to enable" for a gateway this very run had just restarted on
the new code (#112634). An unknown row without a sha whose pid is NOT in the
pre-restart snapshot is a successor still booting: keep polling until it turns
`current`; at the deadline the row is flagged `identity_pending` (carried into
the receipt) and the matrix prints restart-aware copy pointing at
`hermes gateway status`. Surviving pre-restart pids settle as before, so a
genuinely pre-stamping gateway does not hold the poll open.
Existing dead-unit test updated to the `show` output shape; three invariant tests
added (scope mismatch keeps waiting, relaunched pid is waited for / survivor is
not / deadline flag, restart-aware copy).
Fixes#112466Fixes#112634
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.
- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
be untrue — the process provably runs old code and the Desktop app does not respawn it after
a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
`unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
origin/main.
Fixes#111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.
Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).
Fixes#111272
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".
Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.
Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.