Commit Graph

86 Commits

Author SHA1 Message Date
kshitijk4poor
8afaab3703 fix(gateway): restart wait survives a non-finite drain; tighten its tests
Follow-up to the salvaged restart-wait commits:

- A drain or cron timeout of .inf now means "wait indefinitely" instead of
  an OverflowError from the integer stop envelope, which crashed
  `hermes gateway restart` and made `hermes update` silently fall back to
  its 45s floor. The fleet "draining (up to Ns)" lines and the drain
  progress report format the budget instead of int()-ing it, so an
  unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
  the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
  replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
  envelope and that configured cron reaches the CLI wait, instead of
  re-deriving the formula; the negative wording assertion on the pending
  footer is dropped (change-detector).
2026-09-26 20:07:47 +05:30
calvinnwq
074ead267d fix(update): classify a Desktop SSH serve as its remote client's, not manual-serve
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.

Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
2026-09-25 12:09:37 -05:00
ethernet
a36dd91ef1 refactor(update): split _marker_only_restart_obsolete below CC 30
The fleet-restart-pending discharge parsed the marker inventory and
probed a gateway-less host inline (CC 36). _marker_owed_gateways now
turns the inventory into the owed set (raising ValueError, which the
caller's existing handler already maps to "keep the marker"), and
_discharge_gatewayless_marker owns the #118742 host probe. Every
verdict is unchanged; the caller is at CC 21.
2026-09-24 14:08:22 -04:00
ethernet
9a8cdf96bf Merge origin/main into ethie/pm-clean
- tools/browser_tool_install.py: keep pm-clean's frozen old-updater stub; main's
  UTF-8 decode fix touched only the npx prefetch body it replaces.
- tests/hermes_cli/test_update_scoped_reconciliation.py: keep pm-clean's test
  subset (catch-up rides the PM completion owner) and take main's gateway-less
  host evidence (#120740): the updated seed that holds the host at a running
  gateway, and the two gateway-less matrices for the source change that merged
  cleanly into update_cmd_fleet.py.
2026-09-23 20:04:27 -04:00
Austin Pickett
d94769da64 fix(update): a Desktop-only host settles an inventory-less restart obligation (#120740)
* refactor(update): one predicate for serve rows outside the gateway matrix

The inventory branch of `_marker_only_restart_obsolete` inlined the rule for which
serve/dashboard rows the gateway matrix neither covers nor needs to (supervisor-owned,
or a manual serve handed to its own reminder). The inventory-less branch needs the same
rule for #118742, so it moves to `update_cmd_fleet_gatewayless.runtime_outside_gateway_evidence`
and both branches will read one definition. No behaviour change.

* fix(update): a Desktop-only host settles an inventory-less restart obligation

A host that runs no gateway (the Desktop app alone) can be left with an inventory-less
fleet-restart obligation: an updater that died before recording its inventory, or the
pre-inventory writer. With no owed set, the live gateway matrix is its only evidence, and
on that host the matrix is empty forever, so `_marker_only_restart_obsolete` never settled
and every later `hermes update` exited 1 with "gateways are still off the checkout code"
(#118742).

An empty fleet alone cannot tell that host from one whose gateway the dying update stopped,
so the inventory-less branch now asks the live host, never a historical receipt:
`host_owes_no_gateway_restart` settles only when no profile's `gateway_state.json` claims a
state other than stopped/startup_failed (a gateway that went away without a clean stop keeps
the obligation) and every live runtime sits outside the gateway matrix
(`runtime_outside_gateway_evidence`, shared with the inventory branch). HEAD must still
contain the pulled SHA (`checkout_contains`, same rule as #119367). Probe failures keep it.

Tests: the scoped-reconciliation matrix now holds its host at "the update stopped a gateway"
so it keeps pinning receipt independence; a new host-evidence matrix covers Desktop-only,
clean stop, carried commit, stopped gateway in a named profile, unclassified and
unidentified serves, a gateway row without fleet identity, and a diverged checkout. Two
manual-serve tests that assumed an empty fleet always stays pending now stub the live host
and assert the manual reminder survives the gateway obligation settling.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>

---------

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-23 18:48:33 -05:00
ethernet
c6b358592d Merge origin/main into ethie/pm-clean 2026-09-23 10:43:45 -04:00
teknium1
30056d612b fix(update): a carried local commit no longer strands the fleet-restart obligation
`_marker_only_restart_obsolete` bailed out with "a newer pull moved HEAD" whenever
`checkout_sha != expected_sha`, and otherwise held the fleet to `expected_sha` exactly.
A cherry-picked hotfix on top of the pulled SHA moves HEAD without any pull arming a
fresh obligation, so the recorded one could never discharge: every CLI start warned
and every no-op `hermes update` exited 1 through the armed veto in
`_pending_fleet_restart_needed`, while the gateway verifiably served HEAD.

The bail-out now fires only when HEAD no longer CONTAINS `expected_sha`
(`checkout_contains`, `git merge-base --is-ancestor`, fail-closed on any probe
failure), and the live fleet is held to the checkout SHA — the code it actually runs.
A gateway on genuinely stale code still keeps the obligation armed.

Closes #119367
2026-09-23 07:01:11 -07:00
teknium1
8895b0f119 fix(update): an in-gateway cron update no longer exits STALE for the gateway it runs inside
`hermes update --yes` from a cron job inside the gateway takes the #100179
ancestor branch (`_drain_or_signal_gateway_for_update` -> self-restart request,
fire-and-forget) — the gateway can only restart after the updater exits. The
post-restart fleet matrix then saw that same gateway on the pre-update code_sha,
printed STALE + "Update not complete" and exited 1 on every nightly run (#119597).

The restart phase now records the pids that accepted a self-restart
(`_GatewayRestartOutcome.self_restart_pending_pids`, threaded from the systemd,
manual and launchd paths) and `collect_fleet_versions(self_restart_pending=)`
turns exactly those rows into `restart_pending` — rendered as "restart pending
(deferred until this process exits)" and excluded from the stale_or_down
verdict. Every other stale/down gateway keeps its verdict; the same pid without
the recorded acceptance is still STALE. The receipt keeps the row's old sha, so
the next `hermes update` / CLI startup hint still verifies the restart landed
via the live fleet (`_receipt_reports_stale_runtime` -> `_live_fleet_covers_receipt`).

Closes #119597
2026-09-23 06:51:47 -07:00
ethernet
d2daa9685e Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/update_cmd_fleet.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
2026-09-23 08:47:59 -04:00
teknium1
c500dc6d99 fix: hermes update restarts only the gateways of the home it updates (#93349)
hermes-gateway*/hermes-serve* units, ai.hermes.gateway* LaunchAgents and
`gateway run` processes are account-wide namespaces shared by every Hermes
install on the box. The restart phase enumerated them by name, so a scratch
home's `hermes update` drained and restarted the account's real
hermes-gateway.service and SIGTERMed sibling installs' gateways (live: #93349
comment, 2026-09-23), then warned "Fleet version check returned no rows"
because none of those runtimes belonged to the updating home.

Ownership is now judged from what a runtime actually runs on, never from the
label: the live process environment (`_hermes_home_for_pid`), the unit's
declared Environment=HERMES_HOME, or the plist's pinned HERMES_HOME, compared
against the homes the update's plan inventories (the updating root and its
profiles/<name>). Foreign or unreadable ownership is named in the output and
left alone; it is not a failed restart. Applies to the systemd fleet loop, the
catch-up best-effort restart, the launchd derived-label loop, the manual
gateway sweep, and the pre/post-restart PID snapshots that drive the
fail-closed verdict.
2026-09-23 05:36:30 -07:00
ethernet
67475a62c6 fix: stabilize Linux Python CI coverage
Update stale test assumptions for PM install hints and terminal heartbeat,
keep Unix socket fixtures below AF_UNIX path limits, and remove timing and
host-process races from gateway tests. Read systemd units with BOM-tolerant
UTF-8 so the Windows footgun scan remains clean.
2026-09-21 20:56:29 -04:00
ethernet
4806579d6c Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-21 20:27:58 -04:00
teknium1
aef4d7a70b fix(update): repair a system unit that cannot park on exit 78 during the fleet restart
A pre-2026.7 `Restart=on-failure` system unit has no
RestartPreventExitStatus=78, so a permanent refusal restarted it ~180
times while the regenerated user units parked (#118282). The gateway
refreshes its USER unit at boot (`gateway run` -> refresh_systemd_unit_if_needed
(system=False)) but nothing ever rewrote an installed SYSTEM unit unless
the operator ran `sudo hermes gateway install|start|restart --system`.

The update-time fleet restart is the only routine contact with that unit:
as root it now regenerates the stale unit before the drain; without root
it names the exact repair command. A unit that already parks on 78 is
left alone. HERMES_HOME is restored after the refresh so the rest of the
update keeps running for the invoking profile.

Closes #118282
2026-09-21 16:52:01 -07:00
ethernet
a60bc9cc64 fix(update): completion tail skips the fleet restart it does not owe
Every update route now finishes through update_completion._complete_selected,
which restarted the whole fleet unconditionally -- including the "Already up to
date" route that main sent through the pending-restart catch-up. Net effect:
each cron tick and each profile's `hermes update` drained and re-killed the one
multiplexed gateway.

Port the catch-up path's two live guards into the completion tail:
- host_restart_already_completed(checkout sha): a sibling profile attaches to
  the restart this host already stamped (#95294); the restart phase now stamps
  it via mark_host_restart_completed.
- every planned runtime AND every live fleet row current at the checkout sha
  (#117051, d6b0d37ece). Both are required: the live matrix lists gateways
  only, so a planned serve still on pre-update code keeps the restart.

The already-current route also arms the host obligation with the checkout sha;
an SHA-less arm replaced the standing record and wiped the restarted proof.
2026-09-21 19:03:35 -04:00
ethernet
793605bc78 merge origin/main (7 commits) into ethie/pm-clean
Host-scoped update-restart obligation (061195fac1 / 953b6f6f08 / 3be255eca6)
lands on the PM model: the obligation record, its readers and the legacy
per-home marker compat come in as-is. The catch-up restart path
(`_apply_pending_fleet_restart_catchup` / `_run_pending_fleet_restart`) is
retired here (the fleet restart rides the completion owner), so main's
per-host restart-once guard on that path is not carried; its unit→live
MainPID collapse IS ported into the live post-update systemd pass
(`_restart_systemd_gateway_units`), with the two collapse tests rewritten
against that function (red on the pre-port tree: `_unit_main_pid` absent).
Tests that exercised only the retired catch-up path are dropped.

Desktop: main's shared log-rotation planner replaces the inline constants
in main.ts; the merge keeps our machine-profile import beside its import.

utf-8 → utf-8-sig on the three new BOM-intolerant reads (footguns lint).
2026-09-21 10:51:15 -04:00
teknium1
953b6f6f08 fix(update): never disarm the host restart obligation, and never discharge unknown terms
Review findings on the host-scoped update→restart obligation.

- update_cmd_fleet: an unwritable host state dir (read-only HERMES_GATEWAY_LOCK_DIR,
  container UID that does not own $HOME) made write_host_obligation return False and the
  caller ignored it, so an interrupted update left ZERO obligation — stale code, no
  warning, no catch-up restart. The return is now propagated: the legacy per-home marker
  (still read by every reader) carries it, and a host that can write neither says so.
- update_cmd_fleet::_obligation_fields: a PRESENT but unparseable/foreign-versioned host
  record no longer falls through to the legacy marker; terms nobody can read cannot be
  discharged by another record's terms.
- update_cmd_fleet::_restart_identity_sha: zip/pip/Docker installs resolve no checkout
  SHA, so the restart-once stamp was "" and could never match — every profile's update
  re-killed the one shared multiplexer. Falls back to the record's expected_sha, then the
  receipt's post-update identity.
- update_host_obligation: any main_pid probe error is unproven identity (keep its own
  restart), never an aborted restart pass.
- run_notifications: the online notice dedupes per home CHAT, so two served profiles
  sharing one chat get one message (accounting stays per profile); transport resolution is
  isolated per profile, so one broken adapter no longer starves the rest of the fan-out.
- run_adapters / run_profile_reconcile: _profile_configs is pruned with the served set, so
  a failed or removed profile no longer owes a notice nothing can deliver and
  .restart_pending.json is unlinked.

Tests cover the new format's own hazards: unwritable record dir, foreign-version record
with a legacy marker present, non-git install, shared home chat, broken adapter, pruned
config, plus a parity test for the duplicated host-state-dir resolver.
2026-09-21 07:14:37 -07:00
teknium1
061195fac1 FLEET: one host-scoped update-restart obligation, one restart per host
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.

- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
  gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
  unit->live-MainPID collapse rule. The legacy per-home marker stays readable
  and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
  idempotent per host (a completed restart onto the checkout SHA is never
  repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
  restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
  profile's home channels; the marker survives until each was reached.
2026-09-21 07:14:37 -07:00
ethernet
a537981a90 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-21 05:07:33 -04:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
teknium1
e7e1928174 fix(update): the invoking profile's launchd restart is claimed only on a pid that changed
`hermes update` on macOS printed "✓ Restarted ai.hermes.gateway" for the
profile running the update whenever launchd reported *any* pid for the label —
including the pre-update process the restart was supposed to replace.
f29ee96dd3 closed this for SIBLING labels only
(`_wait_for_launchd_service_pid(label, old_pid=old_pid, ...)`); the invoking
profile kept `wait_for_launchd_gateway_supervision` →
`_launchctl_label_supervising_process`, a predicate that is true for "launchctl
list exits 0 and some positive pid" and never compares against the pid seen
before the restart.

- `_launchctl_supervised_pid()` exposes the pid `_launchctl_label_supervising_process`
  already parsed; the boolean is now a thin wrapper over it.
- `wait_for_launchd_gateway_supervision(..., old_pid=...)` accepts a fresh pid
  only, mirroring the sibling loop's contract. `old_pid=None` keeps the old
  "any supervised pid" meaning for callers with no pre-restart observation.
- `_restart_launchd_gateway_after_update()` snapshots the pid before
  `launchd_restart()` and passes it in; the failure line now says launchd is
  not supervising a NEW process. The snapshot is verification-only, so the
  no-`launchctl list`-gating invariant of #74973 is untouched.

Also from round-2 review of this PR:
- the supervised-serve discharge test parametrized 5 supervisors, 3 of which
  the inventory writer can never put on a serve/dashboard row; cut to the two
  it can emit and documented the parity-only members.
- `_marker_only_restart_obsolete` now names who still owns a discharged
  launchd-supervised serve (update-time `report_unaccounted_runtimes`, exit 1).
2026-09-20 20:13:57 -07:00
teknium1
e9197afd6d fix(update): a supervised serve backend no longer pins the fleet-restart warning on forever
A `fleet_restart_pending` marker records the pre-update plan's runtimes as its inventory,
including `serve`/`dashboard` rows. `_marker_only_restart_obsolete` rejected every non-gateway
inventory row outright, so on any host that runs a dashboard or the Desktop `serve` backend the
marker could never discharge: `hermes version` kept printing "a previous `hermes update` ... did
not restart running gateways" after every gateway was already current on the pulled SHA.

Receipts already draw this boundary (`_receipt_owed_gateways`, #115090) and so does the restart
phase for the Desktop backend (#111494); the marker path was the one place left counting a
supervisor-owned serve row as evidence against the gateways it does not cover. A manual-serve row
still needs its durable handoff and an unclassified backend still fails closed.

Also corrects `_owed_stale_serve_rows`'s docstring, which justified the Desktop exclusion by the
"armed forever" consequence this change removes.
2026-09-20 20:13:57 -07:00
teknium1
442636bf44 fix(update): settle latest.json only when the live fleet discharges the pending restart
The catch-up rewrote the receipt's fleet matrix before checking whether that
actually cleared the debt, so an exit-1 run (marker still owes a gateway,
blocked serve_restart_pending storage) left latest.json mutated —
tests/hermes_cli/test_update_scoped_reconciliation.py pins the receipt as
byte-identical on the failure path. Decide on the in-memory settled copy and
write only when the pending restart is discharged (#117051).
2026-09-20 18:52:45 -07:00
teknium1
29c4ec8fa7 fix(update): settle a stuck failed receipt once the live fleet serves the checkout code
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).

In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.

Fixes #117051
2026-09-20 18:52:45 -07:00
teknium1
d6b0d37ece fix(update): pending fleet restart leaves gateways already on the checkout code alone
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").

Part of #117051
2026-09-20 18:52:45 -07:00
fangliquanflq
561658d26d fix(update): exempt gateways from separate checkouts 2026-09-20 15:23:04 -07:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
liuhao1024
ca7200fc40 fix(update): classify launchd-owned serve/dashboard rows as launchd, not manual-serve
A KeepAlive LaunchAgent backend's recorded ledger spawner (the bootstrap
shell) is long dead, so _collect_ledger_runtimes read the row as
manual-serve: the plan proposed a respawn-argv restart that fights the
job's own KeepAlive, and a stale survivor was reported as "manual" with
no kickstart hint.

Classify via the existing loaded-job matcher
(_loaded_launchd_backend_jobs / _launchd_job_owning_backend) before the
spawner probe, record the job's domain/label in the row detail, keep the
recovery-partition skip reason and the receipt's gateway-coverage matrix
consistent with the new supervisor value, and name the launchctl
kickstart command in the stale-survivor warning on macOS.
2026-09-20 00:22:24 -07:00
liuhao1024
7c295a8e97 fix(update): restart this install's legacy hash-suffixed launchd gateways
Units labelled before the profile-name suffix scheme (ai.hermes.gateway-<hash>
from the historical _profile_suffix) are invisible to the label derivation, so
the macOS update restart pass left them on pre-update code with no warning.
legacy_launchd_labels_for_install() credits such a plist only when attribution
to THIS install is provable — its ProgramArguments run this install's venv
python, or its pinned HERMES_HOME is one of this install's homes — keeping the
#41403 boundary (a sandboxed HERMES_HOME never touches another install's
fleet). Unreadable or unattributable plists are skipped fail-closed.

Fixes #115254
2026-09-20 00:03:06 -07:00
teknium1
d6ef05e275 fix(update): inventory-less markers settle on a fleet current at the checkout; supervised serve rows never veto gateway discharge
Two further dead ends in the fleet_restart_pending class that the salvaged
commits left open:

* An inventory-less marker whose expected_sha is no longer HEAD. Every pull
  that does not come from `hermes update` (deploy scripts, `git pull` by hand,
  a patch-preservation wrapper) skips both the arm and the clear, so a marker
  left by an old update survives for weeks with an expected_sha the checkout
  will never be at again; `checkout_sha != expected_sha` kept it pending
  forever (#115638). Such a marker records no owed set, so the fleet being
  current on the checkout is the whole of the evidence its warning can be
  about: settle it against checkout HEAD instead. Inventoried markers keep the
  moved-HEAD gate — their owed set belongs to the SHA they recorded.

* `_receipt_owed_gateways` returned None for any serve/dashboard row that was
  not a manual-serve row already handed to its reminder — so a host running a
  Desktop-hosted, systemd, or launchd `hermes dashboard`/`serve` carried a veto
  in every receipt and the gateway warning could never be discharged by
  evidence (#115090). Those backends are their supervisor's to restart and say
  nothing about gateway coverage; only an unclassified backend or a failed
  manual transfer still makes coverage unverified.

Tests: the empty-inventory shape moves out of the stays-pending list
(#115311), the identity test's "marker" case gets the inventory its comment
describes (an owed identity at a SHA the fleet does not serve), and the
plan-order test's unsupported row becomes a genuinely unclassified serve
row rather than a Desktop-hosted one.
2026-09-19 23:52:01 -07:00
KnowWayle999
f491c795f9 fix(update): clear the fleet-restart obligation on empty gateway inventories
A pre-update plan that records zero runtimes (Desktop-hosted `serve` with no
gateway services, or a fleet of only manually-deferred serve/dashboard
processes) armed `fleet_restart_pending` with an empty inventory. Nothing
navigates the marker's fail-closed discharge path for an empty `owed` set —
`_marker_only_restart_obsolete` returned False before ever probing, so every
later `hermes update` on the host reported "Fleet restart incomplete" and
exited 1 with no gateway to restart.

- `_write_fleet_restart_pending_marker`: never arm a marker for an explicit
  empty inventory (`runtimes=[]`); `None` still arms the conservative
  "unknown obligation" breadcrumb.
- `_marker_only_restart_obsolete`: build the owed set before the SHA/probe
  gates; an inventory owing nothing is discharged and cleared immediately,
  without a fleet probe. Inventories that DO owe restarts keep the existing
  fail-closed probe-until-proven-current path; legacy/malformed inventories
  stay fail-closed.

Regression tests: marker not armed for `runtimes=[]`; pre-fix empty-inventory
marker discharged without touching the fleet probe; an already-up-to-date
`hermes update` carrying the legacy marker exits 0 with no restart run.

Fixes #115311
2026-09-19 23:52:01 -07:00
Yagna Vudathu
58ca941785 fix(update): discharge inventory-less fleet-restart markers on live-fleet evidence
_marker_only_restart_obsolete() stayed fail-closed forever when the marker
carried no inventory= line (legacy markers, or a pull that raced a missing
pre-update plan), so every later CLI call printed the interrupted-update
warning even after the fleet provably served expected_sha (#115638).

An inventory-less marker records no owed set, so it now settles on the same
live-fleet bar as an inventoried one minus the owed-coverage check: at least
one row, every row a current gateway under a known profile at expected_sha,
checkout not moved past the marker. Malformed or unsupported inventories stay
fail-closed, as do empty probes and stale/down rows.

Tests: discharge + stale-keep for inventory-less markers; legacy-marker arms
of the scoped-reconciliation matrix now pin live-evidence discharge with the
receipt left intact; explicit inventory=null settles the same way.
2026-09-19 23:52:01 -07:00
ethernet
45c77a40cf fix: read text files with encoding=utf-8-sig in code merged from main
Same rule as 6e294f543d for the reads that arrived with the origin/main merge;
the Windows footguns lane is blocking.
2026-09-19 01:38:09 -04:00
ethernet
82a5affdd3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/backup.py
#	tests/hermes_cli/test_gateway_restart_loop.py
#	website/docs/developer-guide/web-search-provider-plugin.md
#	website/docs/getting-started/installation.md
#	website/docs/getting-started/updating.md
#	website/docs/index.mdx
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/docker.md
#	website/docs/user-guide/windows-wsl-quickstart.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/plugins/index.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/web-search-provider-plugin.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/index.mdx
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/cli-commands.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/environment-variables.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/docker.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/features/plugins.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/security.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/windows-wsl-quickstart.md
2026-09-18 18:41:16 -04:00
ethernet
a6ae6ace51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/js-tests.yml
#	agent/model_metadata.py
#	apps/desktop/electron/main.ts
#	apps/desktop/scripts/bundle-electron-main.mjs
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/settings/gateway-settings.test.tsx
#	apps/desktop/src/app/settings/gateway-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	gateway/shutdown_flush.py
#	hermes_bootstrap.py
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/main.py
#	hermes_cli/managed_uv.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_deps.py
#	hermes_cli/update_cmd_fleet.py
#	hermes_cli/update_cmd_maint.py
#	hermes_cli/update_receipt.py
#	hermes_cli/update_serve_obligations.py
#	hermes_constants.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_managed_uv.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
#	tests/hermes_cli/test_startup_fast_guards.py
#	tests/hermes_cli/test_update_desktop_stale_warning.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/hermes_state/test_hermes_state.py
#	tests/tools/test_tirith_security.py
#	tools/bot_relay.py
#	tools/checkpoint_manager.py
#	tools/write_approval.py
#	website/docs/getting-started/updating.md
#	website/docs/reference/environment-variables.md
2026-09-18 17:26:10 -04:00
teknium1
db302ada44 fix(update): hermes gateway restart onto a moved checkout clears the startup hint
`_update_owes_fleet_restart` held a receipt whose restart phase completed to exactly
`post_update.sha`. A gateway restarted afterwards onto a checkout moved by hand runs code
NEWER than the update pulled, so the remedy the warning itself names (`hermes gateway
restart`) could not clear it until the next `hermes update` wrote a fresh receipt
(#113350 steps 3-4). A completed update is discharged when every owed gateway serves the
code it pulled OR today's checkout; the catch-up predicate `hermes update` runs is
unchanged.
2026-09-18 12:47:23 -07:00
teknium1
bb0c5c4198 refactor(update): one coverage helper for both pending-restart predicates
`_live_fleet_covers_receipt` and `_marker_only_restart_obsolete` each looped the fleet
twice to fold `served_profiles` into the covered identities and re-validated a shape
`_fleet_row` already enforces. One `_fleet_covered_gateways(fleet)` answers the
`(kind, profile)` set (or None for an unidentified row) for both.
2026-09-18 12:47:23 -07:00
KoNit-K
e5fe22646e fix(update): clear fleet restart warning for multiplexers 2026-09-18 12:47:23 -07:00
yoniebans
b263ea5a3a fix(update): P1 verify owned fleet after catch-up restart
Keep the pending marker until a fresh fleet observation covers its owned gateway inventory. Transfer identified manual serve obligations to durable reminders so an already-verified gateway restart does not keep asking for another restart. Failed transfers retain the marker.

Pin the non-deferred alpha/beta gap and verified-restart-survival cases, including failed probes, storage recovery, and current-successor controls. Both regressions fail without the production change.

Reported-by: ehz0ah

Reported-by: cadamec
2026-09-18 17:09:28 +02:00
yoniebans
c0aa3ce354 fix(update): keep manual serve restart obligations visible through gateway recovery
Manual serve/dashboard runtimes with no restart mechanism kept re-arming fleet_restart_pending: every CLI start warned about a completed update. This keeps the false alarm out while never hiding a real obligation: reminders are keyed by (pid, create_time), survive receipt rotation and gateway recovery, and clear only when the incarnation is provably gone or explicitly handed off. The pending marker owns its runtime inventory from the moment it is written; settlement derives coverage only from that inventory, so an older receipt can never discharge a newer update's obligation. Markers without an inventory stay pending. Startup keeps the completed-restart evidence path: a receipt whose restart phase finished vouches for the fleet at its post-update SHA.

Reported-by: adamkrawczyk
Diagnosis credit: KoNit-K (#107237)
Mechanism findings: g3org3yo, kokhlo, fmercurio
Reported-by: andrexibiza (P1 generation ownership)
2026-09-18 17:09:28 +02:00
ethernet
b4a294fff9 Merge origin/main; keep PM as plugin dependency owner
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
2026-09-17 13:52:05 -04:00
teknium1
ae71d85bba fix(update): startup hint no longer blames an update that did restart the fleet
`hermes` printed "A previous `hermes update` pulled new code but did not restart
running gateways" on every launch after an update whose receipt was `failed` for a
post-restart crash: `fleet` was empty (the matrix never ran), so the check fell back to
the pre-pull `plan.runtimes[].code_sha` against today's checkout and the live gateway,
correctly restarted onto the pulled commit, still read as "not restarted" once a manual
`git pull` had moved the checkout again. Reproduced on the maintainer's box against the
#112604 receipt: gateway on the update's `post_update.sha`, checkout two pulls ahead.

The startup hint now holds a receipt whose restart phase completed (`gateway_restart`
present, `incomplete` false, no `phase_error`) to the code it actually pulled: every owed
gateway serving `post_update.sha` discharges the update. Later drift is
`gateway/code_skew.py`'s job, not a warning that misreports the update. The catch-up
predicate `hermes update` itself runs keeps its checkout semantics (the fleet must reach
HEAD), so the same box still gets its gateway restarted on the next update.

`_live_fleet_covers_receipt` compares stamped `code_sha` against the expected sha for
`current` AND `stale` rows (both labels are relative to the checkout, not to what the
update owed); `unknown`/`down` rows still never cover.
2026-09-17 00:02:09 -07:00
teknium1
94ced1a2b2 fix(update): finish hermes update in an interpreter born on the pulled code
`hermes update` started in an interpreter that had imported the PRE-pull tree, then
kept running every post-swap phase (dependency sync, Node/web/Desktop builds,
maintenance, config migration, fleet restart, verification, receipt) in that same
process, lazily importing NEW source into an OLD `sys.modules` graph. Any rename
between the two commits surfaced as an ImportError/AttributeError inside the updater
after the code swap had already succeeded (#87134, #111271, #112465, #112558, #112604).
Each incident added another purge, reload list or per-step isolation, and each moved
the crash to the next module nobody had listed.

The pre-pull process now stops at the swap: it writes the open receipt, the pre-update
fleet plan, the pre-update version/active features and the Windows pause token to a
hand-off file and re-executes `hermes update <same flags> --post-swap <file>` under the
venv interpreter. The child imports exclusively from the pulled tree, resumes the
receipt and owns the rest of the run; the parent relays its exit code. Git and ZIP paths
both hand off. On Windows, when the updater runs from `hermes.exe`, the child is spawned
detached exactly as the shim hand-off already did (the shim cannot be awaited while the
sync must replace it).

With no pulled code ever executing in a pre-pull interpreter, the stale-module layer is
dead and removed: `_purge_stale_hermes_modules`, `_stale_purge_prefixes`,
`_evict_module`, `_STALE_PURGE_*`, `_reload_updated_runtime_modules`,
`_reload_process_scan_modules`, `_reload_config_modules`, `_UPDATE_RUNTIME_RELOAD_MODULES`
and their tests. `_run_config_check_fresh` / `_run_migrate_config_fresh` keep their names
and simply call the config API.

Tests: the hand-off boundary (child argv/env, detached receipt + plan in the payload,
exit-code relay) and the child side (receipt resumed with its history, plan rebuilt,
pre-update snapshots taken from the payload). Mocked updater flows run the tail
in-process through the same payload round-trip (autouse fixture; opt out with
`@pytest.mark.real_post_swap_handoff`).
2026-09-17 00:02:09 -07:00
teknium1
f079458401 fix: post-update fleet settle poll no longer ends early on inconclusive signals
Two ways the 120 s fleet settle window in `_collect_fleet_snapshot` closed before
the relaunched gateway could publish its identity, both turning a clean update
into a wrong verdict:

* `_restarted_units_gone()` asked `systemctl is-active` in the scope named by the
  bookkeeping string. A user unit asked in the system scope answers `inactive`
  (rc 4) exactly like a dead unit, so one scope-mismatched entry aborted the poll
  on its first iteration -> "Fleet version check returned no rows", exit 1,
  `fleet_restart_pending` left behind (#112466). Death is now proven positively:
  `show -p LoadState,ActiveState` in that scope, and only `LoadState=loaded` with
  a non-active state counts; `not-found` (wrong scope, bare/legacy name) is
  inconclusive and keeps waiting like the existing systemctl-missing/slow cases.

* An `unknown` row counted as "row-capable -> done", so the one-shot probe
  returned ~2 s after a relaunched gateway's process appeared but ~10 s before its
  first runtime-status write, and the matrix printed "gateway predates version
  stamping; restart to enable" for a gateway this very run had just restarted on
  the new code (#112634). An unknown row without a sha whose pid is NOT in the
  pre-restart snapshot is a successor still booting: keep polling until it turns
  `current`; at the deadline the row is flagged `identity_pending` (carried into
  the receipt) and the matrix prints restart-aware copy pointing at
  `hermes gateway status`. Surviving pre-restart pids settle as before, so a
  genuinely pre-stamping gateway does not hold the poll open.

Existing dead-unit test updated to the `show` output shape; three invariant tests
added (scope mismatch keeps waiting, relaunched pid is waited for / survivor is
not / deadline flag, restart-aware copy).

Fixes #112466
Fixes #112634
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:04:46 -07:00
teknium1
2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural
7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
teknium1
7c6b8c7a8d fix(update): Desktop-owned serve no longer fails the update or holds fleet_restart_pending
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.

- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
  reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
  be untrue — the process provably runs old code and the Desktop app does not respawn it after
  a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
  `unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
  that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
  hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
  excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
  gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
  in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
  serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
  origin/main.

Fixes #111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:30:32 -07:00
teknium1
f81c33cb6a fix(update): stop the fleet settle poll early when the restarted unit is dead
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
2026-09-15 15:10:41 -07:00
joaomarcos
c99fee4e0a fix(update): wait for fleet state publication after supervised restart
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.

Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).

Fixes #111272
2026-09-15 15:10:41 -07:00
teknium1
b2577df807 fix(dashboard): accept command-scoped NOPASSWD sudo for system gateway actions
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".

Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.

Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
2026-09-15 04:08:00 -07:00