Commit Graph

31 Commits

Author SHA1 Message Date
calvinnwq
074ead267d fix(update): classify a Desktop SSH serve as its remote client's, not manual-serve
The serve a remote Desktop spawns over SSH has no local spawner, so the
inventory read it as manual-serve: an update filed a manual-restart
reminder nobody on this host can discharge, reported the stale process
as unaccounted, and abort recovery could try an argv respawn without the
client's token file and owner nonce.

Classify it as desktop-ssh (using the canonical argv predicate, so rows
written before the ledger carried isolated are covered too) and treat it
like the local Desktop's own serve: skipped by the restart phase,
deferred to its client, never owed by abort recovery. A hand-started
serve --isolated stays manual-serve.
2026-09-25 12:09:37 -05:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
liuhao1024
ca7200fc40 fix(update): classify launchd-owned serve/dashboard rows as launchd, not manual-serve
A KeepAlive LaunchAgent backend's recorded ledger spawner (the bootstrap
shell) is long dead, so _collect_ledger_runtimes read the row as
manual-serve: the plan proposed a respawn-argv restart that fights the
job's own KeepAlive, and a stale survivor was reported as "manual" with
no kickstart hint.

Classify via the existing loaded-job matcher
(_loaded_launchd_backend_jobs / _launchd_job_owning_backend) before the
spawner probe, record the job's domain/label in the row detail, keep the
recovery-partition skip reason and the receipt's gateway-coverage matrix
consistent with the new supervisor value, and name the launchctl
kickstart command in the stale-survivor warning on macOS.
2026-09-20 00:22:24 -07:00
ethernet
a6ae6ace51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/js-tests.yml
#	agent/model_metadata.py
#	apps/desktop/electron/main.ts
#	apps/desktop/scripts/bundle-electron-main.mjs
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/settings/gateway-settings.test.tsx
#	apps/desktop/src/app/settings/gateway-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	gateway/shutdown_flush.py
#	hermes_bootstrap.py
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/main.py
#	hermes_cli/managed_uv.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_deps.py
#	hermes_cli/update_cmd_fleet.py
#	hermes_cli/update_cmd_maint.py
#	hermes_cli/update_receipt.py
#	hermes_cli/update_serve_obligations.py
#	hermes_constants.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_managed_uv.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
#	tests/hermes_cli/test_startup_fast_guards.py
#	tests/hermes_cli/test_update_desktop_stale_warning.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/hermes_state/test_hermes_state.py
#	tests/tools/test_tirith_security.py
#	tools/bot_relay.py
#	tools/checkpoint_manager.py
#	tools/write_approval.py
#	website/docs/getting-started/updating.md
#	website/docs/reference/environment-variables.md
2026-09-18 17:26:10 -04:00
yoniebans
c0aa3ce354 fix(update): keep manual serve restart obligations visible through gateway recovery
Manual serve/dashboard runtimes with no restart mechanism kept re-arming fleet_restart_pending: every CLI start warned about a completed update. This keeps the false alarm out while never hiding a real obligation: reminders are keyed by (pid, create_time), survive receipt rotation and gateway recovery, and clear only when the incarnation is provably gone or explicitly handed off. The pending marker owns its runtime inventory from the moment it is written; settlement derives coverage only from that inventory, so an older receipt can never discharge a newer update's obligation. Markers without an inventory stay pending. Startup keeps the completed-restart evidence path: a receipt whose restart phase finished vouches for the fleet at its post-update SHA.

Reported-by: adamkrawczyk
Diagnosis credit: KoNit-K (#107237)
Mechanism findings: g3org3yo, kokhlo, fmercurio
Reported-by: andrexibiza (P1 generation ownership)
2026-09-18 17:09:28 +02:00
ethernet
b4a294fff9 Merge origin/main; keep PM as plugin dependency owner
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
2026-09-17 13:52:05 -04:00
teknium1
94ced1a2b2 fix(update): finish hermes update in an interpreter born on the pulled code
`hermes update` started in an interpreter that had imported the PRE-pull tree, then
kept running every post-swap phase (dependency sync, Node/web/Desktop builds,
maintenance, config migration, fleet restart, verification, receipt) in that same
process, lazily importing NEW source into an OLD `sys.modules` graph. Any rename
between the two commits surfaced as an ImportError/AttributeError inside the updater
after the code swap had already succeeded (#87134, #111271, #112465, #112558, #112604).
Each incident added another purge, reload list or per-step isolation, and each moved
the crash to the next module nobody had listed.

The pre-pull process now stops at the swap: it writes the open receipt, the pre-update
fleet plan, the pre-update version/active features and the Windows pause token to a
hand-off file and re-executes `hermes update <same flags> --post-swap <file>` under the
venv interpreter. The child imports exclusively from the pulled tree, resumes the
receipt and owns the rest of the run; the parent relays its exit code. Git and ZIP paths
both hand off. On Windows, when the updater runs from `hermes.exe`, the child is spawned
detached exactly as the shim hand-off already did (the shim cannot be awaited while the
sync must replace it).

With no pulled code ever executing in a pre-pull interpreter, the stale-module layer is
dead and removed: `_purge_stale_hermes_modules`, `_stale_purge_prefixes`,
`_evict_module`, `_STALE_PURGE_*`, `_reload_updated_runtime_modules`,
`_reload_process_scan_modules`, `_reload_config_modules`, `_UPDATE_RUNTIME_RELOAD_MODULES`
and their tests. `_run_config_check_fresh` / `_run_migrate_config_fresh` keep their names
and simply call the config API.

Tests: the hand-off boundary (child argv/env, detached receipt + plan in the payload,
exit-code relay) and the child side (receipt resumed with its history, plan rebuilt,
pre-update snapshots taken from the payload). Mocked updater flows run the tail
in-process through the same payload round-trip (autouse fixture; opt out with
`@pytest.mark.real_post_swap_handoff`).
2026-09-17 00:02:09 -07:00
teknium1
f554d1d4d7 docs(updater): say what a deferred desktop serve actually proves
The match_runtime_outcomes docstring claimed a Desktop-supervised serve is
`deferred` only when the survivor probe confirmed its pre-update incarnation is
still alive. The probe itself fails closed: an unreadable process ledger lists
every planned serve as surviving, so the row is still deferred without any
liveness observation. Reword to 'when the survivor probe ran and still lists its
pid' and spell out the fail-closed caveat so the verdict does not overstate what
was verified (review follow-up on #113089).
2026-09-16 16:55:42 -07:00
KoNit-K
81bee00ba5 fix(updater): require evidence before deferring desktop serve 2026-09-16 16:55:42 -07:00
teknium1
7c6b8c7a8d fix(update): Desktop-owned serve no longer fails the update or holds fleet_restart_pending
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.

- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
  reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
  be untrue — the process provably runs old code and the Desktop app does not respawn it after
  a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
  `unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
  that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
  hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
  excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
  gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
  in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
  serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
  origin/main.

Fixes #111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:30:32 -07:00
ethernet
612d542281 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.gitignore
#	Dockerfile
#	agent/onboarding.py
#	apps/desktop/electron/main.ts
#	apps/desktop/electron/pool-stop.ts
#	apps/desktop/src/components/model-picker.test.tsx
#	apps/desktop/src/store/updates.ts
#	apps/desktop/vite.config.ts
#	datagen-config-examples/run_browser_tasks.sh
#	docs/rca-ssl-cacert-post-git-pull.md
#	gateway/run.py
#	hermes_cli/backup.py
#	hermes_cli/credential_lifecycle.py
#	hermes_cli/dashboard_procs.py
#	hermes_cli/doctor_state.py
#	hermes_cli/env_loader.py
#	hermes_cli/gateway_windows.py
#	hermes_cli/local_runtime/endpoint.py
#	hermes_cli/psutil_android.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_windows.py
#	hermes_cli/web_routers/local_models.py
#	hermes_cli/web_server_config.py
#	hermes_cli/web_server_cron.py
#	plugins/memory/hindsight/__init__.py
#	plugins/memory/holographic/__init__.py
#	plugins/memory/honcho/cli.py
#	plugins/memory/mem0/__init__.py
#	plugins/platforms/google_chat/oauth.py
#	plugins/platforms/photon/adapter.py
#	scripts/ci/list_os_marked_tests.py
#	scripts/run_tests.sh
#	tests/agent/test_compression_stall_fallback.py
#	tests/agent/test_create_openai_client_ssl_verify.py
#	tests/gateway/test_google_chat_oauth_dependencies.py
#	tests/hermes_cli/conftest.py
#	tests/hermes_cli/test_cli_init.py
#	tests/hermes_cli/test_gateway_migrate_multiplex.py
#	tests/hermes_cli/test_psutil_android_extract.py
#	tests/hermes_cli/test_relaunch.py
#	tests/hermes_cli/test_update_check.py
#	tests/hermes_cli/test_update_handoff_desktop_rebuild.py
#	tests/hermes_cli/test_worktree_gc.py
#	tests/scripts/desktop_update/test_desktop_update_windows_python_handoff.py
#	tests/scripts/desktop_update/test_desktop_update_windows_retry_policy.py
#	tests/scripts/desktop_update/test_desktop_update_windows_timestamp.py
#	tests/scripts/install/test_install_autostash_conflict_recovery.py
#	tests/scripts/install/test_install_clone_throttle_fallback.py
#	tests/scripts/install/test_install_commit_pin_rollback.py
#	tests/scripts/install/test_install_diverged_update.py
#	tests/scripts/install/test_install_lockfile_churn.py
#	tests/scripts/install/test_install_macos_launcher.py
#	tests/scripts/install/test_install_no_initial_commit.py
#	tests/scripts/install/test_install_ps1_ascii_only.py
#	tests/scripts/install/test_install_ps1_browser_install.py
#	tests/scripts/install/test_install_ps1_managed_node_swap.py
#	tests/scripts/install/test_install_ps1_native_stderr_eap.py
#	tests/scripts/install/test_install_ps1_node_path_for_npm.py
#	tests/scripts/install/test_install_ps1_python_fallback_venv.py
#	tests/scripts/install/test_install_ps1_resolver_strictmode.py
#	tests/scripts/install/test_install_ps1_uv_install_fallback.py
#	tests/scripts/install/test_install_ps1_uv_powershell_host.py
#	tests/scripts/install/test_install_ps1_venv_process_tree.py
#	tests/scripts/install/test_install_ps1_venv_recreate_safety.py
#	tests/scripts/install/test_install_ps1_venv_rename_abort.py
#	tests/scripts/install/test_install_ps1_venv_transaction_boundary.py
#	tests/scripts/install/test_install_ps1_web_server_syntax_probe.py
#	tests/scripts/install/test_install_scripts_computer_use.py
#	tests/scripts/install/test_install_sh_acp_launcher.py
#	tests/scripts/install/test_install_sh_bootstrap_marker.py
#	tests/scripts/install/test_install_sh_browser_install.py
#	tests/scripts/install/test_install_sh_install_method_stamp.py
#	tests/scripts/install/test_install_sh_node_deps_failure.py
#	tests/scripts/install/test_install_sh_node_deps_workspaces.py
#	tests/scripts/install/test_install_sh_node_global_prefix.py
#	tests/scripts/install/test_install_sh_node_npm_check.py
#	tests/scripts/install/test_install_sh_node_prerelease.py
#	tests/scripts/install/test_install_sh_node_probe.py
#	tests/scripts/install/test_install_sh_node_tarball_without_xz.py
#	tests/scripts/install/test_install_sh_pythonpath_sanitization.py
#	tests/scripts/install/test_install_sh_reuse_supported_python.py
#	tests/scripts/install/test_install_sh_root_fhs_uv_python_path.py
#	tests/scripts/install/test_install_sh_setup_wizard_tty_probe.py
#	tests/scripts/install/test_install_sh_symlink_stomp.py
#	tests/scripts/install/test_install_sh_termux_network_prereqs.py
#	tests/scripts/install/test_install_sh_termux_python_bounds.py
#	tests/scripts/install/test_install_sh_uv_lock_config.py
#	tests/scripts/install/test_install_unmerged_index.py
#	tests/scripts/test_run_tests_parallel.py
#	tests/test_managed_runtime_resolution.py
#	tests/test_project_metadata.py
#	tests/tools/test_browser_use_cli.py
#	tests/tools/test_tts_pythonpath_fallback.py
#	tests/tui_gateway/test_hosted_room_driver_runtime.py
#	tests/tui_gateway/test_tui_gateway_server.py
#	tools/lazy_deps.py
#	tools/voice_mode.py
#	uv.lock
#	website/docs/developer-guide/macos-bundle-updates.md
#	website/docs/developer-guide/pm-audit-status.md
#	website/docs/developer-guide/shared-bundle-builds.md
#	website/docs/developer-guide/source-update-completion.md
#	website/docs/developer-guide/stable-releases.md
2026-09-14 15:38:34 -04:00
teknium1
9188e708b3 fix(gateway): served_profiles bind to a verified gateway identity, not bare PID existence
`live_default_gateway_pid()` trusted `gateway.pid` + `_pid_exists`, so a stale
default record whose PID an unrelated process had recycled kept its old
`served_profiles` authoritative: `hermes -p X gateway start` exited 78 and
`status` said "running via multiplexer" for a gateway long gone (review of
#108352, finding D). The salvaged #110167 fallback inherited the same bare
check for the pid-file branch.

One helper now answers "which live gateway owns this home?" for every reader:
`gateway.status.live_gateway_pid_for_home` = scoped `get_running_pid` (pid file
+ runtime lock, start-time reuse guard, live gateway command line, home match)
then `get_runtime_status_running_pid(..., expected_home=home)` (honours
`gateway_state` stopped/startup_failed). `gateway_multiplex_served`,
`gateway_migrate._live_gateway_pid` and the `hermes update` inventory's
gateway_state.json fallback (#109680: a `stopped` record + recycled PID
fabricated a phantom runtime, so the update exited partial) all route through
it. Tests that impersonated a gateway with this pytest PID now wear a gateway
command line instead of stubbing `_pid_exists`.
2026-09-13 15:41:01 -07:00
ethernet
ce49cdbc59 merge: reconcile upstream main with pm audit closeout
Merge upstream 5e645791ac.

Retain the PM feature-flag owner and add upstream connection options.
Use the deny-only window-open policy while trusted external links keep
the existing IPC path. Keep both session-import and external-link copy.
Preserve captured timeout output when adding terminal yield handoff.
Quickstart tests patch the explicit upstream model-assignment owner.
Migrate incoming legacy OS markers to the branch's platforms gate.

Desktop renderer and Electron typechecks passed. Targeted Electron tests
passed (42 tests), Python conflict checks passed (26 tests, 3 skips),
and the plugin-compat import checker passed. CI owns the broad merge gate.
2026-09-06 13:17:07 -04:00
ethernet
3c08d16ba7 fix(pm): close runtime publication and updater audit gaps
Dependency publication now recovers interrupted config/facts changes before
activation and leases live generations during collection. Receipts retain
update correlation and failed steps across nested command boundaries.
Doctor and desktop surfaces report those failures through shared owners.

Move checkout updates out of the desktop facade. Stage a detached Windows
relaunch waiter before shutdown, with bounded handshake and process-birth
checks. Keep packaged lifecycle tests isolated from the installed app.

Native verification exposed two production races: cron maintenance imported
the interactive CLI and rewrote TERMINAL_CWD, and install-ID reads collided
with first publication. Use the existing owners and locks. Plugin checks
now run at startup and each due-gated housekeeping tick, not after 60 ticks.

Share updater-test mutation boundaries and remove collection-root fixtures.
Separate cold MCP startup from command latency and give the real HTTP drip
test enough time to reach body handling.

Root npm check passed, including packaging. The fixed-tree Windows Python
run reported 44557 passed, one failed, and 1404 skipped, plus one retry-only
HTTP test. Those final failures now pass in a 35-test bounded batch. A real
isolated gateway wrote startup and periodic plugin-check receipts.

Full final-tree CI, bundled Sandbox deployment, and actual App Installer
relaunch remain unverified. docs/pm-audit-status.md records these limits.
2026-09-06 11:45:41 -04:00
mengtanx
b4b6235239 fix(update): credit launchd ai.hermes.gateway in fleet reconciliation (#103679)
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.

Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
2026-09-06 20:47:52 +05:30
ethernet
92686159d1 fix(pm): integrate audited runtime and lifecycle repairs
Prepare dependency generations before selecting them. Keep shipped tool
bytes separate from writable additions, and store facts beside their entries.
Validate proposed plugin sets before config publication. Restore the previous
config if the facts write fails.

Consolidate duplicate updater, backup, setup, and voice helpers. Repair
launcher selection, dependency consumers, download ownership, update feeds,
and native Windows process and file handling.

Verification: 206 changed/prior-failing Python files reported 4630 passed,
one failed, and 330 skipped. Fix the remaining Hindsight fixture boundary.
The final targeted rerun reported 234 passed and two skipped. The store
review regression batch reported 83 passed and one skipped. Desktop
TypeScript checks, 56 selected Electron tests, 24 release tests, and the
removed-import/compatibility guards passed.

This is an integration checkpoint, not full audit acceptance. The complete
Python suite has not run on this fixed tree. Crash-atomic plugin publication,
generation cleanup, receipt correlation, and packaged lifecycle acceptance
remain open in docs/pm-audit-status.md.
2026-09-05 22:36:48 -04:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
64262d10b3 refactor(hermes_cli/update): split collect_runtime_inventory into phase collectors; fold outcome ladders; drop unused RuntimeRecord.to_dict 2026-09-02 23:56:01 -07:00
Teknium
f36355859b refactor(hermes_cli/update): share profile-home + socket-identity probes between receipt and inventory; fold refusal builder; collapse try/except-pass 2026-09-02 21:49:26 -07:00
Teknium
2c7f5d12b5 refactor(hclib): cron/update/install lifecycle — cron status, update_* receipts and recovery, install repair, service manager 2026-09-02 14:45:25 -07:00
Teknium
f9bca5a0d0 fix(update): reconcile serve/dashboard runtimes in their own vocabulary and escalate survivors (#100479)
Widen the two salvaged fixes (#100490, #100493) to the whole class:

- match_runtime_outcomes: serve/dashboard rows never borrow gateway
  bookkeeping at ANY site — not just the bare hermes-gateway unit name
  (#100490) but also relaunched_profiles / externally_supervised_profiles
  and the profile-substring unit match (hermes-gateway-work credited the
  'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
  units (exact names, scope prefix tolerated) or, when the caller passes
  the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
  feed the Phase-2 reconciliation, so a surviving unmanaged serve is
  'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
  remedy instead of 'hermes gateway restart', which cannot reach it.

Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
2026-09-02 00:42:53 -07:00
chelsealong
806878612a fix(update): stop crediting unmanaged serve runtimes with a gateway's restart
match_runtime_outcomes() treats any default-profile runtime as covered
once the bare "hermes-gateway" unit restarts, regardless of the
runtime's own kind. An sshd-spawned `serve --isolated` backend (no
systemd unit, supervisor "manual-serve") shares the default profile
and gets silently marked "restarted" even though its own PID was never
touched — so the #91277 Phase 2 unaccounted-runtime tripwire never
fires for it and `hermes update` reports success while it keeps
running pre-update code (#100479).

Restrict the "hermes-gateway" special case to kind == "gateway" so a
serve/dashboard runtime under the same profile falls through to
"unaccounted" instead of borrowing the gateway's outcome.
2026-09-02 00:42:53 -07:00
joaomarcos
27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
Casey
790e1eb6bd fix(update): pause SCM-supervised Windows gateway services before venv mutation
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.

This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:

- gateway/status.py: expose service-ownership discovery for gateway
  runtimes (find_windows_gateway_services maps validated gateway PIDs
  through process ancestry to running SCM service PIDs, with
  create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
  venv mutation and restart them afterward. Stops wait for a stable
  SCM 'stopped' state AND for the original descendant processes to
  exit (service 'Stopped' is not proof the child released its
  handles). Failure to prove ownership, stop a service, or restart it
  fails closed; rollback restores attempted services, and rollback
  failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
  or a service that will not reach a stable state abort the update
  before any file mutation.

Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.

Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:45:31 -07:00
Teknium
4860978115 feat(update): image/package-managed installs refuse in-place updates through one shared gate (#91277 Phase 3)
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.

A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.

Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
2026-08-26 11:41:04 -07:00
Teknium
27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
Teknium
18b7fc82b6 feat(update): the plan is now the restart worklist — every planned runtime must be accounted for (#91277 Phase 2)
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).

- update_inventory: restart_via becomes a machine-readable mechanism id
  (systemd|launchd|desktop|manual) — THE policy table as data; display
  derived via describe_restart_mechanism. match_runtime_outcomes()
  reconciles every planned runtime against the restart phase's
  bookkeeping (restarted/stopped/failed/unaccounted);
  report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
  land in the receipt (runtime_outcomes); any unaccounted runtime
  escalates exactly like a STALE/DOWN fleet row (exit 1).

Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
2026-08-23 04:45:29 -07:00
Teknium
7a54ab22e6 fix(gateway): control-socket hardening from #92447 post-merge review
- bind under umask 0o177 so the socket is never world-connectable, even
  pre-chmod (review pt 3)
- verb handlers run in an executor: state-file reads stay off the
  adapter event loop (pt 2)
- inventory dedupes one multiplex gateway answering identify for several
  homes — one runtime record per pid, with regression test (pt 1)
- v1 wire contract (one request per connection) documented in the module
  docstring (pt 5); /tmp-unwritable skip in the short-home test

Live-verified: perms 600 at bind, identify 4.3ms via executor path, 20
rapid queries healthy.
2026-08-22 18:05:26 -07:00
Teknium
60b6269142 feat(gateway): gateway-owned control socket — identify/status verbs, fleet consumers prefer it over scans (#92091 step 1)
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:

- identify: pid, profile, hermes_home, code_sha/code_version (#91283
  stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself

Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.

Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
  identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
  socket, and takes the gateway's own supervisor declaration instead of
  inferring it from PID scans

Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).

Part of #91277 (fleet-update reliability). Design: #92091.
2026-08-22 16:45:00 -07:00
Teknium
0aecadc17c feat(update): hermes update --plan — read-only fleet inventory + plan phase in every update
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.

- hermes_cli/update_inventory.py (new): side-effect-free runtime
  inventory — install kind via detect_install_method (git / docker / nix
  / apt, updatable-in-place or not, with the correct external update
  command for image/package-managed installs), all profiles, every live
  gateway with its supervisor (systemd / launchd / manual via the
  fleet-wide _get_service_pids), running code_sha/code_version from the
  #91283 gateway_state.json stamps, and the restart mechanism each
  runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
  docker/nix refusal gates so image-managed installs get a useful
  'not updatable in place + right command' report instead of a bare
  refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
  ('plan' key) and prints a one-line fleet summary, so post-mortems can
  compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
  cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
  dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
  never-raises, JSON round-trip for the receipt, print output shapes,
  receipt integration.
2026-08-21 04:23:13 -07:00