One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].
Fixes#117246
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.
Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.
Fixes#115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
Follow-up to the cherry-picked #116014:
- Pin per the dependency policy (pre-1.0: `<0.(minor+2)`): httptools
`>=0.6.3,<0.9` (floor = uvicorn[standard]'s own floor), uvloop
`>=0.15.1,<0.24`. `watchfiles>=0.20,<2` already complied.
- Copy uvicorn's own uvloop marker (win32, cygwin, PyPy) plus
`sys_platform != 'android'` so `pip install '.[all]'` on those
hosts does not fail on the extra either.
- `tools/lazy_deps.py` mirrors the `web` extra for the lazy dashboard
install: it also requested `uvicorn[standard]`, so a Termux user
opening the dashboard would have hit the same uvloop build at first
use. The web_server install hint follows.
- `uv lock` regenerated; the lock delta is exactly the pyproject delta.
- Two invariant tests: no Termux-reachable extra (or core, or the lazy
dashboard feature) requests uvloop; `[all]` still does, off Android.
- Docs: troubleshooting entry in the Termux guide.
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
Closes the second #97208 atom (qingchuan-x): after a Windows hand-off child refused the
dependency sync (Desktop backend held the venv), the next `hermes update` found git current,
passed the core-import probe — the old release imports fine — and printed
"✓ Already up to date!" while the installed distribution still read hermes-agent 0.20.6
against a 0.21.2 checkout with lagging pins.
`_venv_dependency_set_stale` asks the venv's own interpreter for the installed hermes-agent
version and compares it with pyproject's; a mismatch on the current-checkout path takes
the same repair route as an unhealthy venv. Unknown states (no venv, not installed as a
distribution, probe failure) are never stale, so dev checkouts are untouched.
Review follow-up on the #101600 fix:
- The wait moves from `_cmd_update_impl` to `cmd_update`, BEFORE `UpdateLock.acquire()`.
The child used to run under the parent's marker (handoff pid / ancestor claim); the
parent's exit — which the child now deliberately waits for — released that marker, so
the whole Windows dependency install and gateway resume ran with no update lock. Waiting
first lets the child claim a marker of its own for the tail.
- `SHIM_PARENT_PID_ENV` names the `hermes.exe` launcher ancestor (the process that holds
the shim image open and exits only after reaping this interpreter) when psutil sees one,
else this pid. `_windows_shim_in_process_chain` is split into `_venv_shim_matcher` +
`_windows_shim_ancestor` so the holder pid shares the matcher; docs sentence adjusted.
- The direct-call wait test is replaced by one driving `cmd_update` in a real child with an
older parent: proves the phase starts after the parent died, the child owns the lock, the
handed-off pause token is adopted (no re-discovery) and the env is consumed. Removing the
wait call, the token adoption, or moving the wait back after the lock each turns it red.
`hermes update` run from `venv\Scripts\hermes.exe` hands the dependency install to a
child under the venv Python because the shim cannot replace itself. Two races were left
in that hand-off (#101600, three independent Windows reproductions):
- The child never waited for the shim-run parent. The shim quarantine is a single
`os.rename` with sub-second retries; when the child reached it while the parent was
still alive the rename failed with PermissionError, `ShimQuarantineError` deferred the
whole install and the receipt was marked failed although the checkout had advanced.
- On the legacy re-exec (`_abort_dependency_sync_if_self_locked` ->
`_reexec_dependency_sync_off_windows_shim`, still live for a current checkout with an
unhealthy venv -- i.e. the retry after any failed update) the PARENT resumed the paused
gateways after spawning the child, holding hermes.exe open for the whole relaunch and
raising `RuntimeError: Windows gateway relaunch after update was not verified alive`;
the child then re-ran pause discovery, found that freshly relaunched gateway before its
pid file existed and force-killed it as "without profile mapping".
Now every child spawned off the shim gets `HERMES_UPDATE_SHIM_PARENT_PID` and, at the top
of `_cmd_update_impl`, waits (bounded, 30 s, `(pid, create_time)` identity via psutil)
for that process to exit before scanning holders, pausing gateways or renaming shims. The
legacy re-exec carries the Windows pause token to the child in
`HERMES_UPDATE_GATEWAY_RESUME` and disarms the parent's copy, so the parent exits at once
and the child resumes exactly the fleet the parent stopped instead of rediscovering it --
the same ownership transfer the post-swap hand-off already does (94ced1a2b2).
The wait is host-independent and proven with real processes on Linux; the shim paths run
on the windows-latest lane through the existing `_is_windows`-patched suite.
Fixes#101600
Refs #98010#103031
Supersedes #101630
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
Explain that the matrix shows one row per gateway process, that `served_profiles`
carries coverage for the satellites, and that `hermes gateway restart` now clears the
pending-restart hint on such a fleet.
Two follow-ups on the salvaged commits:
- `_ensure_default_soul_md`: widen the cyclic-only (ELOOP) branch to any symlink the seed
cannot write through — a link dangling into a missing directory raises ENOENT on the same
write and bricked home init the same way (HomeInitializationError -> exit 75 -> supervisor
relaunch storm). Shape: write through the link first (a working link stays operator
wiring), and only on OSError seed the default in place of the link via mkstemp + replace.
Tests trimmed to two invariants (cyclic/dangling replaced; resolving link preserved).
- `_run_pre_update_backup`: the quick snapshot stays best-effort by design (8ed599dc05:
"a broken backup never blocks the update"), but a failure was swallowed at DEBUG level, so
the user only learned from the receipt afterwards that no recovery point existed. Print a
stdout warning with the reason and continue; together with the receipt skip/fail split the
receipt now says "failed" only when a snapshot was requested and not produced.
- docs: updating.md describes the warning and the receipt split.
The updating guide describes the rename-over-previous-build step; mention that a
scanner briefly holding release/win-unpacked is ridden out with short retries
(#112544) so the paragraph matches the promotion behaviour.
Rework of the salvaged #59942 hook so it fixes the whole #52339 class without
regressing what the Desktop updater already gets right:
- macOS only. The Windows arm rmtree'd a possibly-running NSIS install
(partial deletion under a lock) and windows.ps1 already owns that swap;
Linux packages stay with their package manager.
- No re-sign. The rebuilt release/ bundle already carries the stable local
signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -`
on the installed copy replaced it with a fresh ad-hoc cdhash and reset every
TCC grant. ditto preserves the signature, so nothing is signed here.
- Running bundles are reported, not swapped: Electron loads app.asar and helper
apps lazily, so renaming the bundle away and deleting the old tree crashes
the live app. The detached updater waits for exit; a terminal `hermes update`
with the app open now prints what to do instead.
- Failures are printed as warnings; the old `Path | None` return read every
failure as "Desktop app up to date".
- The refresh also runs on the "build stamp current" path, so a stale
/Applications copy left by an earlier update heals on the next `hermes update`
even when there is nothing to rebuild.
- Core is host-independent (_install_rebuilt_macos_bundles takes paths as data);
the two invariant tests run on every OS instead of `skipif(darwin)` tests that
ran nowhere.
Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app
running from apps/desktop/release/ now refreshes the /Applications copy Finder
launches (the Discord report: new shell right after the update, old shell on
the next Dock launch).
hermes_cli/AGENTS.md: the update pipeline no longer runs pulled code in the pre-pull
interpreter; the receipt crosses the hand-off; purge/reload/isolate fixes are not to be
reintroduced. Stale comments in update_receipt / update_abort_recovery / update_cmd_maint
that explained the old shape are rewritten to say why the code is the way it is now.
evals/update_pipeline/post_swap_handoff_ab.sh: installs a real clone at a given SHA,
publishes a target commit whose new module-level import (a symbol added to a module the
old updater kept cached) and tagged completion line show which interpreter ran the tail,
runs `hermes update` against a disposable HOME with no npm and --no-gateway-restart, and
prints one VERDICT line. base (6005aa1f): exit 1, "Could not check config version", old
completion line. fix: exit 0, config check runs, "[completion printed by: pulled-code]",
one receipt carrying post_swap_pid.
The updater and Desktop preflight now accept the uv-default .venv layout;
say so where the Windows guards are documented, including which directory
wins when both exist.
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The flat-install block only ignored the bare cron/executions.db. All three
cron stores (executions, deliveries, notepad) are opened through
sqlite_util.open_db in WAL mode, so executions.db-wal/-shm exist whenever
the scheduler is live, and deliveries.db / notepad.db were not ignored at
all. `git stash push --include-untracked` therefore still unlinked the
live WAL/SHM and the whole deliveries/notepad stores under the running
gateway - the same mechanism this PR closes for the root state.db.
Switch the cron entries to the by-class shape already used at the root
(/cron/*.db plus -wal/-shm/-journal/retired-wal sidecars) and also ignore
the cron lock/heartbeat/output files and the per-launch root markers
(.update_check, gateway-starts.log, .clean_shutdown, active_profile,
.hermes_history, slack_tokens.json) that otherwise force every update
into the stash step. `git ls-files -i -c --exclude-standard` is unchanged
(no tracked file newly hidden). The test's FLAT_INSTALL_RUNTIME_STATE
gains one representative per new class; red before, green after.
Review finding: cron/executions.db-wal, cron/deliveries.db and cron/notepad.db were unignored and swept by the flat-install autostash.
`backups/` is where the pre-update snapshot the updater restores a swept
state.db FROM lives — leaving it unignored means the recovery copy is
swept together with the live database. `vault.key`/`vault.json.enc` are
the local secret vault.
Docs: website/docs/getting-started/updating.md explains that on a flat
install (checkout root == $HERMES_HOME) runtime state is git-ignored and
never enters the autostash.
The scope-dispatching cron worker runs inside the container there, so a
host user manager would start for nothing (as PR #110641 by @liuhao1024
also gated it). nix-setup.md gains the note operators need when they
declare the user themselves.
platform-support.md already lists "macOS on x86 (Intel) processors" under
Unsupported, but the two pages users actually land on don't reflect it:
installation.md recommends the macOS installer without qualification, and
desktop.md says the app "runs on macOS, Windows, and Linux".
Cross-reference the existing policy from both pages so Intel users find it
before downloading rather than after "Bad CPU type in executable".
No change in platform support is proposed or implied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
What `hermes update` does, blockers and fixes, URL change for
inbound-port profiles, the post-create restart reminder, rollback, and
the `gateway migrate` reference row.
Each fix verified against main @ ee84ccd on 2026-09-08.
- windows-native: the troubleshooting entry told users to set HERMES_GATEWAY_FORCE_STARTUP (no code reads it), query a task named HermesGateway (hermes_cli/gateway_windows.py names it Hermes_Gateway), and described the Startup-folder fallback as a cmd.exe shortcut (it writes a .vbs run via wscript.exe). Closes#88077.
- faq: 'hermes config set HERMES_MODEL ...' does not change the default model; 'model.default' does. Closes#65855.
- messaging/index, slash-commands, irc: gateway settings are read from ~/.hermes/config.yaml (gateway/config.py), not gateway-config.yaml. Closes#65857, closes#78276.
- telegram: reaction lifecycle is 👀 then 👍/👎 (plugins/platforms/telegram/adapter.py on_processing_complete), docs said ✅/❌. Closes#78698.
- plugins: bundled memory providers win on a name collision (plugins/memory/__init__.py docstring: bundled, user, project, entry point, first seen wins); the page said user plugins override. Model providers keep the documented last-writer-wins. Closes#100281.
- architecture, index: terminal backend count is seven (tools/terminal_tool.py: local, docker, singularity, modal, daytona, vercel_sandbox, ssh); two surfaces said six. Closes#78252.
- quickstart: the Portal quick path was called 'free'; the login is free, the inference is billed to the subscription. Wording now matches integrations/nous-portal. Closes#78254.
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Salvage #101887 after native Actions 34097643131 reproduced WinError 32 using a ready Electron app with its cwd inside the live release. Reuse the existing install-scoped process cleanup before promotion and wait after forced termination, preserving rollback.
Co-authored-by: fangliquan <fangliquan@qq.com>
Adapt the config-only portion of #104347; omit its environment flag and unrelated docs. Explicit update commands remain independent.
Co-authored-by: Rohith Pariki <rohithpariki@gmail.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Native Windows run 34096838164 reports false success for absent and corrupt executables, missing bundle files, missing chunks and missing or stale stamps. Reuse the existing build identity and PE validators, and check interpreter presence before waiting for Desktop. Preserve dependency recovery and exit-2 refusal behavior.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Port the runtime verification portion of #104692 after native run 34095483533 reproduced ok=true for a zero-exit controlled child that removed its runtime module. Artifact/build-stamp validation remains unaddressed.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Salvage the missing-target guard from #104692. Native Windows run 34094671567 returned exit zero for the absent maintained script. Full runtime/artifact completion remains separate.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Keep externally managed directory links and permissions intact during home
initialization. Refuse missing targets rather than creating directories on an
unmounted volume's underlying filesystem. Report link, target, mount and access
guidance through doctor while preserving config.yaml.
Extract the home initialization phase into config_home, and memoize successful
resolved aliases so plugin discovery cannot repeat chmod after losing the
symlink spelling. Live Linux doctor PTY A/B verified directory and root links,
plain paths, missing targets, mount-style missing paths and file conflicts.
Targeted invariant tests are queued under the campaign's shared serial lock;
this checkpoint is not a unit-suite or merge-readiness claim.
Inspired by #104774 and #103735; deliberately does not auto-create external
targets or silently ignore an unavailable sessions directory.
Co-authored-by: ca-shrimp <320556551+ca-shrimp@users.noreply.github.com>
Co-authored-by: Craig Richardson <craigrichardson@Craigs-Mac-mini.local>
`hermes update` → `hermes desktop --build-only` → `npm run pack` packed
electron-builder's output IN PLACE: before-pack.mjs wipes
`release/<platform>-unpacked` (or the mac `Hermes.app`) before the Electron
unpack/asar/rename, so any failure after that point — corrupt cached zip,
blocked download, missing dep, disk full — left the user with NO app and the
update reporting "partially complete" over an empty release/ (#86443).
Fix the class, not the predicate: cmd_gui now passes
`-c.directories.output=apps/desktop/.staging-<pid>-<ts>` to the pack, runs
the existing verification (packaged-exe probe, macOS re-sign, Windows PE
integrity gate) against the STAGED tree, and only then promotes it:
`release/<unpacked>` → `.previous`, `<staging>/<unpacked>` → `release/<unpacked>`,
drop `.previous`. A rename failure between the two steps restores `.previous`.
On any failure the staging dir is removed and the live app is untouched.
- `_purge_electron_build_cache` / `_ensure_desktop_exe_launchable` /
`_desktop_macos_relaunchable_fixup` take the output dir so the corrupt-zip
retry purge and the integrity self-heal only ever clear the staging tree,
never `release/*-unpacked`.
- `.gitignore` the staging dir so a killed build cannot dirty the checkout.
- Docs: updating.md describes the stage-and-swap Desktop rebuild step.
Live repro (real `_rebuild_desktop_after_update` → real `hermes desktop
--build-only` subprocess, fake npm whose pack wipes appOutDir then fails):
before — `release/linux-unpacked/hermes` gone after the failed rebuild;
after — marker intact, no `.staging-*` left, rebuild returns False; a
passing pack swaps the new app into `release/`.
Closes#86443
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
Co-authored-by: deathxdefeat <deathxdefeat@users.noreply.github.com>
The locked dependency tree now carries @babel/* 8.x, which requires
node ^22.18.0 || >=24.11.0. Our engines.node arm said ^24.0.0 and the
installer gates (node_satisfies_build / Test-NodeVersionOk) accepted any
Node 24 — so a system Node 24.0–24.10 cleared every gate we own and then
failed 'npm install' with EBADENGINE under engine-strict=true.
- Raise the 24 arm to ^24.11.0 in root + desktop package.json and the
package-lock.json mirrors
- Tighten node_satisfies_build (install.sh) and Test-NodeVersionOk
(install.ps1) to 24.11+; update user-facing wording
- Add invariant tests: every engines.node arm floor must satisfy every
locked dependency's engines.node, and the installer gates must encode
the same floors as the manifest — so the next babel-style floor bump
turns into a CI red instead of a user install outage
- docs: correct stale 'Node.js v22' provisioning claim
GSC page-level data shows configuration (171K impr, pos 5.3), quickstart
(135K, 5.3), providers (123K, 5.3), web-dashboard (68K, 6.5), docker
(37K, 6.3) and the desktop app page (98K, 3.9) all losing ranking
headroom because their title/H1 are bare nouns instead of the terms
people search.
- Frontmatter title + H1 now carry the query terms on all six pages
- Desktop docs page links back to the new marketing /desktop product
page, joining the two official properties Google sees for the query
Done by Hermes Agent (deepseek-v4-pro via nous), Nous Research.
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:
- identify: pid, profile, hermes_home, code_sha/code_version (#91283
stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself
Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.
Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
socket, and takes the gateway's own supervisor declaration instead of
inferring it from PID scans
Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).
Part of #91277 (fleet-update reliability). Design: #92091.
Adapts the in-place branch update from PR #89507 (@willfrombr) onto the
switch-by-default behavior: the deterministic switch path remains the
default so non-interactive updates (desktop, gateway, cron) never dead-end
on a merge conflict, and deliberate custom-branch users opt in with
updates.parked_branch_strategy: update_in_place. --switch-branch overrides
the in-place strategy for one run (deep feature branches that must not
accumulate update merge commits). Docs + config comments + tests cover
all three routes.
Co-authored-by: Willian Santos <285090322+willfrombr@users.noreply.github.com>
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.
Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
Home Manager separates an installation from a daemon. This module put
both under `services.hermes-agent`, and `installPackage` added a program
to the PATH from a service module.
`programs.hermes-agent` now installs the command line application and
the desktop application. `services.hermes-agent` keeps the state, the
configuration and the daemons, and stays the authority: the new module
reads `hermesHome` and the backend address from it. A person can enable
one without the other, which is a machine with an application and no
gateway, or a headless gateway with no display.
The desktop application needs this split to work correctly. A launcher
that starts from the desktop menu reads no shell profile, thus the
HERMES_HOME that `home.sessionVariables` exports reaches an interactive
shell only. Home Manager writes `systemd.user.sessionVariables` to
environment.d, and this module puts no HERMES_HOME there, because that
file applies to each user unit. The application then opens ~/.hermes
while the services use `hermesHome`, and the person sees no sessions and
no keys. Thus the launcher carries the value itself, through a new
`extraEnv` argument on the desktop package.
The application also gets the Nix agent package, with
HERMES_DESKTOP_HERMES. The usual distribution of the Electron
application carries its own Hermes runtime and downloads more at the
first start. `hermesDesktop` is a passthru of the agent and pins
`finalAttrs.finalPackage`, so an override of `extraPythonPackages` or
`extraDependencyGroups` reaches both. One machine thus has one runtime.
`backend.sessionTokenFile` connects the application to the backend of
the service. Without it the module runs `hermes serve` and the
application starts a backend of its own, which gives two backends on one
HERMES_HOME. The backend reads the file into
HERMES_DASHBOARD_SESSION_TOKEN. The launcher reads the same file into
HERMES_DESKTOP_REMOTE_TOKEN, beside a HERMES_DESKTOP_REMOTE_URL that
names the address of the service.
Measurements against a live `hermes serve` on loopback show why that
shape is the correct one:
- `_resolve_session_token()` reads HERMES_DASHBOARD_SESSION_TOKEN, and
`_has_valid_session_token` accepts that value as a Bearer credential.
A request without it gets 401, and a request with the wrong value
gets 401.
- The /api/ws socket accepts a query parameter only. A header gets 403,
and `?token=` connects. Hermes Desktop builds exactly that URL, in
`apps/desktop/electron/connection-config.ts`. Thus a test of the HTTP
leg alone is a false positive.
- `resolveDesktopRemoteRoute` throws when the URL is set and the token
is not. Thus the two variables travel together or not at all.
The token enters no Nix store path. `makeWrapper --set` and a systemd
`Environment=` value both write a literal into the store, which all
users can read. Thus each side reads the file at start time. The
launcher does it through a new `extraRun` argument on the desktop
package, and the backend through the launcher script that
`backend.waitFor` already uses. launchd has no EnvironmentFile, so a
script is the one shape that works on Linux and on Darwin.
`backendArgv` gives the plain argv only when nothing must run before
the backend.
`services.hermes-agent.installPackage` is removed. It defaulted to true,
so a person who never named it still got the command line. A silent
removal thus gives them a machine with no `hermes` and no message. The
module refuses a configuration that sets it, and the text names the
exact replacement for the value they gave.
Checks:
- the launcher carries HERMES_HOME
- the launcher reports HERMES_MANAGED only when the services own the
configuration, because no activation writes a marker without them
- the launcher pins the agent package that `programs.enable` installs
- the launcher names the backend of the service, and gives a token
beside the URL
- the backend reads the session token
- each side reads the file at start time, and the token is no `--set`
value
- `programs.enable` alone starts no service
- `installPackage` is refused, with a message that names the
replacement, and its absence evaluates
Each check reads the wrapper of the real package, and not an option
value. Each one was tested with a mutation that breaks the behavior it
asserts.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).