Review follow-up: compute the parked-branch case once in the rollback,
fold the baseline comments into one, reuse pull() and a new complete()
helper in the tests, and say in the docs that a rollback restores the
commit the checkout ran before the pull on every update path.
Review follow-ups for the branch-switch rollback:
- If the parked branch cannot be checked out again (for example, another
worktree holds it), restore its commit detached. Before this, the install
stayed on the broken update branch, and the recovery hint repeated the
failing command.
- Any `merge-base --is-ancestor` result other than "contained" keeps the
strict baseline. rc 128 used to fall back to the lenient one, which
contradicted its own comment.
- `_pull_updates` returns the baseline it verified, and
`_apply_pulled_update` checks against it. Before, that function
recomputed the baseline with a second rev-parse and merge-base against a
hardcoded `origin/<branch>`, which a fork push in between could move.
- Drop a `commit_count == -1` check that could never be false.
- Update the rollback description in `updating.md`.
- Tests share the fixture `git` helper, a repo-init helper and a `pull()`
wrapper instead of repeating those calls.
The default browser_exec tool ran the `browser-use` CLI from a PM side
environment (browser-use==0.13.10 in <home>/environments/browser-use),
provisioned by the installers and `hermes update`. Sealed Desktop payloads
skip that step, so the Desktop app never had it and silently fell back to
the built-in tools; the side env was also per-profile and 225 MB.
The CLI's execution path is only `browser_harness.run.main()`; the
browser-use agent framework (anthropic/openai/google-api pins, 93 MB of
googleapiclient) is never imported. browser-harness itself is 2.6 MB of
pure Python whose pins (Pillow 12.3.0, websockets 15.0.1) already match
Hermes's own, so it becomes a core dependency and runs on sys.executable:
- pyproject/uv.lock: browser-harness==0.1.13 (+ cdp-use, fetch-use).
- _find_cli() returns [sys.executable, -m, browser_harness.run]; the child
env points PYTHONPATH at the harness site dir (the Desktop store
interpreter boots without a venv and the harness daemon re-runs
sys.executable), replacing whatever the agent inherited.
- The side-env provisioning (install_cli, the update/installer step) goes.
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".
- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
installed copy or None. Every caller already pm.ensure()s on None, so a
missing runtime is now provisioned instead of silently borrowing the
user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
another npm.
- source_build.source_product_current: run the freshness reader only with
PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
artifact; delete the "uv on PATH if new enough" developer shortcut.
Computer use and browser use are meant to work out of the box. The PM rewrite
(3d12e86ef1) and the MSIX installer rework (47f4ab3a17) dropped the
install-time cua-driver fetch (7060ac7bed) and the Browser Use CLI install
(baa6b2e34d). On a fresh install the computer_use check_fn therefore stayed
False, so the tool never reached the model and its lazy ensure could not fire,
and browser_exec quietly fell back to the built-in tools.
- cua-driver is a default PM package, so the installers, a bare
`hermes pm install` and `hermes update` carry it on every target it builds
for. Adds the missing Android gap (the lock has no bionic artifact).
- The shared default-tool step, which the installers (via source completion)
and `hermes update` both run, provisions the Browser Use CLI for the default
and explicit Browser Use backends. `--skip-browser` declines it along with
agent-browser, and `off`/Camofox never use it.
- Installers regain --skip-computer-use / -SkipComputerUse (recorded as
`--without cua-driver`).
- The update message stops calling every default "browser tools".
Review follow-ups on the Windows skip:
- Most people who hit the boot loop launched Desktop from the Start menu and
never open a terminal. The in-app update also rebuilds the app: the
Windows shim waits for Desktop to exit, and the already-up-to-date path
still completes with desktop=True. The notice and updating.md now name
Update now in Settings -> About next to `hermes desktop`.
- cmd_gui took the skip's None as "no launchable app was found". A build
that fails raises instead of returning None, so None there only means the
skip, and `hermes desktop` now reopens the app that was kept.
Tests fold to the two invariants, each red when its half of the fix is
reverted: the stop spares its ancestor Desktop on every platform, and a
Windows packaged build under its own Desktop is skipped (now driven through
the real _desktop_ancestor_in with a fake process tree, instead of a stub).
click-session flake (#97982): a bare scrollIntoView() smooth scroll could be dropped under load, and the script
slept a fixed 3000ms before reading state. Extract click-session-helpers.mjs:
instant centered scroll, a bounded poll-for-composer loop instead of the
fixed sleep, and correct nested CDP envelope unwrapping (the old read logged
undefined).
Intel-Mac installer docs (#99033): the Hermes-Setup.dmg bootstrap installer
is built for Apple Silicon only, so Intel Macs hit "not supported on this
Mac". The desktop release pipeline already builds a native darwin-x64
bundle, so the docs now scope the arm64 limit to the bootstrap installer and
name the darwin-x64 bundle (or the CLI plus `hermes desktop`) as the Intel
path, in the desktop README and the platform-support build-targets section.
cua-driver autostart opt-in (#97389): Windows installs registered the
cua-driver-serve scheduled task on every install, with no opt-out, and
treated the task as an install-readiness requirement. Gate the
install-ready check and _repair_cua_driver_autostart_windows on the new
computer_use.autostart config key (default false = on-demand, fails
closed), extract the registration PowerShell into a testable helper, and
document the opt-in (EN + zh-Hans). Windows-only code path: unit tests
cover the registration args and the config gate; live Windows
verification pending.
The APT package does not work at the moment and a fix is in progress.
Warn at the top of the install guide so users do not burn time on steps
that will fail.
Sitting a plugin out after a fetch failure left the recorded stamp stale on
purpose, so every launch resynced, and the spawned-process guard in
prepare_launch raised "dependency sync left this install out of date".
One retry covers a blip; after that the plugin is disabled with the reason,
and hermes plugins enable restores it. Only requires_hermes still sits out,
which boot skips the same way.
An update disabled any plugin that failed a trial build or its
requires_hermes check. Both can be about us, not the plugin: an untagged
source checkout reads as an older release (#122054), so requires_hermes
misjudges a fine plugin, and a download failure says nothing about the
plugin's code. Those now sit the plugin out of the build: config stays
untouched and it rejoins once the cause clears.
Disabling still happens on evidence about the plugin: requires-python vs
the pinned interpreter, manifest_version, an invalid declaration, a uv
resolution conflict, or its own build backend failing (new BuildFailure,
keyed on uv's 'The build backend returned an error').
enabled_member_dirs now skips a requires_hermes misfit instead of
raising. Boot's currency check raised on it before any sync could run, so
the launch path never reached the update sync. The loader skips such a
plugin anyway; admission still refuses enabling one.
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.
Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
A source checkout without hermes_cli/source_check.py predates release
channels, so the only line it can be on is git. The desktop treated the
missing probe as unsupported and parked the user on a manual
`hermes update --help` card ("This checkout predates desktop
source-channel checks"), so an older non-bundled install could never
update itself from the app.
Report such a checkout as tracking main with an update available and let
apply take the normal git handoff. That update brings in the probe, so
later checks resolve normally. A probe that exists but fails still throws.
The flag used to exit 1 as retired. Now that PM installs the browser tools
by default, it maps to `pm.cli install --without agent-browser`, which later
installs and `hermes update` honour.
- tools/browser_tool_install.py: keep pm-clean's frozen old-updater stub; main's
UTF-8 decode fix touched only the npx prefetch body it replaces.
- tests/hermes_cli/test_update_scoped_reconciliation.py: keep pm-clean's test
subset (catch-up rides the PM completion owner) and take main's gateway-less
host evidence (#120740): the updated seed that holds the host at a running
gateway, and the two gateway-less matrices for the source change that merged
cleanly into update_cmd_fleet.py.
* refactor(update): one predicate for serve rows outside the gateway matrix
The inventory branch of `_marker_only_restart_obsolete` inlined the rule for which
serve/dashboard rows the gateway matrix neither covers nor needs to (supervisor-owned,
or a manual serve handed to its own reminder). The inventory-less branch needs the same
rule for #118742, so it moves to `update_cmd_fleet_gatewayless.runtime_outside_gateway_evidence`
and both branches will read one definition. No behaviour change.
* fix(update): a Desktop-only host settles an inventory-less restart obligation
A host that runs no gateway (the Desktop app alone) can be left with an inventory-less
fleet-restart obligation: an updater that died before recording its inventory, or the
pre-inventory writer. With no owed set, the live gateway matrix is its only evidence, and
on that host the matrix is empty forever, so `_marker_only_restart_obsolete` never settled
and every later `hermes update` exited 1 with "gateways are still off the checkout code"
(#118742).
An empty fleet alone cannot tell that host from one whose gateway the dying update stopped,
so the inventory-less branch now asks the live host, never a historical receipt:
`host_owes_no_gateway_restart` settles only when no profile's `gateway_state.json` claims a
state other than stopped/startup_failed (a gateway that went away without a clean stop keeps
the obligation) and every live runtime sits outside the gateway matrix
(`runtime_outside_gateway_evidence`, shared with the inventory branch). HEAD must still
contain the pulled SHA (`checkout_contains`, same rule as #119367). Probe failures keep it.
Tests: the scoped-reconciliation matrix now holds its host at "the update stopped a gateway"
so it keeps pinning receipt independence; a new host-evidence matrix covers Desktop-only,
clean stop, carried commit, stopped gateway in a named profile, unclassified and
unidentified serves, a gateway row without fleet identity, and a diverged checkout. Two
manual-serve tests that assumed an empty fleet always stays pending now stub the live host
and assert the manual reminder survives the gateway obligation settling.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
---------
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
updating.md never said what requires-python >=3.11,<3.15 with every dep
gated on >=3.14 means for users: only 3.14 runs Hermes, the wider range
only lets pre-PM venvs run one more hermes update, and pip/brew installs
on older interpreters get a dependency-less package.
security.md described truststore but not that the startup CA guard,
SSLConfigurationError and HERMES_SKIP_SSL_GUARD are gone, or that the
CA env vars are still read by plain requests/urllib calls.
pm.plugins_state.read_home_selection raises on any unparsable profile
config.yaml and that blocks dependency prep for every home; plugins.md
and profiles.md now say so.
termux.md only described the canary suite while release_artifacts.py and
the bundled-release workflow publish releases/termux/stable (suite
hermes-stable) for tagged releases. Lead with stable, show how to opt
into canary, and name the channels in the README instead of a bare
"prerelease".
google_chat.md still routed --install-deps through tools.lazy_deps and
HERMES_LAZY_INSTALL_TARGET; the plugin calls pm.sync_venv and the Docker
venv is read-only. windows-native.md named hermes_cli/dep_ensure.py and
install.ps1 -Ensure, neither of which exists; features call pm.ensure.
desktop.md and updating.md promised an automatic corrupt-zip retry and
npmmirror fallback for the Electron download; build_prepared_desktop runs
the build once and a failed product aborts the update, so document the
manual recovery (clear the @electron/get cache, set ELECTRON_MIRROR,
rebuild) instead.
Host-scoped update-restart obligation (061195fac1 / 953b6f6f08 / 3be255eca6)
lands on the PM model: the obligation record, its readers and the legacy
per-home marker compat come in as-is. The catch-up restart path
(`_apply_pending_fleet_restart_catchup` / `_run_pending_fleet_restart`) is
retired here (the fleet restart rides the completion owner), so main's
per-host restart-once guard on that path is not carried; its unit→live
MainPID collapse IS ported into the live post-update systemd pass
(`_restart_systemd_gateway_units`), with the two collapse tests rewritten
against that function (red on the pre-port tree: `_unit_main_pid` absent).
Tests that exercised only the retired catch-up path are dropped.
Desktop: main's shared log-rotation planner replaces the inline constants
in main.ts; the merge keeps our machine-profile import beside its import.
utf-8 → utf-8-sig on the three new BOM-intolerant reads (footguns lint).
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].
Fixes#117246
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.
Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.
Fixes#115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
Follow-up to the cherry-picked #116014:
- Pin per the dependency policy (pre-1.0: `<0.(minor+2)`): httptools
`>=0.6.3,<0.9` (floor = uvicorn[standard]'s own floor), uvloop
`>=0.15.1,<0.24`. `watchfiles>=0.20,<2` already complied.
- Copy uvicorn's own uvloop marker (win32, cygwin, PyPy) plus
`sys_platform != 'android'` so `pip install '.[all]'` on those
hosts does not fail on the extra either.
- `tools/lazy_deps.py` mirrors the `web` extra for the lazy dashboard
install: it also requested `uvicorn[standard]`, so a Termux user
opening the dashboard would have hit the same uvloop build at first
use. The web_server install hint follows.
- `uv lock` regenerated; the lock delta is exactly the pyproject delta.
- Two invariant tests: no Termux-reachable extra (or core, or the lazy
dashboard feature) requests uvloop; `[all]` still does, off Android.
- Docs: troubleshooting entry in the Termux guide.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
Closes the second #97208 atom (qingchuan-x): after a Windows hand-off child refused the
dependency sync (Desktop backend held the venv), the next `hermes update` found git current,
passed the core-import probe — the old release imports fine — and printed
"✓ Already up to date!" while the installed distribution still read hermes-agent 0.20.6
against a 0.21.2 checkout with lagging pins.
`_venv_dependency_set_stale` asks the venv's own interpreter for the installed hermes-agent
version and compares it with pyproject's; a mismatch on the current-checkout path takes
the same repair route as an unhealthy venv. Unknown states (no venv, not installed as a
distribution, probe failure) are never stale, so dev checkouts are untouched.
Review follow-up on the #101600 fix:
- The wait moves from `_cmd_update_impl` to `cmd_update`, BEFORE `UpdateLock.acquire()`.
The child used to run under the parent's marker (handoff pid / ancestor claim); the
parent's exit — which the child now deliberately waits for — released that marker, so
the whole Windows dependency install and gateway resume ran with no update lock. Waiting
first lets the child claim a marker of its own for the tail.
- `SHIM_PARENT_PID_ENV` names the `hermes.exe` launcher ancestor (the process that holds
the shim image open and exits only after reaping this interpreter) when psutil sees one,
else this pid. `_windows_shim_in_process_chain` is split into `_venv_shim_matcher` +
`_windows_shim_ancestor` so the holder pid shares the matcher; docs sentence adjusted.
- The direct-call wait test is replaced by one driving `cmd_update` in a real child with an
older parent: proves the phase starts after the parent died, the child owns the lock, the
handed-off pause token is adopted (no re-discovery) and the env is consumed. Removing the
wait call, the token adoption, or moving the wait back after the lock each turns it red.
`hermes update` run from `venv\Scripts\hermes.exe` hands the dependency install to a
child under the venv Python because the shim cannot replace itself. Two races were left
in that hand-off (#101600, three independent Windows reproductions):
- The child never waited for the shim-run parent. The shim quarantine is a single
`os.rename` with sub-second retries; when the child reached it while the parent was
still alive the rename failed with PermissionError, `ShimQuarantineError` deferred the
whole install and the receipt was marked failed although the checkout had advanced.
- On the legacy re-exec (`_abort_dependency_sync_if_self_locked` ->
`_reexec_dependency_sync_off_windows_shim`, still live for a current checkout with an
unhealthy venv -- i.e. the retry after any failed update) the PARENT resumed the paused
gateways after spawning the child, holding hermes.exe open for the whole relaunch and
raising `RuntimeError: Windows gateway relaunch after update was not verified alive`;
the child then re-ran pause discovery, found that freshly relaunched gateway before its
pid file existed and force-killed it as "without profile mapping".
Now every child spawned off the shim gets `HERMES_UPDATE_SHIM_PARENT_PID` and, at the top
of `_cmd_update_impl`, waits (bounded, 30 s, `(pid, create_time)` identity via psutil)
for that process to exit before scanning holders, pausing gateways or renaming shims. The
legacy re-exec carries the Windows pause token to the child in
`HERMES_UPDATE_GATEWAY_RESUME` and disarms the parent's copy, so the parent exits at once
and the child resumes exactly the fleet the parent stopped instead of rediscovering it --
the same ownership transfer the post-swap hand-off already does (94ced1a2b2).
The wait is host-independent and proven with real processes on Linux; the shim paths run
on the windows-latest lane through the existing `_is_windows`-patched suite.
Fixes#101600
Refs #98010#103031
Supersedes #101630
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
Explain that the matrix shows one row per gateway process, that `served_profiles`
carries coverage for the satellites, and that `hermes gateway restart` now clears the
pending-restart hint on such a fleet.
Two follow-ups on the salvaged commits:
- `_ensure_default_soul_md`: widen the cyclic-only (ELOOP) branch to any symlink the seed
cannot write through — a link dangling into a missing directory raises ENOENT on the same
write and bricked home init the same way (HomeInitializationError -> exit 75 -> supervisor
relaunch storm). Shape: write through the link first (a working link stays operator
wiring), and only on OSError seed the default in place of the link via mkstemp + replace.
Tests trimmed to two invariants (cyclic/dangling replaced; resolving link preserved).
- `_run_pre_update_backup`: the quick snapshot stays best-effort by design (8ed599dc05:
"a broken backup never blocks the update"), but a failure was swallowed at DEBUG level, so
the user only learned from the receipt afterwards that no recovery point existed. Print a
stdout warning with the reason and continue; together with the receipt skip/fail split the
receipt now says "failed" only when a snapshot was requested and not produced.
- docs: updating.md describes the warning and the receipt split.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
The updating guide describes the rename-over-previous-build step; mention that a
scanner briefly holding release/win-unpacked is ridden out with short retries
(#112544) so the paragraph matches the promotion behaviour.
Rework of the salvaged #59942 hook so it fixes the whole #52339 class without
regressing what the Desktop updater already gets right:
- macOS only. The Windows arm rmtree'd a possibly-running NSIS install
(partial deletion under a lock) and windows.ps1 already owns that swap;
Linux packages stay with their package manager.
- No re-sign. The rebuilt release/ bundle already carries the stable local
signing identity from _desktop_macos_relaunchable_fixup; a deep `codesign -s -`
on the installed copy replaced it with a fresh ad-hoc cdhash and reset every
TCC grant. ditto preserves the signature, so nothing is signed here.
- Running bundles are reported, not swapped: Electron loads app.asar and helper
apps lazily, so renaming the bundle away and deleting the old tree crashes
the live app. The detached updater waits for exit; a terminal `hermes update`
with the app open now prints what to do instead.
- Failures are printed as warnings; the old `Path | None` return read every
failure as "Desktop app up to date".
- The refresh also runs on the "build stamp current" path, so a stale
/Applications copy left by an earlier update heals on the next `hermes update`
even when there is nothing to rebuild.
- Core is host-independent (_install_rebuilt_macos_bundles takes paths as data);
the two invariant tests run on every OS instead of `skipif(darwin)` tests that
ran nowhere.
Covers the Desktop-button path too: posix.sh runs `hermes update`, so an app
running from apps/desktop/release/ now refreshes the /Applications copy Finder
launches (the Discord report: new shell right after the update, old shell on
the next Dock launch).