The current source checker reports updateAvailable with behind null when GitHub
compare cannot count staged commits; the app-update predicate demanded an integer
behind and refused every HEAD->NEXT leg. Require behind > 0 only for the historical
shape without updateAvailable.
v2026.6.19's DMG bootstrap runs the same install.sh that writes .install_method,
so its macOS leg saw a dirty tree. Share the Linux driver's guarded exclude through
source-driver.sh and apply it before every app-driven macOS update.
Every leg installed a release tag and updated to HEAD, so nothing ran
HEAD's installer on an empty machine (where all four 2026-09-23
install.ps1 breaks lived) and nothing exercised the updater we ship
today -- tag legs run the OLD build's updater handing off to HEAD.
The generator appends a HEAD -> NEXT start after the sampled tags, so the
column runs wherever the update legs run (dispatch and stable-release).
NEXT is a reserved update ref: the drivers mint a child of the install
commit that adds one marker file, written to the object store only, which
the local bare clone carries into serve.git.
On Windows the HEAD leg takes every git.exe dir off PATH and installs no
remote get-url shim: Get-PinnedGit returns any git on PATH, so either one
skipped pinned-git staging. launch-from-spec's HEAD observer now uses the
driver's real git so it cannot poll '' forever on that leg.
The app-update driver requires sole ownership of Electron's single-instance
lock. A competing packaged process prevents Playwright from reaching app ready
and hides update coverage behind a launch timeout.
Reuse the existing quiescence helper before the app-update launch and the
post-update checkpoint.
The linux hermes-update legs from v2026.3.12 and v2026.5.29.2 end with the
tree at HEAD and no .hermes/bin/* launcher, so the post-update checkpoint
fails at "no usable installed command at HEAD" (macOS has the same check:
"no installed command after update"). The update itself is fine -- the sha
check right above it passes.
Those releases cannot complete in-process: their update path reaches no
retired-hook seam (3.12 has none; 5.29.2's only one is gated to darwin +
cua-driver on PATH). What completes them is the NEXT ordinary startup:
hermes_bootstrap runs prepare_launch() before importing anything, which syncs
PM, publishes the launchers and re-execs into the managed interpreter.
So drive that startup before the strict checkpoint, without the
lazy-install ban, and only when the launcher is missing:
- a healthy update still gets judged strictly by the checkpoint below;
- `--version` probes keep HERMES_DISABLE_LAZY_INSTALLS, so a probe can never
complete an unfinished update -- only a real startup may.
source_hermes_for_startup() hands out the legacy venv command for exactly
that startup; source_hermes stays strict for every probe.
Verified: bash -n and shellcheck (-S warning) clean, no new warnings vs HEAD.
The driver's probe saw OPENAI_BASE_URL present immediately before AND after the
snapshot, while the report called the same key an ADDITION -- two descriptions of
one file that cannot both be true. verify-user-state.py gains an `env-keys` mode
that lists every .env IT judges, with absolute paths and key names (names only),
and the driver prints that view directly after its own snapshot probe.
One run then answers whether the snapshot recorded the root .env, a profile-scoped
one (profiles/<p>/.env is judged too -- the local check prints both), or a file
rewritten in between. Verified locally: the mode prints
judged <home>/.env: MOCK_API_KEY OPENAI_BASE_URL
judged <home>/profiles/p/.env: OPENAI_API_KEY
and tests/scripts/test_verify_user_state.py stays 20/20.
bash resolves functions in execution order, so the .env probes called
env_key_names at the user-state snapshot (~line 319) before its definition
(~line 464): "command not found", exit 127 under set -e, and the driver died
silently -- every leg red, including two that had been green for many runs. It is
now declared with the early helpers (line 189).
Audited all three helpers added here the same way, which is the check that catches
this class: definition line < first non-comment use line. close_running_desktop
(476 < 492) and collect_install_side_logs (359 < 454) were already correct.
A lost provider write surfaces only as "the upgrade changed the user's own state"
minutes later. The run now prints the KEY NAMES (never values) after the provider
configure, after the snapshot, and after the update, and asserts OPENAI_BASE_URL
survived both the configure and the snapshot -- so the step that clobbers it is
named in the failing log instead of inferred. Observed signature: .env at 26403
bytes (the OLD installer's .env.example copy) with no OPENAI pair, and the update
then ADDING both keys (+70 bytes).
The leg reported OPENAI_API_KEY/OPENAI_BASE_URL as ADDITIONS after the snapshot:
the snapshot had missed them, so something between the mock start and
preserve_before_upgrade rewrote .env (the checkpoint's app launch, or the CLI step
that produces user state). Rather than guess which, the provider is now
re-pointed immediately BEFORE the snapshot -- the snapshot records what the
upgrade actually starts from -- and the app-update branch no longer calls
mock_start at all, because that point is inside the window the verifier judges and
any rewrite there is attributed to the upgrade.
Two mocks per leg made the checkpoint assertion unsatisfiable. The app's own log
shows it talking to http://127.0.0.1:43475/v1 and completing a turn there
(finish_reason=stop, api_calls=1), while the driver's mock -- whose transcript we
collect -- listened on 46723 and the smoke asserted against that URL. 43475 was
the checkpoint's mock, started by source-desktop-smoke before the driver's own
mock_start at line 292; the app stayed configured for it.
The mock now starts BEFORE the first desktop checkpoint, so there is one mock per
leg and HERMES_E2E_MOCK_URL is set when the checkpoint's smoke looks it up. It is
also written into the provider config before preserve_before_upgrade snapshots
the home, so no provider state moves inside the verified window.
close_running_desktop called `log`, which exists in scripts/install.sh but NOT in
this driver (which defines step/fail/log_group/ok). Under `set -e` the 127 killed
the driver silently right before the new-phase checkpoint -- "line 451: log:
command not found ... Process completed with exit code 127" -- with no assertion
printed, on a leg whose update, user-state and plugin checks had all passed.
printf to stderr matches fail()'s style and depends on no helper. bash -n cannot
catch an undefined command, which is why the syntax check passed for the first
version.
The previous commit stopped the app-update branch from starting a second mock,
which fixed the user-state failure but broke that leg's pre-update chat: it
failed with "The mock must receive this checkpoint prompt after the send" while
mock.log showed ZERO requests. The second start was load-bearing for the CONFIG
write, not just the server -- the app reads config.yaml/.env, not
HERMES_E2E_MOCK_URL, and the fresh harness alone left it pointing at a dead
endpoint.
mock_start now reuses a mock whose pidfile is still alive and always rewrites
the provider config for the live URL. Reusing the same URL keeps .env
byte-identical, so the user-state verifier no longer sees an equal-size
OPENAI_BASE_URL swap after its snapshot, and the app-update branch keeps its call.
Verified against the real code: two consecutive mock_start calls report "already
live", return an identical URL, and leave .env's OPENAI_BASE_URL unchanged.
The app-driven update leg failed the post-update smoke with "electron.launch:
Process failed to launch!", and Playwright's own call log showed why: the child
DID boot (inspector attached, "[hermes] install stamp: c47e76d73b09"), then the
websocket closed at code 1006 and the process exited 0. A clean immediate exit
right after boot is Electron's single-instance lock -- the pre-update instance
that drove the update is still alive, so the second process quits.
The smoke clicks "Update now" in its own window and the updater may relaunch it,
so the instance survives the update. The post-update checkpoint must own the only
instance: close anything still running from this install before launching.
pgrep absent degrades to a no-op rather than failing the leg.
The app-update branch started a SECOND mock inference server after
preserve_before_upgrade had already snapshotted the home. A fresh instance takes
a new port, so .env's OPENAI_BASE_URL was rewritten -- same byte length,
different value -- and the user-state verifier reported "the upgrade changed the
user's own state ... 0 deleted, 1 modified": the harness's own write, blamed on
the upgrade.
The new key-name diagnostic named it in one cycle:
.env: variables changed=OPENAI_BASE_URL
The provider is already configured for the run above (or comes from the caller's
HERMES_E2E_MOCK_URL), so the branch asserts a mock exists and reuses it. If it
died it now fails loudly, rather than silently reconfiguring provider state
inside the window under verification.
The app-driven update's transcript lands in the app's own HERMES_HOME
(logs/desktop-update-handoff.log, plus the XDG userdata copy) and the driver
already knew to snapshot those places -- but the snapshot ran AFTER
`fail "app-driven update exited $rc"`, so the one leg that needs the transcript
was the one leg that never collected it. The block is now a callable that the
app-update branch invokes the moment its observer returns, with a flag so the
later call stays a no-op.
arm_source_redirect installs a git PATH shim that deliberately reports the
OFFICIAL origin for `remote get-url origin` so fork detection stays quiet;
the file:// redirect is only visible through the real binary it exports in
HERMES_E2E_REAL_GIT. The check asserted on the shimmed view, so it could
never pass. Observed: 'origin transport https://github.com/NousResearch/
hermes-agent.git is not redirected'.
The check hardcoded only the HTTPS URL, but installs clone over SSH in
some environments, so a correct install failed with 'origin is configured
as git@github.com:... not the official URL'. Both forms are official; the
invariant is that the CONFIGURED url is an official one while the
TRANSPORT url is redirected, not which of the two forms it is.
The install/update legs asserted plenty about the code -- the checkout
landed, the version bumped, the desktop artifact exists -- and nothing
about the user's own state. An upgrade that ate auth.json or truncated
state.db would have passed every leg.
Adds a read-only, stdlib-only verifier (snapshot/verify) plus the hooks
that drive it around the real upgrade, on the POSIX and Windows drivers.
The state it defends is produced through the ordinary CLI
(hermes chat -q / auth add / profile create), never seeded by the
harness, and each action asserts it actually landed so a leg cannot
'pass' while testing nothing.
Judged: config.yaml, .env, auth.json, state.db, gateway_state.json and
the user's trees. state.db is compared by row counts, not bytes -- a
live SQLite file moves for benign reasons. The bundled skills/ tree is
recorded but never judged (the product re-syncs it), and plugins/** is
left to verify-plugin-preservation.py.
Share the real composer, provider-witness and completed-reply check across
post-build bundle smoke and desktop-bearing install/update checkpoints.
Keep native automatic-relaunch proof separate from post-update chat.
Download receipt-bound artifacts without release credentials and install
DMG, ZIP, MSIX and universal MSIXBUNDLE on each native architecture.
Split Windows assembly from feed publication; publish tested bytes only.
Bind candidate smoke results into the manifest used by stable promotion.
Verify historical/source provenance without assuming a version IPC commit,
strip CI identity from source build children, and use the actual Electron
PID rather than Playwright's Windows launcher wrapper.
Validation: real Linux Electron chat and sequential OLD/NEW source smoke
with preserved history; 145 targeted Python tests and 14 JS tests passed;
TypeScript, shell/PowerShell parsing and workflow checks passed.
Native macOS/Windows deployment and historical upgrades need Actions proof.
Run historical updater completion in a fresh interpreter so cached imports
cannot revive retired dependency installers. Share Git and ZIP completion,
carry receipt and recovery state, and preserve child exit status.
Route plugin admission, binary acquisition, desktop launch and build paths
through PM. Replace redundant helpers and tests with real worker, package,
publication and launch checks. Keep the shipped compatibility surface fixed.
Targeted Python and desktop checks pass. Native update journeys and fresh
production image qualification remain pending. This is a checkpoint before
those acceptance runs.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Add a Hermes-owned upgrade-preservation contract to the install E2E
harness: a tagged upgrade must not delete or modify anything under the
active home's plugins/** or any profile's plugins/<name>/plugins tree.
- tests/install/e2e-assets/verify-plugin-preservation.py: standalone
stdlib-only READ-ONLY verifier. snapshot records every entry (kind,
size + sha256, link target, recursive fingerprint of symlinked
external targets so the externally-owned sidecar witness is covered)
including empty dirs and the roots themselves (lstat, so dangling
root links are still scanned); verify fails on deletion or
modification, treats unreadable paths as hard errors, and refuses an
empty snapshot as inconclusive. Runs under python3/python on all
three driver platforms.
- tests/install/e2e-assets/preserve-plugins.sh: POSIX/macOS hooks.
After install: seed controlled non-dependency directory fixtures
(mnemosyne-wrapper marker + payload + symlinked runtime, second
profile plugin tree, external witness outside the home; no pyproject
in the scanned root, nothing downloaded) and snapshot. After update:
verify; violation fails the leg. Seeding never clobbers a populated
wrapper without the expected marker.
- installer-script-e2e.sh / macos-desktop-e2e.sh / windows-e2e.ps1:
wire before/after hooks into the update leg, and add --update-ref
(-UpdateRef on Windows) so a leg can target an actual stable tag
instead of HEAD; HEAD-upgrade legs keep the HEAD label. Matrix and
workflows unchanged.
- tests/scripts/test_verify_plugin_preservation.py: 17 unit tests on a
real temp filesystem covering clean survival, file/marker deletion
and modification, whole-root deletion, empty-dir deletion, symlink
repoint, external witness tamper, read-only guarantee, empty-snapshot
refusal, unreadable-path hard error (POSIX), and the exact CLI
round-trip the drivers use.
- tests/install/README.md: document the contract, the hooks, and the
stable-to-stable rule.
Share zoom preparation across both launchers and stage the helper with each driver. Use the Appearance preference bridge and verify page zoom rather than display DPR.
Real Electron regressions fail with transient zoom on focus/navigation and pass with persistence. Repeated click-throughs, onboarding unit tests, E2E typecheck and lint passed. The historical onboarding timeout and full install/update matrix remain unverified.
On CI runners the rebuilt app cannot self-relaunch (chrome-sandbox ownership), so it parks on a reopen-to-finish overlay and never exits; a bare app.close() waits on it forever. Close with a bounded ladder instead (graceful, SIGTERM, SIGKILL), sweep the app's descendant processes so the spawned backend cannot keep writing into the install dir, then relaunch from the same captured spec and require the updated app to present a live UI. The driver gets explicit exits plus an unref'd self-deadline so it cannot outlive its own test, and the posix driver quiesces the install dir (processes whose cwd is inside it) before the head desktop smoke because the in-app update's npm is detached from the Electron process tree.
ts_prefix captured its own start per pipe, so each log's [+MM:SS]
was relative to that log's creation - the install log started at
+0:00, the app-update log at +0:00 too, and no single offset could
align all files with the recording. The drivers now stamp TS_BASE=
$SECONDS once at start and ts_prefix stamps every line relative to
it (falling back to its own start when unset): all logs in a leg
share the driver's clock, and playback.html's one offset slider
(recording start vs driver start) aligns every file at once. The ps1
twin was already driver-relative (TsPrefixStart is captured at
dot-source time, near the driver's top).
Verified: two pipes in one shell - the second starts at [+00:03],
continuing the driver clock instead of resetting to [+00:00].
A static single-file player (tests/install/e2e-assets/playback.html):
?zip=<artifact zip url> unzips in-browser (JSZip), plays the screen
recording with a timer pinned top-left, and renders every *.log with
video<->log sync: the video follows the driver's transcript, clicking
a log line seeks the video. A sync-offset slider aligns the recording
start (ffmpeg comes up first) with the driver's relative clock.
Sync axis: drivers now prefix every transcript line with [+MM:SS]
relative to driver start (ts-prefix.sh / ts-prefix.ps1, pipe-safe
under pipefail / relaxed EAP). Browsers cannot play Matroska, so each
leg remuxes recording.mkv -> recording.mp4 (-c copy, no re-encode)
before the artifact upload, on all three OSes.
Verified end-to-end in a real browser against a generated artifact
zip: zip load, mp4 playback, timer, tab switching, follow-sync at
t=6/t=12, click-to-seek, autoplay policy (expected NotAllowedError on
synthetic play; real clicks fine).
Also fixes the shim fail message's dead variable ( ->
observed_git_url) in both posix drivers.
The app-update legs died on the onboarding overlay - a fullscreen div
that intercepts every click (the Settings click timed out under it).
The old plan seeded a fake provider key, which lies: the overlay
vanishes but the app is broken. Instead the driver now runs the
desktop E2E suite's own mock inference server
(tests-js/scripts/mock-server.ts, zero deps, bare-node type
stripping >=22.18) and configures it into HERMES_HOME byte-for-byte
like the dev:mock flow: config.yaml provider + MOCK_API_KEY env. The
app boots genuinely configured - no overlay, real chat surface.
e2e-assets/mock-provider.{sh,mjs} own start/stop (pid + url files;
the wrapper lives until SIGTERM - gating on stdin-close made the
server die instantly, a background process's stdin is already EOF)
and the config write. Wired into the posix script driver's
hermes-desktop-app-update arm and the macos driver's update phase
(both app-update methods). The Playwright flow keeps its defense-in-
depth: the real escape hatch ('I'll choose a provider later') and the
verified 'Open settings' selector.
Probed locally: models + streamed/non-streamed completions answer,
server stops cleanly on kill.
windows-desktop-gui-e2e.ps1 and windows-installer-script-e2e.ps1 fold
into tests/install/windows-e2e.ps1 with orthogonal -InstallMethod and
-Route axes: the install phase dispatches on one, the update phase on
the other, and shared workroot state carries how OLD landed - so any
implemented update method can follow any implemented install method.
Implementing a new pair is now a driver function plus a gate edit,
never a new job.
The run workflow collapses to ONE inner job whose if: is the
implemented-pairs table. Newly cheap pairs go live with the merge:
desktop-installer@latest -> hermes-update / installer-script /
installer-script+desktop / hermes-desktop-app-update
installer-script(+desktop) -> hermes-desktop-app-update
installer-script+desktop -> open-app-update (the -IncludeDesktop
install registers real Start Menu / Desktop shortcuts)
Only desktop-installer@latest as an UPDATE method stays a declared
TODO. scripts/windows_e2e_harness.ps1 executes the parse/parameter/
dispatch checks under pwsh before any Windows runner spins up.
git show | grep -qF exits at grep's first match; install.sh is ~140KB
with the flag strings in the first few KB, so git show takes SIGPIPE
on its next write and the pipeline reports 141 under set -o pipefail.
The probe answered NO for flags the ref HAS - timing-dependent, green
without pipefail (every local check), red on the runner.
It hid while a probe miss just meant omitting --skip-browser; the
first probe where NO is a hard failure (--include-desktop) exposed it
on its first CI leg. Buffer git show into a variable and grep the
string: git always completes, grep judges bytes.
Playwright must own the spawn (it needs the inspection pipe), but
hermes desktop is not just build+launch - stamp checks, integrity
gates, sandbox fixups, and a constructed child environment. So the
driver intercepts the product's own launch: a sitecustomize.py on
PYTHONPATH (opt-in via HERMES_E2E_CAPTURE_LAUNCH) wraps subprocess.run,
captures argv/cwd/env at the spawn site, and fakes success instead of
spawning; launch-from-spec.mjs then _electron.launch-es exactly that
spec and clicks Settings -> About -> Update now. Completion is product
state, not a Playwright event: the handoff result file or the checkout
reaching the expected sha (source installs write no result file).
Ships with the driver, so it works unchanged on every sampled OLD ref
- no product flag, no pre-flag fallback split. Both launch shapes are
matched (npm exec electron / packaged exe under apps/desktop/release);
npm BUILD calls pass through untouched. Exit 0 without a capture fails
the leg: a version that never reached its launch must not pass.
Probe-the-probe: scripts/launch_capture_probe.sh runs control rows
(no opt-in, non-launch argv) and both treatment shapes - all green
locally. Gate flips on the shared run workflow for linux/macos;
windows adopts the same path with the driver restructuring.
The one-liner with its desktop stage opted in (--include-desktop /
-IncludeDesktop) is a real install kind, distinct on both sides:
on windows the stage builds Hermes.exe AND registers Start Menu /
Desktop shortcuts - a second path to a hand-launchable app - while
on linux/macos it builds into the checkout and registers no OS
entry point.
Declared on every OS and driven by both script drivers: the drivers
pass the flag through (hard failure if the ref predates it - the
tag-has-desktop gate already skips pre-desktop tags upstream) and
assert the built app exists under apps/desktop/release afterwards.
The run-workflow gates run +desktop pairs only on desktop-bearing
tags; app-update pairs from +desktop installs stay declared TODOs.
Each script-driver leg now proves the installed CLI can build the
desktop app, after the install phase and again after the update.
--build-only runs the full desktop pipeline and stops before the
launch - the same call hermes update makes. Old releases that predate
the flag skip the phase after a --help probe of the installed binary.
Actually launching the app is a TODO: it needs the spawn-interception
launcher and, on linux runners, a virtual display.
Also adds tests/install/README.md describing how the test family
works: the four layers, the git-redirect isolation, the phases, the
probe-do-not-assume rule for old versions, skips, triggers, artifacts.
All three drivers (installer-script-e2e.sh, windows-installer-script
-e2e.ps1, windows-desktop-gui-e2e.ps1) now emit the complete
install/update transcript into the job log wrapped in ::group::/
::endgroup:: - collapsed by default, one click to expand, win or
lose. Replaces the tail-50-only-on-failure pattern: a green
install's transcript is how you diagnose the leg that fails next,
and the artifact download was the only way to see it before. The
GUI driver's bootstrap-installer.log / desktop-update-handoff.log
tails become full folded dumps too.
The POSIX sibling of windows-desktop-gui-e2e.ps1, sharing its staging
trick: bare-clone the checkout to serve.git, park main at OLD, point
every git process at it with url.<file://serve.git>.insteadOf in a
driver-owned GIT_CONFIG_GLOBAL. The installer and updater run
byte-for-byte against their real URLs; no bwrap, no MITM proxy, no
TLS interception - a disposable CI runner IS the sandbox, so the
same driver can run on macos-latest unchanged.
install.sh is not curl'd: the install leg runs the copy shipped AT
the OLD ref (what a user who installed then actually executed), the
installer-script update leg runs HEAD's copy (what the website
serves at update time). Flags are probed per-ref (--skip-browser is
newer than sampled tags); HOME is isolated because old installers
hardcode ~/.hermes; .skip_upstream_prompt suppresses the updater's
fork prompt on the file:// origin; the dirty-tree guard checks
tracked files only (-uno) since untracked files cannot leak into a
bare clone.
Verified locally end-to-end: v0.20.2 installed via its own
install.sh (uv, managed Python, Node, venv; hermes --version OK),
served main advanced, hermes update landed the checkout on HEAD
with a working hermes. bash -n + shellcheck clean.