Finish bootstrap uv before PM replaces its store entry. Keep failure
receipts stdlib-only and align the cryptography requirement and override
with the locked version.
Let bundle builders declare launch paths and update ownership. Remove
payload discovery, Store probing, and the unused develop command.
Derive Nix Python from the PM lock and share its provenance stamp.
Document setup, activation, optional dependencies, and distribution
ownership. Targeted Windows tests, relocated runtime launches, Electron
bundling, and bilingual docs builds pass. Native Nix and signed-package
acceptance remain CI gates.
Run the entire CI workflow before Docker build and tests. Require Nix,
native payload smoke tests, install/update E2E and signed-package upgrade
acceptance before publishing. Keep Desktop Playwright E2E deferred.
Archive tested Docker images and signed bundle candidates with provenance
and hashes. Publishers consume those exact artifacts without rebuilding.
Advance stable channels only after all required publications succeed.
Keep canaries on their separate path and reject direct stable-builder
publication that bypasses the gate.
Move shared release transport, manifests and gates to Python. Keep native
Electron adapters in JS and share feed/MIME facts as JSON. Replace the
R2/feed JS implementation and move its protocol tests to Python.
Verified targeted Python and JS tests, real loopback transport and CLI
execution, temporary Git admission, workflow graph lint, and typechecks.
No live stable release was run. Native signing, package upgrades and real
registry/Store promotion still need their release-run receipts. Separate
services cannot promote atomically. A promotion failure keeps the run red.
- driver now stages the manifest via the parent-owned common
bundle-inputs.mjs (bundle-inputs.json written where require_manifest
expects it) and creates WORK_ROOT/LOG_DIR/HOME dirs up front
- HOME is set to the sandbox (HOME_SANDBOX was exported but never used)
- the installed executable is derived from CFBundleExecutable in the
bundle's own Info.plist (packaged name is 'Hermes Bundled', no rename
assumed) and persisted in an install receipt the update phase reads —
no hardcoded Contents/MacOS/Hermes
- codesign -dv display output is read from STDERR via spawnSync combined
streams (execFileSync stdout was always empty)
- identity is verified as the exact CFBundleIdentifier and an explicit
teamId is required and verified exactly on both sides (mirrors the
common resolver's rules); dropped the com.nousresearch.* prefix check
- relaunch proof additionally requires the watcher's appBin to equal the
receipt path
- lipo -archs verifies the installed bundle matches the leg's arch
- feed server binds port 0 and writes a readiness JSON with the actual
port; the driver configures desktop_feed_base_url from it instead of a
fixed 8791; path traversal check uses path.relative with a real
separator-aware rejection
- feed channel derives from the NEW tag (canary tags serve
releases/darwin/canary/<arch->canary-mac.yml); manifest validator
rejects channel-crossing pairs
- sha512 of the release zip is streamed, no whole-file buffer
- serve/watcher pids are globals killed by a top-level EXIT trap
(trap referenced a local under set -u; watcher leaked on failure)
- final PASS line uses the update phase's new_tag (old_tag was out of
scope there)
- tests exercise helper behavior (validators, streamed hash, safeJoin,
per-channel feed layout) instead of source-shape assumptions
Native CI-only arm (Darwin + GITHUB_ACTIONS) proving a real user route:
install the actual signed OLD release zip into an isolated .app location,
verify codesign/team/version/stamp against the parent resolver's
normalized manifest, configure updates.desktop_feed_base_url at a
loopback-only static feed built from the actual signed NEW zip in the
production update-feed contract (update-feed.cjs darwinFeed layout,
per-arch <arch>-stable-mac.yml), click the real About -> Update now under
Playwright, and let the detached external watcher own the automatic
relaunch proof: old pid exits, a NEW pid/birth appears at the same
installed path with no driver launch, the swapped bundle re-verifies
against the NEW manifest side, the relaunched backend answers
/api/health, and isolated user state + plugin fixtures survive.
No internal apply calls, no source artifacts, no re-signing, no public
feed writes; Squirrel.Mac's signature gate stays in charge.
Add a Hermes-owned upgrade-preservation contract to the install E2E
harness: a tagged upgrade must not delete or modify anything under the
active home's plugins/** or any profile's plugins/<name>/plugins tree.
- tests/install/e2e-assets/verify-plugin-preservation.py: standalone
stdlib-only READ-ONLY verifier. snapshot records every entry (kind,
size + sha256, link target, recursive fingerprint of symlinked
external targets so the externally-owned sidecar witness is covered)
including empty dirs and the roots themselves (lstat, so dangling
root links are still scanned); verify fails on deletion or
modification, treats unreadable paths as hard errors, and refuses an
empty snapshot as inconclusive. Runs under python3/python on all
three driver platforms.
- tests/install/e2e-assets/preserve-plugins.sh: POSIX/macOS hooks.
After install: seed controlled non-dependency directory fixtures
(mnemosyne-wrapper marker + payload + symlinked runtime, second
profile plugin tree, external witness outside the home; no pyproject
in the scanned root, nothing downloaded) and snapshot. After update:
verify; violation fails the leg. Seeding never clobbers a populated
wrapper without the expected marker.
- installer-script-e2e.sh / macos-desktop-e2e.sh / windows-e2e.ps1:
wire before/after hooks into the update leg, and add --update-ref
(-UpdateRef on Windows) so a leg can target an actual stable tag
instead of HEAD; HEAD-upgrade legs keep the HEAD label. Matrix and
workflows unchanged.
- tests/scripts/test_verify_plugin_preservation.py: 17 unit tests on a
real temp filesystem covering clean survival, file/marker deletion
and modification, whole-root deletion, empty-dir deletion, symlink
repoint, external witness tamper, read-only guarantee, empty-snapshot
refusal, unreadable-path hard error (POSIX), and the exact CLI
round-trip the drivers use.
- tests/install/README.md: document the contract, the hooks, and the
stable-to-stable rule.
Preserve the PM runtime-repair module boundary. Port the incoming stderr-streaming fix without restoring the deleted managed_uv downloader. Targeted runtime/progress tests and root JS checks passed; desktop typechecks passed.
Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
Playwright 1.58 accepts a Promise-valued waitForFunction predicate before its false result. Await each zoom read explicitly; cover delayed responses and fresh-install startup.
Share zoom preparation across both launchers and stage the helper with each driver. Use the Appearance preference bridge and verify page zoom rather than display DPR.
Real Electron regressions fail with transient zoom on focus/navigation and pass with persistence. Repeated click-throughs, onboarding unit tests, E2E typecheck and lint passed. The historical onboarding timeout and full install/update matrix remain unverified.
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.
The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.
e2e-screen-record comment: hhttps -> https.
^(10|[1-9])$ replaces the two-step guard: the [0-9]+ regex accepted
leading zeros that bash arithmetic then read as octal (010 passed as 8,
08 errored). README cost figures corrected to the generator's real
expansion: 41 legs/tag, 82 at the default 2 tags, update route 8/tag,
first matrix overflow at 15 tags (270 windows entries).
tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).
Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.
README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
The pre-#97052 prompt wedge fires AFTER the updater has already moved
the checkout (stash cleanup, reset, bootstrap refresh all precede the
upstream-remote prompt), so checkout movement cannot distinguish a
wedged updater from a working one. Two CI rides confirmed the check
never fires. The 35-minute wait bound already caps these legs.
The 12-minute tripwire compared against Get-InstalledHead's result, but
on a wedged updater that read can throw (the venv shim lock is held), so
head stayed empty, never equal to the start sha, and the tripwire never
fired... the legs ran to the full 35m bound. Unreadable now counts as
unmoved, gated on two consecutive no-progress polls so a single bad read
cannot fail a healthy leg.
A working updater moves the checkout off the starting sha within its
first minutes; a wedged one sits before its git step and emits nothing
(buffered stdout, no exit), so only disk state can tell them apart
early. Snapshot the sha when the wait starts and fail with the wedge
diagnosis if it has not moved by 12 minutes, instead of waiting out the
full bound.
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.
The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
The last declared TODO: a user whose install is stale re-downloads
Hermes-Setup.exe and clicks Install over the existing install, the GUI
twin of re-running the one-liner. Windows shows the full installer UI on
a re-run (the already-installed fast path is macOS-only), so the existing
AHK install drive applies unchanged; install.ps1's repository stage
fetches the existing checkout forward to what main serves, now HEAD.
Invoke-PhaseInstallGui gains an update mode instead of a parallel copy:
the phase label, proof dir, and expected-sha assertion become parameters,
and the update-is-available assert stays install-only. The bootstrap log
rotates before the re-run so the AHK's completion fallback cannot match
the install phase's old completion line.
A dmg user can also update from the terminal (hermes update) or by
re-running the install one-liner; the dmg arm only routed the two
app-button methods, leaving three declared cells as permanent TODOs.
Port the POSIX driver's method blocks into the macos driver's update
phase (hermes-update with the --yes probe, installer re-run with per-ref
flag probing, the +desktop built-app assert) and open the workflow gate.
The dmg bootstrap driver also learns to recover instead of waiting out
its bound when a stage fails: the bootstrap parks on an error screen
with a Retry button (seen live: HTTP 429 downloading install.sh under
full-matrix runner load), so the driver watches the bootstrap log for
real failure shapes, requires two consecutive error probes before
retargeting the click at Retry's measured position, unlatches when the
log goes healthy, and gives up with the true cause after three retries.
Driver rule learned three times in this suite (lsof +D, git show and
find piped to grep -q): under set -euo pipefail, never feed grep -q
from a pipe... grep exits at first match, the producer takes SIGPIPE,
and a TRUE condition reads as failure. Capture to a variable or test
paths directly.
Verified end to end: run 33506177130, macos slice 13/13 green.
Hermes-Setup is a Tauri app that boots to a setup-choice screen and waits
for a click on Install Hermes before any install work starts; run bare it
blocked until the 120-minute job cap (both dmg legs, every run). Launch it
in the background with the driver's env, read the window geometry via
System Events (position and size need no assistive grant), post a real
CGEvent click with cliclick at the button's measured position (65% of
window height; System Events' own click needs assistive access the runners
deny), then wait for the full install to land: checkout, venv console
script, AND the built Hermes.app, since the cleanup trap would otherwise
kill the installer before its desktop-build stage. Bounded at 45 minutes
with desktop screenshots on every phase and failure. Adds a macos-desktop
dispatch route so this arm iterates without the full matrix.
Verified end to end: run 33407401698, both dmg legs green (first ever),
macos slice 10/10.
Fires only on the failure path, at most every fifth iteration: what wins elementsFromPoint at the button's center, cluster and bar rects with computed z-index, the titlebar CSS vars, and window size with dpr. Makes a CI-only click interception attributable from the log alone.
On CI runners the rebuilt app cannot self-relaunch (chrome-sandbox ownership), so it parks on a reopen-to-finish overlay and never exits; a bare app.close() waits on it forever. Close with a bounded ladder instead (graceful, SIGTERM, SIGKILL), sweep the app's descendant processes so the spawned backend cannot keep writing into the install dir, then relaunch from the same captured spec and require the updated app to present a live UI. The driver gets explicit exits plus an unref'd self-deadline so it cannot outlive its own test, and the posix driver quiesces the install dir (processes whose cwd is inside it) before the head desktop smoke because the in-app update's npm is detached from the Electron process tree.
The boot-time auto-check can fail transiently and latch the error UI even though a fresh check succeeds. Click Check now whenever it is clickable, re-test Update now, 3 minute ceiling. On final failure, pull the update status over the same IPC the About panel uses so the log carries the real git error instead of the UI's generic one.
The app ships a 90% UI-zoom default, so a fresh install renders at devicePixelRatio 0.9 and Playwright's input coordinates land ~10% off target on the CI runners: clicks aimed at the titlebar settings gear hit the bar beside it, which read as a phantom overlay interception and failed every app-update leg. Set zoom to 100% through webContents after boot, retrying until dpr reads 1 because the boot path re-applies the default asynchronously.
The overlay mounts in phases (boot-progress card, then the provider picker) and the settings gear reports visible while sitting under it, so a one-shot dismiss probe and a visibility gate both race it. Alternate short-timeout dismiss clicks with settings clicks until a settings click lands; a landed click is proof the overlay is gone. Also pick the window that renders the app UI instead of trusting firstWindow(), which can return a helper webContents.
The get-url shim self-check referenced REPO_DIR, which is never defined
(the script defines REPO_ROOT). Under set -u the stage phase dies
immediately, so every macos desktop-installer@latest leg fails before
install. Introduced in 368164d7 with the shim itself.
Verified: bash -n clean; shim + self-check block replayed standalone
with the fix and passes.
ts_prefix captured its own start per pipe, so each log's [+MM:SS]
was relative to that log's creation - the install log started at
+0:00, the app-update log at +0:00 too, and no single offset could
align all files with the recording. The drivers now stamp TS_BASE=
$SECONDS once at start and ts_prefix stamps every line relative to
it (falling back to its own start when unset): all logs in a leg
share the driver's clock, and playback.html's one offset slider
(recording start vs driver start) aligns every file at once. The ps1
twin was already driver-relative (TsPrefixStart is captured at
dot-source time, near the driver's top).
Verified: two pipes in one shell - the second starts at [+00:03],
continuing the driver clock instead of resetting to [+00:00].
The merged view interleaves every *.log in time order across the
whole leg, filename-prefixed per line (green .file span), and syncs
like any other tab: the video highlights the right file's line at the
right moment, clicking a line seeks the video. Untimed lines inherit
their file's previous timestamp so intra-file order survives the
sort; the merged tab is first and the default.
Verified in a real browser with an interleaved synthetic zip (two
logs, overlapping timelines): time-order merge, prefixes, follow-sync
at t=22 highlighting the app-update.log line, click-to-seek to t=30.
Windows transcripts were ZERO bytes: ts-prefix.ps1 formatted with
{0:D2}, but Floor() returns a double and the D specifier is
integer-only - it threw per line, and under the driver's relaxed EAP
every line errored into the void. {0:00} fixes it (custom numeric
format works on doubles). Reproduced the exact pipeline locally
(empty file + Format specifier invalid), verified the fix produces
prefixed merged stdout+stderr with exit code intact. That is also
why the log timeline never auto-synced: there was nothing in the
files to sync.
The GitHub artifact URL 307s to /suites/... server-side and strips
the ?zip= query param. The player now reads the zip URL from a
#zip= HASH param (client-side, survives the redirect) with ?zip=
as fallback; the hash path was verified in a real browser against a
real leg zip (auto-fetch + boot).
Per ethie's design, one player artifact for the whole run: new
leg-player job uploads playback.html (archive:false) before the
matrix legs, the report job needs it, and each ran cell gets TWO
links - 📼 to the player with #zip=<that leg's logs zip> and ⬇️ to
the raw zip. Per-leg player uploads removed from all three run
workflows.
Each leg uploads playback.html as a single-file artifact (archive:
false) before the driver runs, so it exists even on failure. The
results chart now links every leg that RAN (pass or fail, not skip)
to its player with ?zip= pointing at that leg's logs artifact.
Leg<->artifact mapping: the generator mints a leg_id per matrix entry
(sanitized matrix name, exported legId()), every run workflow names
its artifacts install-e2e-{player,logs}-<leg-id>, and the report job
feeds the run's artifact name->id list to the results renderer, which
rebuilds the leg id from the parsed job name. GitHub does not link
jobs to artifacts, so the deterministic name is the join key.
Empirical finding: GitHub artifact downloads are auth-gated (the
download URL 307s to /suites/... which is 404 anonymous), so a
locally-opened player page cannot fetch the zip cross-origin. The
player now degrades gracefully: ?zip= fetch failure renders a real
download link for the zip (a normal click carries the user's session)
plus a drag-and-drop / file-picker path, and no-param opens as a pure
drop target. Verified in a real browser against a real artifact URL.
Verified: generator emits leg_id, results renderer emits
✅/❌ [📼](...?zip=...) only on ran cells, npm run check PASS,
install tests 36/36, strict tsc PASS, actionlint x4 PASS.