Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
Playwright 1.58 accepts a Promise-valued waitForFunction predicate before its false result. Await each zoom read explicitly; cover delayed responses and fresh-install startup.
Share zoom preparation across both launchers and stage the helper with each driver. Use the Appearance preference bridge and verify page zoom rather than display DPR.
Real Electron regressions fail with transient zoom on focus/navigation and pass with persistence. Repeated click-throughs, onboarding unit tests, E2E typecheck and lint passed. The historical onboarding timeout and full install/update matrix remain unverified.
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.
The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.
e2e-screen-record comment: hhttps -> https.
^(10|[1-9])$ replaces the two-step guard: the [0-9]+ regex accepted
leading zeros that bash arithmetic then read as octal (010 passed as 8,
08 errored). README cost figures corrected to the generator's real
expansion: 41 legs/tag, 82 at the default 2 tags, update route 8/tag,
first matrix overflow at 15 tags (270 windows entries).
tag-count now reaches the shell via the environment, validated to 1-10
(an apostrophe in the raw interpolation could terminate quoting; above
~14 tags the expansion exceeds GitHub's 256-job matrix limit).
Result-chart cell ranking matches on the leading token: rendered
success/failure cells carry artifact links, so whole-cell indexOf
ranked them -1 and any skip in the map beat a real outcome.
README documents per-run cost, route slice sizes, the tag-count bound,
and a warning against running the GUI drivers outside a disposable VM.
The pre-#97052 prompt wedge fires AFTER the updater has already moved
the checkout (stash cleanup, reset, bootstrap refresh all precede the
upstream-remote prompt), so checkout movement cannot distinguish a
wedged updater from a working one. Two CI rides confirmed the check
never fires. The 35-minute wait bound already caps these legs.
The 12-minute tripwire compared against Get-InstalledHead's result, but
on a wedged updater that read can throw (the venv shim lock is held), so
head stayed empty, never equal to the start sha, and the tripwire never
fired... the legs ran to the full 35m bound. Unreadable now counts as
unmoved, gated on two consecutive no-progress polls so a single bad read
cannot fail a healthy leg.
A working updater moves the checkout off the starting sha within its
first minutes; a wedged one sits before its git step and emits nothing
(buffered stdout, no exit), so only disk state can tell them apart
early. Snapshot the sha when the wait starts and fail with the wedge
diagnosis if it has not moved by 12 minutes, instead of waiting out the
full bound.
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.
The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
The last declared TODO: a user whose install is stale re-downloads
Hermes-Setup.exe and clicks Install over the existing install, the GUI
twin of re-running the one-liner. Windows shows the full installer UI on
a re-run (the already-installed fast path is macOS-only), so the existing
AHK install drive applies unchanged; install.ps1's repository stage
fetches the existing checkout forward to what main serves, now HEAD.
Invoke-PhaseInstallGui gains an update mode instead of a parallel copy:
the phase label, proof dir, and expected-sha assertion become parameters,
and the update-is-available assert stays install-only. The bootstrap log
rotates before the re-run so the AHK's completion fallback cannot match
the install phase's old completion line.
A dmg user can also update from the terminal (hermes update) or by
re-running the install one-liner; the dmg arm only routed the two
app-button methods, leaving three declared cells as permanent TODOs.
Port the POSIX driver's method blocks into the macos driver's update
phase (hermes-update with the --yes probe, installer re-run with per-ref
flag probing, the +desktop built-app assert) and open the workflow gate.
The dmg bootstrap driver also learns to recover instead of waiting out
its bound when a stage fails: the bootstrap parks on an error screen
with a Retry button (seen live: HTTP 429 downloading install.sh under
full-matrix runner load), so the driver watches the bootstrap log for
real failure shapes, requires two consecutive error probes before
retargeting the click at Retry's measured position, unlatches when the
log goes healthy, and gives up with the true cause after three retries.
Driver rule learned three times in this suite (lsof +D, git show and
find piped to grep -q): under set -euo pipefail, never feed grep -q
from a pipe... grep exits at first match, the producer takes SIGPIPE,
and a TRUE condition reads as failure. Capture to a variable or test
paths directly.
Verified end to end: run 33506177130, macos slice 13/13 green.
Hermes-Setup is a Tauri app that boots to a setup-choice screen and waits
for a click on Install Hermes before any install work starts; run bare it
blocked until the 120-minute job cap (both dmg legs, every run). Launch it
in the background with the driver's env, read the window geometry via
System Events (position and size need no assistive grant), post a real
CGEvent click with cliclick at the button's measured position (65% of
window height; System Events' own click needs assistive access the runners
deny), then wait for the full install to land: checkout, venv console
script, AND the built Hermes.app, since the cleanup trap would otherwise
kill the installer before its desktop-build stage. Bounded at 45 minutes
with desktop screenshots on every phase and failure. Adds a macos-desktop
dispatch route so this arm iterates without the full matrix.
Verified end to end: run 33407401698, both dmg legs green (first ever),
macos slice 10/10.
Fires only on the failure path, at most every fifth iteration: what wins elementsFromPoint at the button's center, cluster and bar rects with computed z-index, the titlebar CSS vars, and window size with dpr. Makes a CI-only click interception attributable from the log alone.
On CI runners the rebuilt app cannot self-relaunch (chrome-sandbox ownership), so it parks on a reopen-to-finish overlay and never exits; a bare app.close() waits on it forever. Close with a bounded ladder instead (graceful, SIGTERM, SIGKILL), sweep the app's descendant processes so the spawned backend cannot keep writing into the install dir, then relaunch from the same captured spec and require the updated app to present a live UI. The driver gets explicit exits plus an unref'd self-deadline so it cannot outlive its own test, and the posix driver quiesces the install dir (processes whose cwd is inside it) before the head desktop smoke because the in-app update's npm is detached from the Electron process tree.
The boot-time auto-check can fail transiently and latch the error UI even though a fresh check succeeds. Click Check now whenever it is clickable, re-test Update now, 3 minute ceiling. On final failure, pull the update status over the same IPC the About panel uses so the log carries the real git error instead of the UI's generic one.
The app ships a 90% UI-zoom default, so a fresh install renders at devicePixelRatio 0.9 and Playwright's input coordinates land ~10% off target on the CI runners: clicks aimed at the titlebar settings gear hit the bar beside it, which read as a phantom overlay interception and failed every app-update leg. Set zoom to 100% through webContents after boot, retrying until dpr reads 1 because the boot path re-applies the default asynchronously.
The overlay mounts in phases (boot-progress card, then the provider picker) and the settings gear reports visible while sitting under it, so a one-shot dismiss probe and a visibility gate both race it. Alternate short-timeout dismiss clicks with settings clicks until a settings click lands; a landed click is proof the overlay is gone. Also pick the window that renders the app UI instead of trusting firstWindow(), which can return a helper webContents.
The get-url shim self-check referenced REPO_DIR, which is never defined
(the script defines REPO_ROOT). Under set -u the stage phase dies
immediately, so every macos desktop-installer@latest leg fails before
install. Introduced in 368164d7 with the shim itself.
Verified: bash -n clean; shim + self-check block replayed standalone
with the fix and passes.
ts_prefix captured its own start per pipe, so each log's [+MM:SS]
was relative to that log's creation - the install log started at
+0:00, the app-update log at +0:00 too, and no single offset could
align all files with the recording. The drivers now stamp TS_BASE=
$SECONDS once at start and ts_prefix stamps every line relative to
it (falling back to its own start when unset): all logs in a leg
share the driver's clock, and playback.html's one offset slider
(recording start vs driver start) aligns every file at once. The ps1
twin was already driver-relative (TsPrefixStart is captured at
dot-source time, near the driver's top).
Verified: two pipes in one shell - the second starts at [+00:03],
continuing the driver clock instead of resetting to [+00:00].
The merged view interleaves every *.log in time order across the
whole leg, filename-prefixed per line (green .file span), and syncs
like any other tab: the video highlights the right file's line at the
right moment, clicking a line seeks the video. Untimed lines inherit
their file's previous timestamp so intra-file order survives the
sort; the merged tab is first and the default.
Verified in a real browser with an interleaved synthetic zip (two
logs, overlapping timelines): time-order merge, prefixes, follow-sync
at t=22 highlighting the app-update.log line, click-to-seek to t=30.
Windows transcripts were ZERO bytes: ts-prefix.ps1 formatted with
{0:D2}, but Floor() returns a double and the D specifier is
integer-only - it threw per line, and under the driver's relaxed EAP
every line errored into the void. {0:00} fixes it (custom numeric
format works on doubles). Reproduced the exact pipeline locally
(empty file + Format specifier invalid), verified the fix produces
prefixed merged stdout+stderr with exit code intact. That is also
why the log timeline never auto-synced: there was nothing in the
files to sync.
The GitHub artifact URL 307s to /suites/... server-side and strips
the ?zip= query param. The player now reads the zip URL from a
#zip= HASH param (client-side, survives the redirect) with ?zip=
as fallback; the hash path was verified in a real browser against a
real leg zip (auto-fetch + boot).
Per ethie's design, one player artifact for the whole run: new
leg-player job uploads playback.html (archive:false) before the
matrix legs, the report job needs it, and each ran cell gets TWO
links - 📼 to the player with #zip=<that leg's logs zip> and ⬇️ to
the raw zip. Per-leg player uploads removed from all three run
workflows.
Each leg uploads playback.html as a single-file artifact (archive:
false) before the driver runs, so it exists even on failure. The
results chart now links every leg that RAN (pass or fail, not skip)
to its player with ?zip= pointing at that leg's logs artifact.
Leg<->artifact mapping: the generator mints a leg_id per matrix entry
(sanitized matrix name, exported legId()), every run workflow names
its artifacts install-e2e-{player,logs}-<leg-id>, and the report job
feeds the run's artifact name->id list to the results renderer, which
rebuilds the leg id from the parsed job name. GitHub does not link
jobs to artifacts, so the deterministic name is the join key.
Empirical finding: GitHub artifact downloads are auth-gated (the
download URL 307s to /suites/... which is 404 anonymous), so a
locally-opened player page cannot fetch the zip cross-origin. The
player now degrades gracefully: ?zip= fetch failure renders a real
download link for the zip (a normal click carries the user's session)
plus a drag-and-drop / file-picker path, and no-param opens as a pure
drop target. Verified in a real browser against a real artifact URL.
Verified: generator emits leg_id, results renderer emits
✅/❌ [📼](...?zip=...) only on ran cells, npm run check PASS,
install tests 36/36, strict tsc PASS, actionlint x4 PASS.
A static single-file player (tests/install/e2e-assets/playback.html):
?zip=<artifact zip url> unzips in-browser (JSZip), plays the screen
recording with a timer pinned top-left, and renders every *.log with
video<->log sync: the video follows the driver's transcript, clicking
a log line seeks the video. A sync-offset slider aligns the recording
start (ffmpeg comes up first) with the driver's relative clock.
Sync axis: drivers now prefix every transcript line with [+MM:SS]
relative to driver start (ts-prefix.sh / ts-prefix.ps1, pipe-safe
under pipefail / relaxed EAP). Browsers cannot play Matroska, so each
leg remuxes recording.mkv -> recording.mp4 (-c copy, no re-encode)
before the artifact upload, on all three OSes.
Verified end-to-end in a real browser against a generated artifact
zip: zip load, mp4 playback, timer, tab switching, follow-sync at
t=6/t=12, click-to-seek, autoplay policy (expected NotAllowedError on
synthetic play; real clicks fine).
Also fixes the shim fail message's dead variable ( ->
observed_git_url) in both posix drivers.
The app-update legs died on the onboarding overlay - a fullscreen div
that intercepts every click (the Settings click timed out under it).
The old plan seeded a fake provider key, which lies: the overlay
vanishes but the app is broken. Instead the driver now runs the
desktop E2E suite's own mock inference server
(tests-js/scripts/mock-server.ts, zero deps, bare-node type
stripping >=22.18) and configures it into HERMES_HOME byte-for-byte
like the dev:mock flow: config.yaml provider + MOCK_API_KEY env. The
app boots genuinely configured - no overlay, real chat surface.
e2e-assets/mock-provider.{sh,mjs} own start/stop (pid + url files;
the wrapper lives until SIGTERM - gating on stdin-close made the
server die instantly, a background process's stdin is already EOF)
and the config write. Wired into the posix script driver's
hermes-desktop-app-update arm and the macos driver's update phase
(both app-update methods). The Playwright flow keeps its defense-in-
depth: the real escape hatch ('I'll choose a provider later') and the
verified 'Open settings' selector.
Probed locally: models + streamed/non-streamed completions answer,
server stops cleanly on kill.
The stage shim (git.bat, lying to the product about origin's URL) sits
on PATH, so Invoke-Git's bare 'git' routed the driver's own plumbing
through cmd - whose parser eats unquoted carets. PowerShell only quotes
args containing whitespace, so rev-parse 'v2026.8.3^{commit}' reached
the bat as v2026.8.3{commit}: bad revision.
The shim must stay a .bat: its audience is the product's python callers
(fork detection's remote get-url), which resolve via PATHEXT and never
see a .ps1. So the split is by audience - Invoke-Git pins the resolved
git.exe for every driver call; the shim serves the product, whose
shimmed flows use no caret revs (documented as the accepted hole, with
a loud bad-revision failure if that ever changes). The shim self-check
now probes through PATH, since Invoke-Git deliberately bypasses it.
macos gains the desktop-installer@latest install method: the website's
Hermes-Setup.dmg (verified live), mounted with hdiutil and its app
binary run DIRECTLY - an open-launched app inherits none of the git
redirect env, so direct exec is what keeps the isolation honest while
staying the same binary and first-launch flow.
install-e2e-macos-run.yml takes the windows shape: one workflow, one
inner job per driver arm, native skips. Arm 1 delegates script installs
to the shared OS-agnostic run workflow; arm 2 stages, installs from the
dmg, and drives both app-update methods through launch-from-spec.mjs -
open-app-update launches the installed .app (the double-click surface,
env via Playwright), hermes-desktop-app-update captures the product's
own hermes desktop spawn. Both end on sha asserts, never version
strings.
windows-desktop-gui-e2e.ps1 and windows-installer-script-e2e.ps1 fold
into tests/install/windows-e2e.ps1 with orthogonal -InstallMethod and
-Route axes: the install phase dispatches on one, the update phase on
the other, and shared workroot state carries how OLD landed - so any
implemented update method can follow any implemented install method.
Implementing a new pair is now a driver function plus a gate edit,
never a new job.
The run workflow collapses to ONE inner job whose if: is the
implemented-pairs table. Newly cheap pairs go live with the merge:
desktop-installer@latest -> hermes-update / installer-script /
installer-script+desktop / hermes-desktop-app-update
installer-script(+desktop) -> hermes-desktop-app-update
installer-script+desktop -> open-app-update (the -IncludeDesktop
install registers real Start Menu / Desktop shortcuts)
Only desktop-installer@latest as an UPDATE method stays a declared
TODO. scripts/windows_e2e_harness.ps1 executes the parse/parameter/
dispatch checks under pwsh before any Windows runner spins up.
git show | grep -qF exits at grep's first match; install.sh is ~140KB
with the flag strings in the first few KB, so git show takes SIGPIPE
on its next write and the pipeline reports 141 under set -o pipefail.
The probe answered NO for flags the ref HAS - timing-dependent, green
without pipefail (every local check), red on the runner.
It hid while a probe miss just meant omitting --skip-browser; the
first probe where NO is a hard failure (--include-desktop) exposed it
on its first CI leg. Buffer git show into a variable and grep the
string: git always completes, grep judges bytes.