An install from v2026.3.12 ships `.env` with `LLM_MODEL` set, because that
release's template wrote it. The upgrade runs config migration 12 -> 13, which
clears that dead var -- and the user-state verifier reported the change as the
upgrade modifying the user's own state, failing the leg.
The verifier exists to catch an upgrade taking state away; a var the CURRENT
tree retires is not that. The retired set is parsed out of
`hermes_cli/config_migrations.py` (the `for dead in (...): save_env_value(dead,
"")` shape) rather than restated here, so it cannot drift from the tree, and an
unreadable source retires nothing -- every .env change stays fatal.
Tolerance is deliberately narrow: only keys the tree retires, only when the
upgrade EMPTIED them, and only when no key was added or removed. Clearing a
live key, or deleting a retired one outright, still fails.
state.db was already handled that way; cron/executions.db was not, so the cron
ticker writing rows mid-window would have failed the leg as a byte modification of
user state. Any *.db is now row-judged: byte churn is tolerated, and a table that
LOSES rows still fails (_rows_shrank compares the counted tables, so genuine loss
is caught, as the new test pins against cron_jobs in cron/executions.db).
tests/scripts/test_verify_user_state.py: 23/23.
The last red leg failed with "2 deleted": cron/executions.db-shm and
cron/executions.db-wal -- SQLite's own sidecars, which exist only while a
connection is open. A snapshot taken with one live records them, they disappear
when it closes, and the report read that as the upgrade deleting user state.
-wal/-shm/-journal are tolerated deletions now and reported as such, while the
databases themselves stay judged (state.db by row counts), pinned by a test that
tolerates the sidecars vanishing AND still fails when executions.db goes.
tests/scripts/test_verify_user_state.py: 22/22.
desktop->desktop reported "0 deleted, 1 modified" with MODIFIED .env and NO
`variables ...` line: the per-key digests deliberately ignore comments, blank lines
and ordering, so a rewrite in exactly those parts is invisible to them. The snapshot
now records the file's line/comment/blank counts and its key order, and the report
prints that comparison when no key differs -- the next failure names the shape that
changed instead of leaving a bare "1 modified". Counts and names only, never
content, pinned by a new case (tests/scripts/test_verify_user_state.py 21/21).
cron/ticker_heartbeat and cron/ticker_last_success came back as fatal
modifications: the ticker writes them on its own schedule, inside the verified
window or not. On a cold home they land as ordinary additions (tolerated); once
the ticker exists, the same file moves -- and the harness's own background process
failed the user's upgrade outward. Both names are now tolerated under any cron/
directory, while the user's own cron definitions (cron/jobs.json) still fail if
they move, which the new test pins.
An app-driven upgrade rewrote $HERMES_HOME/.env to a different sha256 at an
IDENTICAL byte count, so the leg reported "0 deleted, 1 modified" and two hashes
of a secrets file with no lead. The snapshot now records per-key VALUE digests
for .env-shaped files and the report names the variables added/removed/changed --
names and equality only, never content, which is also asserted by the new test.
The verifier's judged walk pruned only plugins/**, so a bundled skill under
profiles/<name>/skills/** was judged while the identical tree at the home root
was advisory. Every leg that owns a second profile therefore failed
USER-STATE PRESERVATION on the update's own skills sync (21 modified entries,
all "advisory modified (skills sync)" in the same report).
_walk_root now carries the per-ENTRY classification instead of the per-root
flag, consulting _is_bundled_skill at both roots -- the profiles branch of that
predicate was unreachable before. skills/.archive/** stays judged at both
roots, since the curator's archive holds restorable user skills.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
The demo was scaffolding, not repo content: nothing referenced it but a
docstring.
The tests were 12 against the repo's 1-2 invariant bar. Keep the two that fail
silently when inverted -- the sentinel reaching an exec'd child (the unexported
variable this change fixes) and the prologue's staleness gate, which inverted
either way costs a re-sync per run or a stale environment that looks fine. The
per-OS command pair stays because neither host can observe the other's string.
Dropped the exit-code/message restatement, the no-sentinel cold path, and the
demo-driven harness.
`scripts/install.sh::discard_update_lockfile_churn` and `scripts/install.ps1::Discard-LockfileChurn`
run the same per-directory predicate as `hermes update` did before the previous commit, so an
installer-driven update of a managed checkout (Desktop / bootstrap) reverted the root
`package-lock.json` whenever only `apps/desktop/package.json` was dirty, leaving spec and lock
out of sync for the next `npm ci`. Port the same ownership model: the root lock is kept when the
root manifest or any manifest matching a root `workspaces` glob is dirty; nested lockfiles are
still kept only with their sibling manifest; a manifest outside the graph still does not
protect the root lock.
install.sh reads the globs with sed/grep (no jq dependency) and matches with `case`; install.ps1
uses ConvertFrom-Json and `-like`. Bash side live-A/B'd in a throwaway repo (red on main, green
after; controls unchanged); the PowerShell side is the same shape and could not be executed on
this Linux host (no pwsh).
Repo scripts assume the PM-activated environment, so running one without
activation fails much later with a confusing ImportError. Add the two halves
covering both invocation paths:
- scripts/_activation.py: require_activation() exits immediately, naming the
exact command for the caller's shell (source ./activate on POSIX,
. .\activate.ps1 on a native Windows host), before any heavy import.
- scripts/_hermes-python: the POSIX shebang target. `#!/usr/bin/env -S bash -c
'...'` hands itself the target path through bash -c's $0, sources activate,
then execs the interpreter on the same file -- so tracebacks and __file__
still point at the real script and ./scripts/foo.py works from any cwd with
no manual source.
__HERMES_ACTIVATED changes from a bare "1" to the installed-state file the
environment was composed against, so one value carries activation, which
checkout activated it, and a staleness stamp. The prologue compares that file
against uv.lock / pyproject.toml / pm/lock.json with the `-nt` builtin -- no
process spawn -- and re-activates once when the inherited environment predates
its inputs. pm rewrites that file only on a real sync, so the check settles
back to current rather than re-syncing on every run.
A legacy "1" keeps working: require_activation() tests non-emptiness, and the
prologue's [ -e ] fails on it, so it activates once and upgrades.
The install/update legs asserted plenty about the code -- the checkout
landed, the version bumped, the desktop artifact exists -- and nothing
about the user's own state. An upgrade that ate auth.json or truncated
state.db would have passed every leg.
Adds a read-only, stdlib-only verifier (snapshot/verify) plus the hooks
that drive it around the real upgrade, on the POSIX and Windows drivers.
The state it defends is produced through the ordinary CLI
(hermes chat -q / auth add / profile create), never seeded by the
harness, and each action asserts it actually landed so a leg cannot
'pass' while testing nothing.
Judged: config.yaml, .env, auth.json, state.db, gateway_state.json and
the user's trees. state.db is compared by row counts, not bytes -- a
live SQLite file moves for benign reasons. The bundled skills/ tree is
recorded but never judged (the product re-syncs it), and plugins/** is
left to verify-plugin-preservation.py.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
The termux-main pool deletes a package's previous archive when it rebuilds, so
the runtime-lib pin table and the bionic lock rows rot without warning. The last
rotation broke a build on eight rows at once, and the stager's concurrent
downloads only surfaced whichever 404 won the race.
pm now owns the pin table it repairs: scripts/termux/runtime_libs.json moves to
pm/termux_runtime_libs.json, so pins live in pm/ and scripts consume them — the
direction scripts/ci/archive_inputs.py already reads pm/lock.json in.
`hermes pm update --termux` repins exactly the rows whose archive the pool has
replaced, hashing each replacement against the index SHA256 before writing
url/version/hash together. `--check` reports without writing and exits 1, so a
retired pin can fail a cheap preflight instead of a payload build.
It is a repair, not an update: an alive pin is never moved, because a repin can
land a rebuilt library under a moved soname and the table is the payload's
recursive DT_NEEDED closure. A pin whose package the pool has dropped outright
is reported and left alone. `--termux` runs alone — names/--target/--uv/--npm
are ignored, since repairing foreign-target pins is not a version resolution.
Verified: `pm update --termux --check` against the live pool reports 89 rows
served; a table deliberately pinned to the retired libiconv 1.18-1 repins to
1.19 with the pool's hash through the real network path; 19 new tests; the
tests/pm, tests/ci and tests/scripts suites have the same failure set as the
base commit (91 pre-existing Windows environment failures, none new).
The local --channel command no longer touches R2: it resolves the exact
pushed commit and dispatches the default-branch workflow, whose privileged
allocation step creates the channel and mints the immutable build request.
The disposable allocation path is generalized to cover the unscoped
production preview, gated by a new `channel` workflow input; build legs
consume the same channel-build / channel-request-sha256 job outputs as
before. The anti-tamper gate is now commit_build.admit (maintainer
permission) running inside the allocate step.
Drop the resume path: --resume-channel-build / --request-sha256 and
resume_build() are gone, and allocate_protected no longer recovers a lost
request PUT by sequence. Retrying re-dispatches and mints a fresh sequence
slot; idempotency survives via the deterministic build ID and the existing
immutable-request dedup. Keep the preview allocate `lastAllocation` build-ID
field (concurrent-CAS uniqueness), which is not resume.
channel_public_base now defaults to the documented production origin like
the commit-build path, so a local command names its page without a
hand-set CLOUDFLARE_R2_PUBLIC_URL.
Also drops a stale fork-isolation assertion left behind by the
fork-conditional dispatch removal.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.
/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
The CI desktop-build cache restores by prefix after dependency changes, so
it accumulates wheels for superseded pins; measured 210MB of sediment on a
real build host. stage_uv_cache copied that snapshot wholesale — stale
wheels shipped in every bundle and grew monotonically.
prune_uv_cache_to_lock now deletes archive buckets and wheel/sdist index
entries the shipped uv.lock cannot resolve, after staging. The lock is the
ship contract: CI transport may roll and accumulate; the payload never
ships sediment. Unidentifiable buckets (no dist-info) survive — pruning
fails open for unknown layouts, never for identifiable stale pins.
Repository identity no longer selects behavior: a commit build is the
same direct dispatch from any repository, and R2 disposable scoping is
opt-in via R2_DISPOSABLE_RUN rather than fork-mandated. The fork guard
in the workflow admission step, the fork refusal in R2Scope.configured,
and the client-side fork routing (disposable_dispatch_command) all go.
Disposable namespaces keep their own protections: malformed leases are
refused, and a scoped namespace still belongs to exactly one repository.
Un-slim the bundle's uv cache: prune_uv_cache_to_built (built-only keep-set)
is deleted — every wheel the build resolved ships, so any per-install venv
rebuild (plugin extras, feature changes) resolves offline with zero network.
stage_uv_cache keeps excluding sdist src/ trees (build-only bulk) but now
KEEPS built wheel ZIPs including the copies beside sdist metadata.msgpack
shards: probing showed an offline rebuild whose revision entry lacks its
wheel ZIP fails closed wanting the sdist download. Fixture sdist uses a
stdlib-only PEP 517 backend so the hermetic offline rebuild needs no build
tools.
First launch of a bundled payload paid a cold-compile stall: the launcher
redirects bytecode writes to a user-level cache (signature-breaking on
macOS, read-only mount on AppImage/MSIX), so every import compiled from
source. Now staging bakes the cache into the payload:
- compileall with the payload's OWN staged 3.14 interpreter, unchecked-
hash pycs: repack mtimes cannot invalidate them, a stale source can
never trigger a rewrite, and read-only pycs mean the macOS signature
never observes a change. Dirs stay writable — in-place rebuilds
rmtree the tree; asserted coverage plus unchecked-hash means no
cache-miss write can target them.
- coverage is the perf contract: the bake FAILS if any parseable module
lacks a pyc (empirically 0 unparseable files ship, so compileall is
strict). Probe suite: py_compile/cache_from_source, PEP 552 flags,
multi-root read, stale-source no-rewrite, read-only cache-dir import.
- launcher: the baked marker makes configure() leave sys.pycache_prefix
UNSET — the prefix relocates reads too and would hide the baked pycs.
Payload modules read their source-adjacent cache (Python's default
multi-root lookup); plugin/user modules keep caching beside their own
sources under HERMES_HOME. Unmarked payloads keep the old redirect.
- snapshot(): the sealed payload ships without tests/website/evals/
.github/nix/docker/tests-js (~69MB, 46% of tracked bytes) and without
apps/ui-tui/web/scripts — CI prebuilds those products, and
is_bundled_payload routes sealed updates to the channel updater, so
the rebuild graph never runs in a bundle (linux_desktop_entry degrades
to the themed icon). Frontend product staging keeps the full tree.
- test_bundle_native now stages the FULL relocatable toolchain (a bare
interpreter ELF falls back to its compile-time /install prefix and
cannot create a venv), and runs on the real 3.14 for the first time
this campaign — the whole battery had been running 3.12 against the
3.14-pinned lock.
A fork commit build is now ONE workflow dispatch whose run both allocates
the disposable channel and builds it. Previously allocation printed a
follow-up command that had to be dispatched separately (the GITHUB_TOKEN
recursion wall forced two runs; release.py grew poll/extract machinery to
automate the hop — all deleted now, net -144 lines).
- Lease is the run id alone (no attempt suffix): re-run failed jobs
re-enters the same namespace; succeeded allocate job is skipped and its
outputs persist. r2_scope accepts legacy <id>-<attempt> leases on read.
- allocate-disposable emits job outputs (channel_build, request digest,
lease, public base); validate consumes them in disposable mode and
re-exports a normalized pin every downstream job reads.
- Admission guards unchanged in semantics: forks still cannot run without
a disposable allocation; disposable runs still never touch production
feeds, termux, or the commit-builds page.
- Upstream trusted-controller path (channel_build inputs) byte-identical.
- Fixtures updated for run-id leases; two dead two-dispatch tests and the
fixture's allocation-probe plumbing removed; smoke matrix test skips
banana's new admission job (no toolchain by design).
Fork CI now requires a disposable channel allocation (R2_DISPOSABLE_RUN
guard in desktop-bundled-release.yml), but 'release.py --build-commit REV
--publish' still fired the old direct dispatch, so every fork commit build
died at admission. cmd_build_commit now detects a non-upstream repository
(case-insensitive NousResearch/hermes-agent compare), dispatches the
allocation workflow with disposable_channel/build_commit/bundle_env baked
in (all --bundle-env/--bundle-unset values travel inside the immutable
request; the follow-up never re-passes them), polls the allocation run to
completion (15s interval, 15min budget), extracts the printed follow-up
dispatch from the run logs (channel_disposable's single-line JSON
'command'), validates its shape (gh workflow run of this workflow against
the same repository), and auto-dispatches it with the local maintainer's
gh login — falling back to a clear run-summary pointer when log recovery
fails. Upstream behavior is unchanged. Also fixes a pre-existing TypeError
that masked check_output failures whose CalledProcessError has stderr=None.
The desktop payloads shipped the builder's full uv cache (2.0GB mac,
1.1GB win-x64, 800MB win-arm64) so a mutable-venv rebuild could run
offline. But the bundle contract never needed offline rebuilds — it
needs rebuilds that never invoke a compiler: packages with no
downloadable wheel (sdist-only or platform gaps) would otherwise
demand Xcode CLT/MSVC/Rust on the user's machine.
prune_uv_cache_to_built keeps only what the builder compiled itself,
detected from the cache's own records (sdists-v9 entries, cached
wheels not listed in uv.lock, archive-v0 buckets by dist-info) and
drops the ~245 packages whose wheels PyPI re-serves in seconds.
Validated live on a clean host per target:
- macos arm64 (iris, no uv): 2.0GB -> 3MB; sync online, 0 builds,
pilk/psutil/alibabacloud-tea installed from the slim
- win11 arm64 (promise, no uv/MSVC): 800MB -> 48MB; sync online,
0 dependency builds, all 12 CI-built packages (cryptography,
httptools, brotlicffi, dependency-injector, ruamel-yaml-clib,
obstore, davey, firecrawl-anydoc, pilk, psutil, alibabacloud set)
import clean. The kept set matched the CI build log's Built list
exactly (12/12, no false positives).
Index-dir detection is name-agnostic (pypi/ vs custom --index-url
hashed dir), proven by the rewritten bundle test: sdist-only survives
the slim, downloadable wheels are dropped, the build machine's cache
is untouched, and an online rebuild compiles nothing.
The e2e-screen-record action installed ffmpeg through three different
OS package managers (apt, brew, winget). winget is the flaky leg on
windows-11-arm and serves an x64-gyan build that runs emulated on ARM;
choco's community package wraps the same gyan x64 zip, so a choco
fallback would not fix either problem. PM already ships a locked,
sha256-pinned ffmpeg (martin-riedl posix, BtbN win32 including a NATIVE
winarm64 build), so the recording action now verifies ffmpeg on PATH
instead of installing it, and the pinned binary rides the same
tools-cache as node/python.
- setup-pm learns a `packages` input (extra PM tools beyond the
toolchain roots, e.g. ffmpeg) threaded through setup_toolchain.py's
prepare/install/archive-inputs phases; the tools-cache key gains an
extra-packages fragment so existing keys stay byte-identical.
- The six chat-driver jobs pass `packages: ffmpeg` to setup-pm.
- e2e-screen-record drops the apt/brew/winget install steps, the
ffmpeg actions/cache steps and the save-cache input; Xvfb (headless
linux) and the macOS replayd-approval hack stay.
- The "Install locked chat driver dependencies" step installs only the
tests-js workspace with --omit=dev instead of the whole apps/desktop
tree: the drivers need @playwright/test, zod (previously a phantom
hoisted from @assistant-ui/react), js-yaml and semver only. 64 pure
packages in ~2s vs ~1800 including electron-builder and native
builds; the tree-shaken node_modules runs the real driver modules
(verified by importing desktop-chat-smoke.ts and update-window-chat
end to end).
- tests pin the new contracts; tests/install/README.md updated.
Share the real composer, provider-witness and completed-reply check across
post-build bundle smoke and desktop-bearing install/update checkpoints.
Keep native automatic-relaunch proof separate from post-update chat.
Download receipt-bound artifacts without release credentials and install
DMG, ZIP, MSIX and universal MSIXBUNDLE on each native architecture.
Split Windows assembly from feed publication; publish tested bytes only.
Bind candidate smoke results into the manifest used by stable promotion.
Verify historical/source provenance without assuming a version IPC commit,
strip CI identity from source build children, and use the actual Electron
PID rather than Playwright's Windows launcher wrapper.
Validation: real Linux Electron chat and sequential OLD/NEW source smoke
with preserved history; 145 targeted Python tests and 14 JS tests passed;
TypeScript, shell/PowerShell parsing and workflow checks passed.
Native macOS/Windows deployment and historical upgrades need Actions proof.