The installer ladder stopped at node-deps/path/desktop with its own
semantics while an update ran launchers, product builds and post-build
maintenance, so a fresh install and a finished update ended in different
states: after re-running the installer at HEAD the products had no receipts
and the read-only source acceptance failed.
hermes_cli/source_completion.py now owns that tail -- publish launchers,
build the products, run the maintenance -- and update_completion's
_complete_selected calls it, so there is one implementation. install.sh and
install.ps1 keep the bootstrap stages (prerequisites, repository, venv,
python-deps, config) and hand off to it in a single `products` stage;
--include-desktop selects the desktop product inside that stage instead of
adding a second build stage, and `desktop` stays dispatchable via --stage for
external callers.
Windows keeps its installer-owned PATH publication (expose_cli answers
"windows-installer-owned" on Windows) plus the packaged-artifact probe, ACL
grant and shortcuts. The desktop stage no longer pre-syncs wake/voice: pm
lazy-installs them at first use, as the update path does.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
The demo was scaffolding, not repo content: nothing referenced it but a
docstring.
The tests were 12 against the repo's 1-2 invariant bar. Keep the two that fail
silently when inverted -- the sentinel reaching an exec'd child (the unexported
variable this change fixes) and the prologue's staleness gate, which inverted
either way costs a re-sync per run or a stale environment that looks fine. The
per-OS command pair stays because neither host can observe the other's string.
Dropped the exit-code/message restatement, the no-sentinel cold path, and the
demo-driven harness.
Path.glob raises NotImplementedError for a non-relative pattern, which a string `workspaces` (iterated char by char, so "/") or an absolute entry produces. The (OSError, ValueError, TypeError) catch missed it, so the error escaped to the caller's suppress(Exception) and no lock was reverted at all -- back to autostash every run. Non-list values are now ignored and each pattern is tried on its own so a bad one just owns nothing.
install.sh: read the workspace globs with `while read` instead of an unquoted $(...) so they are never pathname-expanded against the caller's CWD before `case` sees the pattern.
`scripts/install.sh::discard_update_lockfile_churn` and `scripts/install.ps1::Discard-LockfileChurn`
run the same per-directory predicate as `hermes update` did before the previous commit, so an
installer-driven update of a managed checkout (Desktop / bootstrap) reverted the root
`package-lock.json` whenever only `apps/desktop/package.json` was dirty, leaving spec and lock
out of sync for the next `npm ci`. Port the same ownership model: the root lock is kept when the
root manifest or any manifest matching a root `workspaces` glob is dirty; nested lockfiles are
still kept only with their sibling manifest; a manifest outside the graph still does not
protect the root lock.
install.sh reads the globs with sed/grep (no jq dependency) and matches with `case`; install.ps1
uses ConvertFrom-Json and `-like`. Bash side live-A/B'd in a throwaway repo (red on main, green
after; controls unchanged); the PowerShell side is the same shape and could not be executed on
this Linux host (no pwsh).
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:
- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
`git status --ignored=matching`, so a gitignored `.venv/` surfaced as
`!! .venv/` and refused every ZIP fallback on such installs ("the working
tree has uncommitted changes or untracked files") — the live runtime was
being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
place rather than growing a second environment that `project_venv_dir()`
then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
before this commit); the holder-guard test now uses a kernel-runner child
whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
it (red on origin/main); drop the `process.platform`-override vitest case,
which exercised the same resolver as the `.venv` case with a different
directory string.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
The committed lockfile still resolved body-parser 1.20.6 (nested qs 6.15.3),
express 4.22.2 with a top-level qs 6.15.3 and sharp 0.35.3, so the override
bump alone left `npm audit` at 3 findings (2 moderate qs, 1 high sharp) for
anyone installing from the lock — and `hermes doctor` kept flagging the
"WhatsApp bridge deps" row. `npm update --package-lock-only` inside the
existing manifest ranges: body-parser 1.20.8, express 4.22.3, qs 6.16.0,
sharp 0.35.4 -> `npm audit`: found 0 vulnerabilities. No manifest change
beyond the override bump; Baileys stays pinned at 7.0.0-rc13.
#109060 (Sep 12) was the earliest PR to move the override to 1.20.8 (it also
carried a redundant qs override, which 1.20.8 makes unnecessary).
Part of #112382
Co-authored-by: BenKalsky <1568840+BenKalsky@users.noreply.github.com>
The 1.20.6 pin sat inside the vulnerable range it was meant to clear (1.20.5 - 1.20.6); 1.20.8 pulls qs ~6.16.0, clearing the transitive qs advisories.
Repo scripts assume the PM-activated environment, so running one without
activation fails much later with a confusing ImportError. Add the two halves
covering both invocation paths:
- scripts/_activation.py: require_activation() exits immediately, naming the
exact command for the caller's shell (source ./activate on POSIX,
. .\activate.ps1 on a native Windows host), before any heavy import.
- scripts/_hermes-python: the POSIX shebang target. `#!/usr/bin/env -S bash -c
'...'` hands itself the target path through bash -c's $0, sources activate,
then execs the interpreter on the same file -- so tracebacks and __file__
still point at the real script and ./scripts/foo.py works from any cwd with
no manual source.
__HERMES_ACTIVATED changes from a bare "1" to the installed-state file the
environment was composed against, so one value carries activation, which
checkout activated it, and a staleness stamp. The prologue compares that file
against uv.lock / pyproject.toml / pm/lock.json with the `-nt` builtin -- no
process spawn -- and re-activates once when the inherited environment predates
its inputs. pm rewrites that file only on a real sync, so the check settles
back to current rather than re-syncing on every run.
A legacy "1" keeps working: require_activation() tests non-emptiness, and the
prologue's [ -e ] fails on it, so it activates once and upgrades.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
The canary tag regex capped the major at 3 digits while the stable
shape had no cap. The repo's current stable line is CalVer (v2026.9.14),
so canary_tag_for_date cuts v2026.9.15-canary.<ts> — which
handoff.validate_identity then rejected as 'Invalid release handoff
identity', killing every tag-mode stage leg (Windows and Darwin) while
the same-minor stable staged fine.
Lift the cap in the canonical _CANARY_TAG_RE and the Termux mirror
regex; add the stable-vs-canary shape-parity invariant, proven red on
the base regexes.
The termux-main pool deletes a package's previous archive when it rebuilds, so
the runtime-lib pin table and the bionic lock rows rot without warning. The last
rotation broke a build on eight rows at once, and the stager's concurrent
downloads only surfaced whichever 404 won the race.
pm now owns the pin table it repairs: scripts/termux/runtime_libs.json moves to
pm/termux_runtime_libs.json, so pins live in pm/ and scripts consume them — the
direction scripts/ci/archive_inputs.py already reads pm/lock.json in.
`hermes pm update --termux` repins exactly the rows whose archive the pool has
replaced, hashing each replacement against the index SHA256 before writing
url/version/hash together. `--check` reports without writing and exits 1, so a
retired pin can fail a cheap preflight instead of a payload build.
It is a repair, not an update: an alive pin is never moved, because a repin can
land a rebuilt library under a moved soname and the table is the payload's
recursive DT_NEEDED closure. A pin whose package the pool has dropped outright
is reported and left alone. `--termux` runs alone — names/--target/--uv/--npm
are ignored, since repairing foreign-target pins is not a version resolution.
Verified: `pm update --termux --check` against the live pool reports 89 rows
served; a table deliberately pinned to the retired libiconv 1.18-1 repins to
1.19 with the pool's hash through the real network path; 19 new tests; the
tests/pm, tests/ci and tests/scripts suites have the same failure set as the
base commit (91 pre-existing Windows environment failures, none new).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
The local --channel command no longer touches R2: it resolves the exact
pushed commit and dispatches the default-branch workflow, whose privileged
allocation step creates the channel and mints the immutable build request.
The disposable allocation path is generalized to cover the unscoped
production preview, gated by a new `channel` workflow input; build legs
consume the same channel-build / channel-request-sha256 job outputs as
before. The anti-tamper gate is now commit_build.admit (maintainer
permission) running inside the allocate step.
Drop the resume path: --resume-channel-build / --request-sha256 and
resume_build() are gone, and allocate_protected no longer recovers a lost
request PUT by sequence. Retrying re-dispatches and mints a fresh sequence
slot; idempotency survives via the deterministic build ID and the existing
immutable-request dedup. Keep the preview allocate `lastAllocation` build-ID
field (concurrent-CAS uniqueness), which is not resume.
channel_public_base now defaults to the documented production origin like
the commit-build path, so a local command names its page without a
hand-set CLOUDFLARE_R2_PUBLIC_URL.
Also drops a stale fork-isolation assertion left behind by the
fork-conditional dispatch removal.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
d8286ae58f made repository identity irrelevant to dispatch and made R2
disposable scoping opt-in via R2_DISPOSABLE_RUN, but the stable-releases
doc and the --build-commit help still described forks as requiring a
disposable allocation. Describe the single direct dispatch and the opt-in
scoped namespace instead.
The CI cache barely ever restored anything useful, for four independent
reasons found in live run logs and the fork's cache store:
- `uv cache prune --ci` (setup-pm/prune) discards downloaded wheels before
the save, so the snapshot carried only source-built wheels — the next
run "restored" it and still cold-downloaded everything. Upstream main's
logs showed the end state: a 209KB stub cache exact-hitting forever.
- Tool-only jobs (icons-freshness etc.) auto-saved a 1.5KB empty uv cache
under the production exact key; caches are immutable, so the stub won
forever and blocked real saves.
- The npm cache key had no restore-keys, so one lockfile bump missed the
exact key and every npm job went cold.
- `2efa4ff94f` deleted the electron-builder toolchain save but left the
assemble job restoring `eb2-` — a fossil nobody regenerates; assembly
cold-downloads winCodeSign/ATS/dotnet on every release.
Changes:
- pm/cache_lock.py: move prune_uv_cache_to_lock out of
scripts/bundles/native.py (which re-exports it); the lock-exactness
contract now serves both the bundle ship gate and CI caches.
- pm.build_env learns --exact-lock --lock-source: prune cache entries the
project uv.lock cannot resolve, keeping lock-required downloaded wheels
(unlike --ci). Refuses --ci/--prune-cache combinations.
- setup-pm: python-cache auto-save now requires extras (no stubs from
tool-only jobs); key drops the prune flag and bumps to v3 — pruned and
unpruned saves share one namespace since both are lock-exact; pre-save
pruning switched from --ci to --exact-lock; npm cache gains a
lockfile-agnostic restore prefix.
- save-pm-cache: same exact-lock prune before explicit saves.
- desktop-bundled-release: build legs (cache-mode: write) restore+save the
electron-builder toolchain under eb3- keyed on the locked builder
version; assemble restores the same namespace; the dead default-cache
resolution step is removed (assembly resolves no electron artifacts).
- cleanup_pm_toolchain_caches.py: match v2 and v3 smoke keys.
Validation: tests/scripts/test_bundle_native.py 11/11 (incl. both
lock-prune gates), tests/pm failures identical before/after the diff,
tests/scripts/test_bundle_payload.py 5/5, tests/ci cleanup 3/3,
tests-js setup-pm-post 1/1 and the three setup-pm-cache contract tests
updated and green (6 failures in that file pre-date this diff and are
drift between the workflow and its stale assertions); pm.build_env
--exact-lock E2E against a real uv cache copy pruned 3 stale entries and
kept the rest.
pillow-heif's sdist build failed with RequiredDependencyException: the
container toolchain installs pillow's libs but not libheif, so
pkg-config libheif had nothing to answer with. Termux's libheif 1.23.4
ships the headers and libheif.pc directly (no -dev split).
Runtime: the built _pillow_heif extension DT_NEEDEDs libheif.so, which
pulls libx265/libde265/libaom/libx264/libc++_shared. All but libheif
and libde265 were already staged in runtime_libs.json (verified by
reading the .so's DT_NEEDED, not the deb Depends line, which also
lists gdk-pixbuf/glib/rav1e the linker never loads); pin those two.
Import gate: import pillow_heif never fails on a dead link -- its
__init__ swallows the _pillow_heif ImportError into a DeferredError
that fires only on first use. The gate now imports _pillow_heif
directly so a dlopen failure fails the build instead of the phone.
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#89322 fixed the bridge-local normalizeWhatsAppId, but bridge.js has since
moved its id handling to bridge_helpers.js::normalizeWhatsAppId, which still
turned `<user>:<device>@lid` into the malformed `<user>@<device>@lid` for
mentionedJid / quoted participant / reaction keys, and the Python side
(gateway/platforms/whatsapp_common.py::_normalize_whatsapp_id) did the same
':'->'@' swap on botIds. Drop the local duplicate in bridge.js, import the
helper, and strip the `:<device>` suffix on both layers so the bot's own ids
compare equal to the bare ids WhatsApp sends for mentions and quotes.
One invariant test: device-qualified botIds match a bare mentionedId and a
bare quotedParticipant; a plain group message still does not trigger.
normalizeWhatsAppId did String(value).replace(':','@'), which turns a device-qualified id
like '116342762025117:14@lid' into the malformed '116342762025117@14@lid'. The bot's own id
(sock.user.id / sock.user.lid) carries the :<device> suffix while inbound mentionedJid and
contextInfo.participant (quoted message author) do not, so the bot's id never matches its
botIds set -> @mention and reply-to-bot are never detected in groups. Strip the :<device>
suffix instead so all id forms compare consistently.
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.
_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.
Includes a sabotage-verified unit test (fails without the scaler).
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.
Fix the class, not the sites:
* hermes_state: register every successfully constructed SessionDB in a
test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
everything left in the registry after each test. close() is idempotent
and unregisters the pinning atexit hook, so instances become
collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
leak fails fast with MemoryError instead of eating the box.
Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
for registration, idempotent close, and the cross-test sweep.
Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.
Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.
/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
The CI desktop-build cache restores by prefix after dependency changes, so
it accumulates wheels for superseded pins; measured 210MB of sediment on a
real build host. stage_uv_cache copied that snapshot wholesale — stale
wheels shipped in every bundle and grew monotonically.
prune_uv_cache_to_lock now deletes archive buckets and wheel/sdist index
entries the shipped uv.lock cannot resolve, after staging. The lock is the
ship contract: CI transport may roll and accumulate; the payload never
ships sediment. Unidentifiable buckets (no dist-info) survive — pruning
fails open for unknown layouts, never for identifiable stale pins.
Repository identity no longer selects behavior: a commit build is the
same direct dispatch from any repository, and R2 disposable scoping is
opt-in via R2_DISPOSABLE_RUN rather than fork-mandated. The fork guard
in the workflow admission step, the fork refusal in R2Scope.configured,
and the client-side fork routing (disposable_dispatch_command) all go.
Disposable namespaces keep their own protections: malformed leases are
refused, and a scoped namespace still belongs to exactly one repository.
Un-slim the bundle's uv cache: prune_uv_cache_to_built (built-only keep-set)
is deleted — every wheel the build resolved ships, so any per-install venv
rebuild (plugin extras, feature changes) resolves offline with zero network.
stage_uv_cache keeps excluding sdist src/ trees (build-only bulk) but now
KEEPS built wheel ZIPs including the copies beside sdist metadata.msgpack
shards: probing showed an offline rebuild whose revision entry lacks its
wheel ZIP fails closed wanting the sdist download. Fixture sdist uses a
stdlib-only PEP 517 backend so the hermetic offline rebuild needs no build
tools.
One-dispatch runs export HERMES_BUILD_COMMIT to every build leg, and
handoff's argparse defaults pick it up even when the leg names only
--channel-request — so the macOS stage refused the allocation's own
commit as 'another commit'. Match the admit() equality rule: only a
commit different from request[commit] (or any tag) is an override.
The override guard rejected BUILD_COMMIT and BUNDLE_ENV_JSON outright —
correct for the old two-dispatch flow, where the build run carried neither
input. In the one-dispatch flow the validation job inherits the allocation
dispatch's own inputs, so admission refused the exact values the allocation
had pinned and every disposable build died at "Channel requests cannot
select one-off, release, Termux or bundle overrides".
The request is already digest-pinned when admit() sees it, so the
invariant becomes equality: BUILD_COMMIT must match request["commit"] and
BUNDLE_ENV_JSON must parse to exactly request["bundleEnv"]. Any other
value still refuses, with the clearer message that it differs from the
dispatch that allocated this build. Release/tag/Termux overrides keep the
absolute refusal.
First launch of a bundled payload paid a cold-compile stall: the launcher
redirects bytecode writes to a user-level cache (signature-breaking on
macOS, read-only mount on AppImage/MSIX), so every import compiled from
source. Now staging bakes the cache into the payload:
- compileall with the payload's OWN staged 3.14 interpreter, unchecked-
hash pycs: repack mtimes cannot invalidate them, a stale source can
never trigger a rewrite, and read-only pycs mean the macOS signature
never observes a change. Dirs stay writable — in-place rebuilds
rmtree the tree; asserted coverage plus unchecked-hash means no
cache-miss write can target them.
- coverage is the perf contract: the bake FAILS if any parseable module
lacks a pyc (empirically 0 unparseable files ship, so compileall is
strict). Probe suite: py_compile/cache_from_source, PEP 552 flags,
multi-root read, stale-source no-rewrite, read-only cache-dir import.
- launcher: the baked marker makes configure() leave sys.pycache_prefix
UNSET — the prefix relocates reads too and would hide the baked pycs.
Payload modules read their source-adjacent cache (Python's default
multi-root lookup); plugin/user modules keep caching beside their own
sources under HERMES_HOME. Unmarked payloads keep the old redirect.
- snapshot(): the sealed payload ships without tests/website/evals/
.github/nix/docker/tests-js (~69MB, 46% of tracked bytes) and without
apps/ui-tui/web/scripts — CI prebuilds those products, and
is_bundled_payload routes sealed updates to the channel updater, so
the rebuild graph never runs in a bundle (linux_desktop_entry degrades
to the themed icon). Frontend product staging keeps the full tree.
- test_bundle_native now stages the FULL relocatable toolchain (a bare
interpreter ELF falls back to its compile-time /install prefix and
cannot create a venv), and runs on the real 3.14 for the first time
this campaign — the whole battery had been running 3.12 against the
3.14-pinned lock.
A fork commit build is now ONE workflow dispatch whose run both allocates
the disposable channel and builds it. Previously allocation printed a
follow-up command that had to be dispatched separately (the GITHUB_TOKEN
recursion wall forced two runs; release.py grew poll/extract machinery to
automate the hop — all deleted now, net -144 lines).
- Lease is the run id alone (no attempt suffix): re-run failed jobs
re-enters the same namespace; succeeded allocate job is skipped and its
outputs persist. r2_scope accepts legacy <id>-<attempt> leases on read.
- allocate-disposable emits job outputs (channel_build, request digest,
lease, public base); validate consumes them in disposable mode and
re-exports a normalized pin every downstream job reads.
- Admission guards unchanged in semantics: forks still cannot run without
a disposable allocation; disposable runs still never touch production
feeds, termux, or the commit-builds page.
- Upstream trusted-controller path (channel_build inputs) byte-identical.
- Fixtures updated for run-id leases; two dead two-dispatch tests and the
fixture's allocation-probe plumbing removed; smoke matrix test skips
banana's new admission job (no toolchain by design).
Fork CI now requires a disposable channel allocation (R2_DISPOSABLE_RUN
guard in desktop-bundled-release.yml), but 'release.py --build-commit REV
--publish' still fired the old direct dispatch, so every fork commit build
died at admission. cmd_build_commit now detects a non-upstream repository
(case-insensitive NousResearch/hermes-agent compare), dispatches the
allocation workflow with disposable_channel/build_commit/bundle_env baked
in (all --bundle-env/--bundle-unset values travel inside the immutable
request; the follow-up never re-passes them), polls the allocation run to
completion (15s interval, 15min budget), extracts the printed follow-up
dispatch from the run logs (channel_disposable's single-line JSON
'command'), validates its shape (gh workflow run of this workflow against
the same repository), and auto-dispatches it with the local maintainer's
gh login — falling back to a clear run-summary pointer when log recovery
fails. Upstream behavior is unchanged. Also fixes a pre-existing TypeError
that masked check_output failures whose CalledProcessError has stderr=None.
The desktop payloads shipped the builder's full uv cache (2.0GB mac,
1.1GB win-x64, 800MB win-arm64) so a mutable-venv rebuild could run
offline. But the bundle contract never needed offline rebuilds — it
needs rebuilds that never invoke a compiler: packages with no
downloadable wheel (sdist-only or platform gaps) would otherwise
demand Xcode CLT/MSVC/Rust on the user's machine.
prune_uv_cache_to_built keeps only what the builder compiled itself,
detected from the cache's own records (sdists-v9 entries, cached
wheels not listed in uv.lock, archive-v0 buckets by dist-info) and
drops the ~245 packages whose wheels PyPI re-serves in seconds.
Validated live on a clean host per target:
- macos arm64 (iris, no uv): 2.0GB -> 3MB; sync online, 0 builds,
pilk/psutil/alibabacloud-tea installed from the slim
- win11 arm64 (promise, no uv/MSVC): 800MB -> 48MB; sync online,
0 dependency builds, all 12 CI-built packages (cryptography,
httptools, brotlicffi, dependency-injector, ruamel-yaml-clib,
obstore, davey, firecrawl-anydoc, pilk, psutil, alibabacloud set)
import clean. The kept set matched the CI build log's Built list
exactly (12/12, no false positives).
Index-dir detection is name-agnostic (pypi/ vs custom --index-url
hashed dir), proven by the rewritten bundle test: sdist-only survives
the slim, downloadable wheels are dropped, the build machine's cache
is untouched, and an online rebuild compiles nothing.
The e2e-screen-record action installed ffmpeg through three different
OS package managers (apt, brew, winget). winget is the flaky leg on
windows-11-arm and serves an x64-gyan build that runs emulated on ARM;
choco's community package wraps the same gyan x64 zip, so a choco
fallback would not fix either problem. PM already ships a locked,
sha256-pinned ffmpeg (martin-riedl posix, BtbN win32 including a NATIVE
winarm64 build), so the recording action now verifies ffmpeg on PATH
instead of installing it, and the pinned binary rides the same
tools-cache as node/python.
- setup-pm learns a `packages` input (extra PM tools beyond the
toolchain roots, e.g. ffmpeg) threaded through setup_toolchain.py's
prepare/install/archive-inputs phases; the tools-cache key gains an
extra-packages fragment so existing keys stay byte-identical.
- The six chat-driver jobs pass `packages: ffmpeg` to setup-pm.
- e2e-screen-record drops the apt/brew/winget install steps, the
ffmpeg actions/cache steps and the save-cache input; Xvfb (headless
linux) and the macOS replayd-approval hack stay.
- The "Install locked chat driver dependencies" step installs only the
tests-js workspace with --omit=dev instead of the whole apps/desktop
tree: the drivers need @playwright/test, zod (previously a phantom
hoisted from @assistant-ui/react), js-yaml and semver only. 64 pure
packages in ~2s vs ~1800 including electron-builder and native
builds; the tree-shaken node_modules runs the real driver modules
(verified by importing desktop-chat-smoke.ts and update-window-chat
end to end).
- tests pin the new contracts; tests/install/README.md updated.