`stage_repository` treated a failed `pull --ff-only` as "keep local state"
and carried on. Every later stage reads files only the new tree has (`pm/`),
so an install whose checkout could not fast-forward failed further down with
`ModuleNotFoundError: No module named 'pm'` instead of updating.
A release cut off the main line diverges for every user who installed from it:
v2026.5.29.2 is not an ancestor of main, so its checkout can never fast-forward
to a later main. Match the remote the way `hermes update` already does, after
parking the old tip behind refs/hermes-install-backup/<stamp>-<sha> and
stashing any local work so neither is destroyed silently.
Verified against a real diverged repository: before, the checkout stayed on
the old line (HEAD=6a2e091, remote=43eea63) with no backup; after, HEAD equals
the new main, the backup ref and an autostash exist, and both are named in the
installer's output.
wrappingLabel's alignment must reach the cell — setAlignment on the
field alone left the title and stage line left-aligned inside
full-width frames ('Updating Hermes' starting at center-looking-box
left edge). Labels now get full-width frames and cell-level centering.
linger() also switched from a runloop pump to a blocking
NSThread.sleepForTimeInterval: with the animation stopped there are no
timer sources to service, so the pump returned immediately (terminal
states lived <1s instead of the designed 1.5s/15s).
The AppleScript panel rendered a stock NSProgressIndicator spinner and
could not port the shim's curve (AppleScript has no trig). Rewrite as
JavaScript for Automation: the ui.html curve math moves over
character-for-character (real JS), rendered into an 80x80 NSImage per
pump tick and swapped into an NSImageView — the JXA analog of the
page's requestAnimationFrame loop. Title, stage line, dark/light
seeds, terminal glyphs and the lingering rules all match ui.html's
apply().
JXA bridge notes baked into the code: variadic NSArray selectors do
not bridge (use the $() JS-array bridge), color selectors must use
bundled names (colorWithCalibratedRedGreenBlueAlpha), and 'delay' is a
reserved JXA global. spawn sites (posix.sh start_status_panel,
update_stage.ensure_panel) pass -l JavaScript; the .applescript panel
is deleted.
Verified on a macOS host: animates without beachballing, self-exits
~2s after a done publish and ~2s after the status file vanishes.
On a real update the panel stayed on 'Installing the new app' after the
app relaunched: the shim publishes 'done' and then removes the status
file, and the panel treated a missing file as 'keep waiting' — it only
ever exited if a poll happened to land in the half-second window before
the removal. Race lost, panel immortal.
A disappearance AFTER the file has been seen to exist now means done
(the shim removes the file only after publishing a terminal state and
tearing down its UI; an error publish keeps the file alive through the
15s leave-window grace, so a failure cannot be masked). A file that
never existed still means 'no status yet'. stop_ui also resets
UI_PANEL_PID with the other UI handles.
Verified on a macOS host: the panel compiles, survives while the file
exists, and self-exits ~3s after the file vanishes.
The panel polled the status file with 'delay 0.5', which blocks the
main runloop: AppKit never repaints, the label stays on its initial
'Preparing update...' and macOS beachballs the window - observed on
the first real desktop update, where the stage publishes were landing
in the status file but the panel never showed them. Drain
NSDefaultRunLoopMode for the poll interval instead, so label updates
and the progress animation render between ticks.
The shim renders progress from its status JSON file, but everything
after 'Updating code and dependencies' — the PM dependency sync, Node
deps, the TUI/web/desktop builds — ran for minutes without touching
that file, so the window froze on a stale stage (or showed nothing at
all when the old shim skipped its browser window).
hermes_cli/update_stage.py publishes stages to the watching UI,
std-only and never raising. Two discovery paths: the exported
HERMES_UPDATE_STATUS_FILE (current shim), and, for an OLD shim that
never exported it (every old-to-new checkout transition), the marker's
owner pid, which names the shim whose status file is deterministically
/tmp/hermes-update-status.<pid>. On macOS the takeover child also pops
update-panel.applescript from the freshly pulled tree when the old
shim's log shows it started no renderer.
Publishes: PM sync (_update_takeover.prepare, venv_sync.sync), Node
deps and each product build (source_build.build_update_products).
Only 'running' stages are ever written — terminal states stay the
shim's.
On macOS the shim only rendered through Chrome/Chromium honoring the
default-browser rule, so Safari/Firefox users (most of them) watched
the app quit and the update run invisibly - 'shim: no renderer;
skipping UI'. Add update-panel.applescript: an osascript/AppleScriptObjC
panel (NSWindow + label + indeterminate NSProgressIndicator, accessory
policy, no Dock icon) that polls the same status JSON write_status
emits - no HTTP server, no browser, no other-app scripting, hence no
TCC automation prompt.
start_status_panel is best-effort: osascript absent or the panel dying
instantly (headless session) falls back to today's UI-less behavior.
The panel self-exits after a terminal state (done after 1.5s, error/
manual after the leave-window grace so the message is readable) and
stop_ui KILLs it on early teardown, matching the SIG_IGN-survives-exec
contract of the HTTP server. Status fields are extracted with grep -
NSJSONSerialization's by-ref |error|: label does not parse under
osascript, and write_status output is flat JSON we control.
Verified on a macOS host: script compiles clean under osacompile; the
status reader returns running/done/missing/garbage correctly; GUI
rendering itself requires a WindowServer session (headless ssh cannot
show it) - first real desktop update exercises that path.
installer-script+desktop -> installer-script failed as "desktop output is
missing, stale, or damaged": the leg installs with the desktop, then re-runs the
plain one-liner, which built only tui/web -- while the driver's verifier still
expects the desktop product it installed (EXPECT_DESKTOP comes from the install
method). The artifacts live inside the tree, so an update makes them stale rather
than absent, and a desktop build left over from the previous code is exactly what
the freshness receipt rejects.
The products stage now selects the desktop when --include-desktop/-IncludeDesktop
is given OR the checkout already carries a built app. Verified: install.sh syntax
clean and the new predicate returns absent/present against real temp trees;
install.ps1 parses clean.
A checkout-sized tar sits silent for a minute plus, which reads like a
hang. bsdtar (macOS) and GNU tar disagree on progress options, so poll
the growing archive and overwrite the line every 2s; the loop doubles
as a liveness signal and the final wait still propagates tar's exit.
The backup root is typically the same internal disk, so the size saving
buys nothing; measured on an M1 over a 3.3G checkout, gzip made the
backup 5x slower (74s vs 14s). Store hermes-home.tar and
electron-userdata.tar uncompressed.
Hand-off kit for proving an existing source install can move to a
branch through the real update surfaces: pre (backup + arm a
transport-level insteadOf redirect at a serve.git of the target ref),
then 'hermes update', then post (rollback + restore-exactness report).
PLAN.md holds the design; smoke-test.* is the maintainer self-check.
The installer ladder stopped at node-deps/path/desktop with its own
semantics while an update ran launchers, product builds and post-build
maintenance, so a fresh install and a finished update ended in different
states: after re-running the installer at HEAD the products had no receipts
and the read-only source acceptance failed.
hermes_cli/source_completion.py now owns that tail -- publish launchers,
build the products, run the maintenance -- and update_completion's
_complete_selected calls it, so there is one implementation. install.sh and
install.ps1 keep the bootstrap stages (prerequisites, repository, venv,
python-deps, config) and hand off to it in a single `products` stage;
--include-desktop selects the desktop product inside that stage instead of
adding a second build stage, and `desktop` stays dispatchable via --stage for
external callers.
Windows keeps its installer-owned PATH publication (expose_cli answers
"windows-installer-owned" on Windows) plus the packaged-artifact probe, ACL
grant and shortcuts. The desktop stage no longer pre-syncs wake/voice: pm
lazy-installs them at first use, as the update path does.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
The demo was scaffolding, not repo content: nothing referenced it but a
docstring.
The tests were 12 against the repo's 1-2 invariant bar. Keep the two that fail
silently when inverted -- the sentinel reaching an exec'd child (the unexported
variable this change fixes) and the prologue's staleness gate, which inverted
either way costs a re-sync per run or a stale environment that looks fine. The
per-OS command pair stays because neither host can observe the other's string.
Dropped the exit-code/message restatement, the no-sentinel cold path, and the
demo-driven harness.
Path.glob raises NotImplementedError for a non-relative pattern, which a string `workspaces` (iterated char by char, so "/") or an absolute entry produces. The (OSError, ValueError, TypeError) catch missed it, so the error escaped to the caller's suppress(Exception) and no lock was reverted at all -- back to autostash every run. Non-list values are now ignored and each pattern is tried on its own so a bad one just owns nothing.
install.sh: read the workspace globs with `while read` instead of an unquoted $(...) so they are never pathname-expanded against the caller's CWD before `case` sees the pattern.
`scripts/install.sh::discard_update_lockfile_churn` and `scripts/install.ps1::Discard-LockfileChurn`
run the same per-directory predicate as `hermes update` did before the previous commit, so an
installer-driven update of a managed checkout (Desktop / bootstrap) reverted the root
`package-lock.json` whenever only `apps/desktop/package.json` was dirty, leaving spec and lock
out of sync for the next `npm ci`. Port the same ownership model: the root lock is kept when the
root manifest or any manifest matching a root `workspaces` glob is dirty; nested lockfiles are
still kept only with their sibling manifest; a manifest outside the graph still does not
protect the root lock.
install.sh reads the globs with sed/grep (no jq dependency) and matches with `case`; install.ps1
uses ConvertFrom-Json and `-like`. Bash side live-A/B'd in a throwaway repo (red on main, green
after; controls unchanged); the PowerShell side is the same shape and could not be executed on
this Linux host (no pwsh).
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:
- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
`git status --ignored=matching`, so a gitignored `.venv/` surfaced as
`!! .venv/` and refused every ZIP fallback on such installs ("the working
tree has uncommitted changes or untracked files") — the live runtime was
being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
place rather than growing a second environment that `project_venv_dir()`
then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
before this commit); the holder-guard test now uses a kernel-runner child
whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
it (red on origin/main); drop the `process.platform`-override vitest case,
which exercised the same resolver as the `.venv` case with a different
directory string.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
The committed lockfile still resolved body-parser 1.20.6 (nested qs 6.15.3),
express 4.22.2 with a top-level qs 6.15.3 and sharp 0.35.3, so the override
bump alone left `npm audit` at 3 findings (2 moderate qs, 1 high sharp) for
anyone installing from the lock — and `hermes doctor` kept flagging the
"WhatsApp bridge deps" row. `npm update --package-lock-only` inside the
existing manifest ranges: body-parser 1.20.8, express 4.22.3, qs 6.16.0,
sharp 0.35.4 -> `npm audit`: found 0 vulnerabilities. No manifest change
beyond the override bump; Baileys stays pinned at 7.0.0-rc13.
#109060 (Sep 12) was the earliest PR to move the override to 1.20.8 (it also
carried a redundant qs override, which 1.20.8 makes unnecessary).
Part of #112382
Co-authored-by: BenKalsky <1568840+BenKalsky@users.noreply.github.com>
The 1.20.6 pin sat inside the vulnerable range it was meant to clear (1.20.5 - 1.20.6); 1.20.8 pulls qs ~6.16.0, clearing the transitive qs advisories.
Repo scripts assume the PM-activated environment, so running one without
activation fails much later with a confusing ImportError. Add the two halves
covering both invocation paths:
- scripts/_activation.py: require_activation() exits immediately, naming the
exact command for the caller's shell (source ./activate on POSIX,
. .\activate.ps1 on a native Windows host), before any heavy import.
- scripts/_hermes-python: the POSIX shebang target. `#!/usr/bin/env -S bash -c
'...'` hands itself the target path through bash -c's $0, sources activate,
then execs the interpreter on the same file -- so tracebacks and __file__
still point at the real script and ./scripts/foo.py works from any cwd with
no manual source.
__HERMES_ACTIVATED changes from a bare "1" to the installed-state file the
environment was composed against, so one value carries activation, which
checkout activated it, and a staleness stamp. The prologue compares that file
against uv.lock / pyproject.toml / pm/lock.json with the `-nt` builtin -- no
process spawn -- and re-activates once when the inherited environment predates
its inputs. pm rewrites that file only on a real sync, so the check settles
back to current rather than re-syncing on every run.
A legacy "1" keeps working: require_activation() tests non-emptiness, and the
prologue's [ -e ] fails on it, so it activates once and upgrades.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
The canary tag regex capped the major at 3 digits while the stable
shape had no cap. The repo's current stable line is CalVer (v2026.9.14),
so canary_tag_for_date cuts v2026.9.15-canary.<ts> — which
handoff.validate_identity then rejected as 'Invalid release handoff
identity', killing every tag-mode stage leg (Windows and Darwin) while
the same-minor stable staged fine.
Lift the cap in the canonical _CANARY_TAG_RE and the Termux mirror
regex; add the stable-vs-canary shape-parity invariant, proven red on
the base regexes.
The termux-main pool deletes a package's previous archive when it rebuilds, so
the runtime-lib pin table and the bionic lock rows rot without warning. The last
rotation broke a build on eight rows at once, and the stager's concurrent
downloads only surfaced whichever 404 won the race.
pm now owns the pin table it repairs: scripts/termux/runtime_libs.json moves to
pm/termux_runtime_libs.json, so pins live in pm/ and scripts consume them — the
direction scripts/ci/archive_inputs.py already reads pm/lock.json in.
`hermes pm update --termux` repins exactly the rows whose archive the pool has
replaced, hashing each replacement against the index SHA256 before writing
url/version/hash together. `--check` reports without writing and exits 1, so a
retired pin can fail a cheap preflight instead of a payload build.
It is a repair, not an update: an alive pin is never moved, because a repin can
land a rebuilt library under a moved soname and the table is the payload's
recursive DT_NEEDED closure. A pin whose package the pool has dropped outright
is reported and left alone. `--termux` runs alone — names/--target/--uv/--npm
are ignored, since repairing foreign-target pins is not a version resolution.
Verified: `pm update --termux --check` against the live pool reports 89 rows
served; a table deliberately pinned to the retired libiconv 1.18-1 repins to
1.19 with the pool's hash through the real network path; 19 new tests; the
tests/pm, tests/ci and tests/scripts suites have the same failure set as the
base commit (91 pre-existing Windows environment failures, none new).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
The local --channel command no longer touches R2: it resolves the exact
pushed commit and dispatches the default-branch workflow, whose privileged
allocation step creates the channel and mints the immutable build request.
The disposable allocation path is generalized to cover the unscoped
production preview, gated by a new `channel` workflow input; build legs
consume the same channel-build / channel-request-sha256 job outputs as
before. The anti-tamper gate is now commit_build.admit (maintainer
permission) running inside the allocate step.
Drop the resume path: --resume-channel-build / --request-sha256 and
resume_build() are gone, and allocate_protected no longer recovers a lost
request PUT by sequence. Retrying re-dispatches and mints a fresh sequence
slot; idempotency survives via the deterministic build ID and the existing
immutable-request dedup. Keep the preview allocate `lastAllocation` build-ID
field (concurrent-CAS uniqueness), which is not resume.
channel_public_base now defaults to the documented production origin like
the commit-build path, so a local command names its page without a
hand-set CLOUDFLARE_R2_PUBLIC_URL.
Also drops a stale fork-isolation assertion left behind by the
fork-conditional dispatch removal.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
d8286ae58f made repository identity irrelevant to dispatch and made R2
disposable scoping opt-in via R2_DISPOSABLE_RUN, but the stable-releases
doc and the --build-commit help still described forks as requiring a
disposable allocation. Describe the single direct dispatch and the opt-in
scoped namespace instead.
The CI cache barely ever restored anything useful, for four independent
reasons found in live run logs and the fork's cache store:
- `uv cache prune --ci` (setup-pm/prune) discards downloaded wheels before
the save, so the snapshot carried only source-built wheels — the next
run "restored" it and still cold-downloaded everything. Upstream main's
logs showed the end state: a 209KB stub cache exact-hitting forever.
- Tool-only jobs (icons-freshness etc.) auto-saved a 1.5KB empty uv cache
under the production exact key; caches are immutable, so the stub won
forever and blocked real saves.
- The npm cache key had no restore-keys, so one lockfile bump missed the
exact key and every npm job went cold.
- `2efa4ff94f` deleted the electron-builder toolchain save but left the
assemble job restoring `eb2-` — a fossil nobody regenerates; assembly
cold-downloads winCodeSign/ATS/dotnet on every release.
Changes:
- pm/cache_lock.py: move prune_uv_cache_to_lock out of
scripts/bundles/native.py (which re-exports it); the lock-exactness
contract now serves both the bundle ship gate and CI caches.
- pm.build_env learns --exact-lock --lock-source: prune cache entries the
project uv.lock cannot resolve, keeping lock-required downloaded wheels
(unlike --ci). Refuses --ci/--prune-cache combinations.
- setup-pm: python-cache auto-save now requires extras (no stubs from
tool-only jobs); key drops the prune flag and bumps to v3 — pruned and
unpruned saves share one namespace since both are lock-exact; pre-save
pruning switched from --ci to --exact-lock; npm cache gains a
lockfile-agnostic restore prefix.
- save-pm-cache: same exact-lock prune before explicit saves.
- desktop-bundled-release: build legs (cache-mode: write) restore+save the
electron-builder toolchain under eb3- keyed on the locked builder
version; assemble restores the same namespace; the dead default-cache
resolution step is removed (assembly resolves no electron artifacts).
- cleanup_pm_toolchain_caches.py: match v2 and v3 smoke keys.
Validation: tests/scripts/test_bundle_native.py 11/11 (incl. both
lock-prune gates), tests/pm failures identical before/after the diff,
tests/scripts/test_bundle_payload.py 5/5, tests/ci cleanup 3/3,
tests-js setup-pm-post 1/1 and the three setup-pm-cache contract tests
updated and green (6 failures in that file pre-date this diff and are
drift between the workflow and its stale assertions); pm.build_env
--exact-lock E2E against a real uv cache copy pruned 3 stale entries and
kept the rest.
pillow-heif's sdist build failed with RequiredDependencyException: the
container toolchain installs pillow's libs but not libheif, so
pkg-config libheif had nothing to answer with. Termux's libheif 1.23.4
ships the headers and libheif.pc directly (no -dev split).
Runtime: the built _pillow_heif extension DT_NEEDEDs libheif.so, which
pulls libx265/libde265/libaom/libx264/libc++_shared. All but libheif
and libde265 were already staged in runtime_libs.json (verified by
reading the .so's DT_NEEDED, not the deb Depends line, which also
lists gdk-pixbuf/glib/rav1e the linker never loads); pin those two.
Import gate: import pillow_heif never fails on a dead link -- its
__init__ swallows the _pillow_heif ImportError into a DeferredError
that fires only on first use. The gate now imports _pillow_heif
directly so a dlopen failure fails the build instead of the phone.
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#89322 fixed the bridge-local normalizeWhatsAppId, but bridge.js has since
moved its id handling to bridge_helpers.js::normalizeWhatsAppId, which still
turned `<user>:<device>@lid` into the malformed `<user>@<device>@lid` for
mentionedJid / quoted participant / reaction keys, and the Python side
(gateway/platforms/whatsapp_common.py::_normalize_whatsapp_id) did the same
':'->'@' swap on botIds. Drop the local duplicate in bridge.js, import the
helper, and strip the `:<device>` suffix on both layers so the bot's own ids
compare equal to the bare ids WhatsApp sends for mentions and quotes.
One invariant test: device-qualified botIds match a bare mentionedId and a
bare quotedParticipant; a plain group message still does not trigger.
normalizeWhatsAppId did String(value).replace(':','@'), which turns a device-qualified id
like '116342762025117:14@lid' into the malformed '116342762025117@14@lid'. The bot's own id
(sock.user.id / sock.user.lid) carries the :<device> suffix while inbound mentionedJid and
contextInfo.participant (quoted message author) do not, so the bot's id never matches its
botIds set -> @mention and reply-to-bot are never detected in groups. Strip the :<device>
suffix instead so all id forms compare consistently.
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.
_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.
Includes a sabotage-verified unit test (fails without the scaler).
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.
Fix the class, not the sites:
* hermes_state: register every successfully constructed SessionDB in a
test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
everything left in the registry after each test. close() is idempotent
and unregisters the pinning atexit hook, so instances become
collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
leak fails fast with MemoryError instead of eating the box.
Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
for registration, idempotent close, and the cross-test sweep.
Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.
Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.
/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
The CI desktop-build cache restores by prefix after dependency changes, so
it accumulates wheels for superseded pins; measured 210MB of sediment on a
real build host. stage_uv_cache copied that snapshot wholesale — stale
wheels shipped in every bundle and grew monotonically.
prune_uv_cache_to_lock now deletes archive buckets and wheel/sdist index
entries the shipped uv.lock cannot resolve, after staging. The lock is the
ship contract: CI transport may roll and accumulate; the payload never
ships sediment. Unidentifiable buckets (no dist-info) survive — pruning
fails open for unknown layouts, never for identifiable stale pins.