A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.
Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.
Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.
Phase 2 of #88715.
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.
The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
Use the configured group policy and group-JID allowlist at Node bridge intake instead of applying the DM sender allowlist to group participants.
Co-authored-by: Martin Gontovnikas <m@gon.to>
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
The bare-flag check asked pytest's parser which tokens it does not know,
but wrapped parse_known_args in the same blanket except that guards
parser construction. A known flag with a missing value (`--tb` alone)
raises pytest.UsageError there, which the except turned into "nothing
unknown", so discovery ran and every per-file pytest died with
"argument --tb: expected one argument".
Keep the fallback around building the parser only; let parse_known_args
run outside it and surface UsageError (and the unknown-token list) as
this runner's own usage error before discovery. One invariant test.
A bare token this runner does not own used to be forwarded to every per-file
pytest, so a typo (`--jbs`, or `--help` before #114065's fix) discovered the
whole suite and each file died with "unrecognized arguments" — hours to learn
about a typo. Validate the bare passthrough tokens against pytest's own
argparse parser (installed plugins loaded), and fail once with this runner's
usage (exit 2) before discovery. argparse handles the attached-short-value
(`-rA`), combined-flag (`-xvs`) and `-k expr` forms, so real pytest flags keep
passing through; tokens after a literal `--` are the caller's explicit choice
and are never validated. If pytest's parser cannot be built, the check is
skipped and behaviour is unchanged.
Follow-up to KoNit-K's `-h`/`--help` interception (#114065). Fixes#114059.
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
Path.glob raises NotImplementedError for a non-relative pattern, which a string `workspaces` (iterated char by char, so "/") or an absolute entry produces. The (OSError, ValueError, TypeError) catch missed it, so the error escaped to the caller's suppress(Exception) and no lock was reverted at all -- back to autostash every run. Non-list values are now ignored and each pattern is tried on its own so a bad one just owns nothing.
install.sh: read the workspace globs with `while read` instead of an unquoted $(...) so they are never pathname-expanded against the caller's CWD before `case` sees the pattern.
`scripts/install.sh::discard_update_lockfile_churn` and `scripts/install.ps1::Discard-LockfileChurn`
run the same per-directory predicate as `hermes update` did before the previous commit, so an
installer-driven update of a managed checkout (Desktop / bootstrap) reverted the root
`package-lock.json` whenever only `apps/desktop/package.json` was dirty, leaving spec and lock
out of sync for the next `npm ci`. Port the same ownership model: the root lock is kept when the
root manifest or any manifest matching a root `workspaces` glob is dirty; nested lockfiles are
still kept only with their sibling manifest; a manifest outside the graph still does not
protect the root lock.
install.sh reads the globs with sed/grep (no jq dependency) and matches with `case`; install.ps1
uses ConvertFrom-Json and `-like`. Bash side live-A/B'd in a throwaway repo (red on main, green
after; controls unchanged); the PowerShell side is the same shape and could not be executed on
this Linux host (no pwsh).
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:
- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
`git status --ignored=matching`, so a gitignored `.venv/` surfaced as
`!! .venv/` and refused every ZIP fallback on such installs ("the working
tree has uncommitted changes or untracked files") — the live runtime was
being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
place rather than growing a second environment that `project_venv_dir()`
then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
before this commit); the holder-guard test now uses a kernel-runner child
whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
it (red on origin/main); drop the `process.platform`-override vitest case,
which exercised the same resolver as the `.venv` case with a different
directory string.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
The committed lockfile still resolved body-parser 1.20.6 (nested qs 6.15.3),
express 4.22.2 with a top-level qs 6.15.3 and sharp 0.35.3, so the override
bump alone left `npm audit` at 3 findings (2 moderate qs, 1 high sharp) for
anyone installing from the lock — and `hermes doctor` kept flagging the
"WhatsApp bridge deps" row. `npm update --package-lock-only` inside the
existing manifest ranges: body-parser 1.20.8, express 4.22.3, qs 6.16.0,
sharp 0.35.4 -> `npm audit`: found 0 vulnerabilities. No manifest change
beyond the override bump; Baileys stays pinned at 7.0.0-rc13.
#109060 (Sep 12) was the earliest PR to move the override to 1.20.8 (it also
carried a redundant qs override, which 1.20.8 makes unnecessary).
Part of #112382
Co-authored-by: BenKalsky <1568840+BenKalsky@users.noreply.github.com>
The 1.20.6 pin sat inside the vulnerable range it was meant to clear (1.20.5 - 1.20.6); 1.20.8 pulls qs ~6.16.0, clearing the transitive qs advisories.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
scripts/check_profile_scope_patterns.py runs the validated hazard regexes in
scripts/ci/profile_scope_patterns.json (18 of the 31 campaign patterns: every one has a scope_hint
and hits <= 50 sites on main; the wider ones are review greps, not lint) against the lines added
vs the PR base and prints file:line, pattern id/class and why. Always exits 0: most shapes have
legitimate sites (a standalone `hermes -p x` process where environ IS the profile), so the
reviewer reads each finding against its scope hint. Wired into lint.yml beside the public-surface
diff with continue-on-error.
Proof: the pre-fix tools/bot_relay.py (`env = dict(os.environ)`, before the served_profile_child_env
change) is flagged as P05/C2; the fixed file and this branch's diff vs main report 0 findings.
Test: a fixture with the hazard is flagged on the right lines, the scoped-builder version is not,
and the line filter hides hits outside the added range.
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
the long-wait rate-limit check read the turn's extract_api_error_context()
dict, which never carries welcome_refusal / welcome_route. They now read
classified.error_context, where _nous_welcome_tier parks them; the guard
records the classifier's reset_at. Tests drive the real classifier and the
real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
account (anon_account_locked) is now retired without replacement, matching
the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
re-inventories, so a provider connected during the cooldown keeps
inference.
- The desktop's setup.ready listener only refreshes an untouched picker
(oauth mode, no local endpoint, idle flow) and re-checks after the
readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
lock that _send re-acquires; the log is copied out first.
Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#89322 fixed the bridge-local normalizeWhatsAppId, but bridge.js has since
moved its id handling to bridge_helpers.js::normalizeWhatsAppId, which still
turned `<user>:<device>@lid` into the malformed `<user>@<device>@lid` for
mentionedJid / quoted participant / reaction keys, and the Python side
(gateway/platforms/whatsapp_common.py::_normalize_whatsapp_id) did the same
':'->'@' swap on botIds. Drop the local duplicate in bridge.js, import the
helper, and strip the `:<device>` suffix on both layers so the bot's own ids
compare equal to the bare ids WhatsApp sends for mentions and quotes.
One invariant test: device-qualified botIds match a bare mentionedId and a
bare quotedParticipant; a plain group message still does not trigger.
normalizeWhatsAppId did String(value).replace(':','@'), which turns a device-qualified id
like '116342762025117:14@lid' into the malformed '116342762025117@14@lid'. The bot's own id
(sock.user.id / sock.user.lid) carries the :<device> suffix while inbound mentionedJid and
contextInfo.participant (quoted message author) do not, so the bot's id never matches its
botIds set -> @mention and reply-to-bot are never detected in groups. Strip the :<device>
suffix instead so all id forms compare consistently.
Greptile's two findings on the original PR were both right.
1. The scaler read test_durations.json from the checkout, but CI ran on
a fresh runner where that file never exists (it is gitignored and the
slicing-era artifact/merge job that produced it is gone). The feature
was inert exactly where the false FLAKY kills happen. tests.yml now
restores the most recent main-saved cache before the run (PRs read
only) and saves it after a green push to main, mirroring the
ci-timings-baseline restore/save pattern already in ci.yaml.
2. _save_durations persisted every file's total subprocess wall,
including the ~cap of a timed-out attempt and the retry-summed wall
of a FLAKY file. With the scaler that compounds: a hang cached at
~300s earns 900s next run, then ~900s cached earns 2700s, until the
job timeout is the only bound. _clean_pass_durations drops failed and
FLAKY files from the write so a file's cached duration is always a
first-attempt-clean measurement; those files keep their previous
known-good entry.
Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.
_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.
Includes a sabotage-verified unit test (fails without the scaler).
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.
Fix the class, not the sites:
* hermes_state: register every successfully constructed SessionDB in a
test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
everything left in the registry after each test. close() is idempotent
and unregisters the pinning atexit hook, so instances become
collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
leak fails fast with MemoryError instead of eating the box.
Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
for registration, idempotent close, and the cross-test sweep.
Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.
Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.
/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
The staleness test regenerates in the Python lane, which has no
node_modules; prettier-dependent output would make the check pass locally
and fail in CI (or the reverse). Single-quoted literals, bare identifier
keys, no trailing commas or whitespace — prettier --check is clean on the
committed file.
215 methods, 13 server→client requests and 67 notifications now have Pydantic
contracts under tui_gateway/contracts/<topic>.py, rendered to
apps/shared/src/gateway-contract.generated.ts (616 types) and
gateway-contract.openrpc.json. tests/contracts/test_generated.py pins both
files to an in-memory regeneration and asserts catalog completeness from the
CODE side (every registered handler / emitted event / sent request has a
contract, nothing orphaned). scripts/ci/classify_changes.py runs the Python
lane when either generated file changes.
Phantom fields the hand-typed TS carried and no emitter ever set:
tool.start.todos, error.reason, voice.transcript.voice_stopped.
tui_gateway/contracts/ is the single source of truth for the JSON-RPC wire:
Params/Result/Payload bases, a registry of METHODS / SERVER_REQUESTS /
EVENTS, runtime validation (unknown/mistyped params answer 4000 with the
field path; a result or payload that violates its model raises under the
test suite and logs once in production), and the server-request contracts
+ shared value shapes as the authoring template. scripts/gen_gateway_contracts.py
renders the tables through a small JSON-Schema-subset walker into
apps/shared/src/gateway-contract.generated.ts and
gateway-contract.openrpc.json (unsupported constructs raise at generation).
Method contracts for the 213 remaining handlers follow in the next commits.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
scripts/dump_desktop_slash_registry.py --check exits 1 (no write) when the
committed apps/desktop/src/lib/desktop-slash-registry.json differs from
desktop_surface_registry(); the docstring now cites the test that actually
fails (test_desktop_slash_registry.py, not test_commands.py). The pytest
runs --check as a subprocess (green on the committed file) and calls
check() on a hand-edited copy (red).
scripts/ci/classify_changes.py: _PY_SKIP includes apps/, so an apps-only
edit of that JSON (or apps/shared/src/gateway-events.json) skipped the
Python lane and the only equality check never ran in CI. A table of
cross-language contract files now forces python:true; two classifier rows
pin it.
34 of the 46 `NO_DESKTOP_SURFACE` rows in desktop-slash-commands.ts were a
byte-for-byte copy of `desktop=` on the matching CommandDef in
hermes_cli/commands.py (7 more were aliases of those rows). The live
`commands.catalog` already carries that metadata; the static list was the
offline fallback and would silently drift on the next registry edit.
Now `hermes_cli/commands.py::desktop_surface_registry()` is the one author
of `/name -> desktop` (aliases included). `scripts/dump_desktop_slash_registry.py`
writes it to apps/desktop/src/lib/desktop-slash-registry.json, which the
desktop imports as its offline fallback (`registryUnavailableSpecs`). Five
names the Python registry has never heard of stay in an explicit
`TS_ONLY_NO_DESKTOP_SURFACE` with the reason WHY: `/density /details /logs
/mouse` are Ink-process-local display toggles (handled in
ui-tui/src/app/slash/commands/core.ts; advertised via `_TUI_EXTRA`), and
`/pets` is the plural typo of the desktop's own `/pet` action. `/switch`
was never a block-list row (it is a `/resume` alias) — not a finding.
Cross-language contract: tests/hermes_cli/test_desktop_slash_registry.py
asserts the committed JSON == desktop_surface_registry() and that every
alias carries its canonical value; desktop-slash-commands.test.ts asserts
every dumped row is unavailable/unsuggested offline with the dumped reason
and that the TS-only set is disjoint from the dump. Both sides fail on
drift (sabotage: flipping one `desktop=` in commands.py -> Python test
"stale"; adding `/clear` to the TS-only list -> vitest disjointness fails;
dropping `registryUnavailableSpecs()` -> 4 existing vitest cases fail).
Behavior change: none for users. `/model` (`desktop="hidden"`) keeps its
local picker spec; `hidden` is a popover flag read from the live catalog.
Sites: apps/desktop/src/lib/desktop-slash-commands.ts::NO_DESKTOP_SURFACE
(46 rows) -> hermes_cli/commands.py::desktop_surface_registry (41 rows via
the dump) + TS_ONLY_NO_DESKTOP_SURFACE (5 rows). apps/desktop/src/AGENTS.md
updated.
Two CI reds from the previous commits.
benchmark_browser_eval.py moved from scripts/ (exempt) to evals/ and its
bare shutil.which("npx") tripped test_no_unreviewed_bare_managed_runtime_
lookups. evals/ are standalone user-invoked benchmark programs, the same
class as scripts/ and skills/, and three other eval files already call
which() the same way; the guard just never saw them because the harness
that carried them lived under scripts/. Exempt the directory.
list_os_marked_tests.py used Path.rglob, which raises FileNotFoundError
when a __pycache__ directory disappears mid-scan (a sibling job in the
same workspace). The managed-runtime guard already switched to os.walk
for the identical TOCTOU; do the same here. Same output, sorted.
scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".
- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
-> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
_PY_SKIP includes apps/, so a change to apps/shared/src/gateway-events.json
(or apps/desktop/src/lib/desktop-slash-registry.json) alone classified as
python:false and tests/tui_gateway/test_gateway_event_contract.py — the
only side that checks emitters against the JSON — never ran. "Either side
drifting turns both red" was true locally and false in CI for that
direction. A table of cross-language contract files now forces python:true
alongside frontend; two classifier rows pin it (they fail without the
allowlist).
reseed_if_terminal created its temp as <auth>.rebootstrap.<pid>.tmp with
O_CREAT|O_EXCL. Boot-hook PIDs inside a container are near-deterministic,
so a run SIGKILL'd between create and replace leaves a same-named file and
every later boot hits FileExistsError - which main() swallows as
"error (ignored)", leaving the terminal-session recovery path dead until
someone deletes the temp by hand.
tempfile.mkstemp in the auth dir gives a random name at 0600 (stdlib only,
matching the script's no-hermes-imports rule); the fsync + os.replace +
unlink-on-failure semantics are unchanged.
Ten hand-rolled "write a token file safely" routines each carried a
different subset of {0600-on-create, fsync, atomic_replace, parent-0700
guard, BaseException cleanup}. Two of them (iron_proxy state files,
the exchanged-JWT store) still opened the temp file at process umask
and chmod'ed afterwards - the exact TOCTOU window the others document
as fixed. None of the bare-os.replace copies got atomic_replace's
Windows-contention retry or EXDEV fallback.
utils gains fsync_dir= (absorbs auth.py's dir fsync), atomic_write_bytes
(vault blob) and mode= on atomic_write_text; the ten sites become 1-3
line callers. mkstemp creates the temp file O_EXCL at 0600 regardless of
umask, so the payload is never umask-readable.
Behavior change: iron_proxy proxy.yaml/mappings.json and the exchanged-JWT
store are now 0600 from creation and fsync'd; every credential write goes
through atomic_replace (symlink-preserving, Windows retry, EXDEV copy).
auth_nous shared store now uses atomic_replace too (it forced os.replace
with no recorded reason). secret_sources cache parent-0700 goes through
the guarded secure_parent_dir instead of an unguarded chmod.
Review finding: with the install's own venv activated (a re-run), `uv python
find <range>` returned venv/bin/python3, which setup_venv then deleted before
handing the dead path to `uv venv`; the stage still exited 0 and printed
"Virtual environment ready". Probe with --system so only base interpreters
qualify, and fail the stage when the venv was not created.
check_python only asked uv for 3.11 and, missing that, downloaded it —
a hard failure on hosts that cannot reach GitHub releases even when a
3.12/3.13 inside requires-python is already installed. Ask uv for the
supported range second and pin the venv onto that interpreter.
Same direction as #10824 (@gnanirahulnutakki), reduced to the shell
path; the PowerShell installer already has this fallback.
Refs #10778
`tar xf node-*.tar.xz` shells out to the xz binary; minimal Debian, DietPi
and WSL images ship tar without it, so extraction died mid-way and the
installer then failed on a missing directory. Select .tar.xz only when
`xz` is on PATH, in both the installer and the runtime node bootstrap.
Same approach as the earlier #4229 (@JoshuaMart) and #39541 (@karnull);
#11278 (@vominh1919) attempted an apt-only install of xz-utils instead.
Refs #11197
Review finding on #109142: the JS bridge's DEFAULT_REPLY_PREFIX (the sender
actually used at runtime), scripts/install.sh, setup-hermes.sh and the site
favicon still carried the old glyph. Python and JS defaults now agree.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).