The termux-main pool deletes a package's previous archive when it rebuilds, so
the runtime-lib pin table and the bionic lock rows rot without warning. The last
rotation broke a build on eight rows at once, and the stager's concurrent
downloads only surfaced whichever 404 won the race.
pm now owns the pin table it repairs: scripts/termux/runtime_libs.json moves to
pm/termux_runtime_libs.json, so pins live in pm/ and scripts consume them — the
direction scripts/ci/archive_inputs.py already reads pm/lock.json in.
`hermes pm update --termux` repins exactly the rows whose archive the pool has
replaced, hashing each replacement against the index SHA256 before writing
url/version/hash together. `--check` reports without writing and exits 1, so a
retired pin can fail a cheap preflight instead of a payload build.
It is a repair, not an update: an alive pin is never moved, because a repin can
land a rebuilt library under a moved soname and the table is the payload's
recursive DT_NEEDED closure. A pin whose package the pool has dropped outright
is reported and left alone. `--termux` runs alone — names/--target/--uv/--npm
are ignored, since repairing foreign-target pins is not a version resolution.
Verified: `pm update --termux --check` against the live pool reports 89 rows
served; a table deliberately pinned to the retired libiconv 1.18-1 repins to
1.19 with the pool's hash through the real network path; 19 new tests; the
tests/pm, tests/ci and tests/scripts suites have the same failure set as the
base commit (91 pre-existing Windows environment failures, none new).
The local --channel command no longer touches R2: it resolves the exact
pushed commit and dispatches the default-branch workflow, whose privileged
allocation step creates the channel and mints the immutable build request.
The disposable allocation path is generalized to cover the unscoped
production preview, gated by a new `channel` workflow input; build legs
consume the same channel-build / channel-request-sha256 job outputs as
before. The anti-tamper gate is now commit_build.admit (maintainer
permission) running inside the allocate step.
Drop the resume path: --resume-channel-build / --request-sha256 and
resume_build() are gone, and allocate_protected no longer recovers a lost
request PUT by sequence. Retrying re-dispatches and mints a fresh sequence
slot; idempotency survives via the deterministic build ID and the existing
immutable-request dedup. Keep the preview allocate `lastAllocation` build-ID
field (concurrent-CAS uniqueness), which is not resume.
channel_public_base now defaults to the documented production origin like
the commit-build path, so a local command names its page without a
hand-set CLOUDFLARE_R2_PUBLIC_URL.
Also drops a stale fork-isolation assertion left behind by the
fork-conditional dispatch removal.
The CI desktop-build cache restores by prefix after dependency changes, so
it accumulates wheels for superseded pins; measured 210MB of sediment on a
real build host. stage_uv_cache copied that snapshot wholesale — stale
wheels shipped in every bundle and grew monotonically.
prune_uv_cache_to_lock now deletes archive buckets and wheel/sdist index
entries the shipped uv.lock cannot resolve, after staging. The lock is the
ship contract: CI transport may roll and accumulate; the payload never
ships sediment. Unidentifiable buckets (no dist-info) survive — pruning
fails open for unknown layouts, never for identifiable stale pins.
Repository identity no longer selects behavior: a commit build is the
same direct dispatch from any repository, and R2 disposable scoping is
opt-in via R2_DISPOSABLE_RUN rather than fork-mandated. The fork guard
in the workflow admission step, the fork refusal in R2Scope.configured,
and the client-side fork routing (disposable_dispatch_command) all go.
Disposable namespaces keep their own protections: malformed leases are
refused, and a scoped namespace still belongs to exactly one repository.
Un-slim the bundle's uv cache: prune_uv_cache_to_built (built-only keep-set)
is deleted — every wheel the build resolved ships, so any per-install venv
rebuild (plugin extras, feature changes) resolves offline with zero network.
stage_uv_cache keeps excluding sdist src/ trees (build-only bulk) but now
KEEPS built wheel ZIPs including the copies beside sdist metadata.msgpack
shards: probing showed an offline rebuild whose revision entry lacks its
wheel ZIP fails closed wanting the sdist download. Fixture sdist uses a
stdlib-only PEP 517 backend so the hermetic offline rebuild needs no build
tools.
First launch of a bundled payload paid a cold-compile stall: the launcher
redirects bytecode writes to a user-level cache (signature-breaking on
macOS, read-only mount on AppImage/MSIX), so every import compiled from
source. Now staging bakes the cache into the payload:
- compileall with the payload's OWN staged 3.14 interpreter, unchecked-
hash pycs: repack mtimes cannot invalidate them, a stale source can
never trigger a rewrite, and read-only pycs mean the macOS signature
never observes a change. Dirs stay writable — in-place rebuilds
rmtree the tree; asserted coverage plus unchecked-hash means no
cache-miss write can target them.
- coverage is the perf contract: the bake FAILS if any parseable module
lacks a pyc (empirically 0 unparseable files ship, so compileall is
strict). Probe suite: py_compile/cache_from_source, PEP 552 flags,
multi-root read, stale-source no-rewrite, read-only cache-dir import.
- launcher: the baked marker makes configure() leave sys.pycache_prefix
UNSET — the prefix relocates reads too and would hide the baked pycs.
Payload modules read their source-adjacent cache (Python's default
multi-root lookup); plugin/user modules keep caching beside their own
sources under HERMES_HOME. Unmarked payloads keep the old redirect.
- snapshot(): the sealed payload ships without tests/website/evals/
.github/nix/docker/tests-js (~69MB, 46% of tracked bytes) and without
apps/ui-tui/web/scripts — CI prebuilds those products, and
is_bundled_payload routes sealed updates to the channel updater, so
the rebuild graph never runs in a bundle (linux_desktop_entry degrades
to the themed icon). Frontend product staging keeps the full tree.
- test_bundle_native now stages the FULL relocatable toolchain (a bare
interpreter ELF falls back to its compile-time /install prefix and
cannot create a venv), and runs on the real 3.14 for the first time
this campaign — the whole battery had been running 3.12 against the
3.14-pinned lock.
A fork commit build is now ONE workflow dispatch whose run both allocates
the disposable channel and builds it. Previously allocation printed a
follow-up command that had to be dispatched separately (the GITHUB_TOKEN
recursion wall forced two runs; release.py grew poll/extract machinery to
automate the hop — all deleted now, net -144 lines).
- Lease is the run id alone (no attempt suffix): re-run failed jobs
re-enters the same namespace; succeeded allocate job is skipped and its
outputs persist. r2_scope accepts legacy <id>-<attempt> leases on read.
- allocate-disposable emits job outputs (channel_build, request digest,
lease, public base); validate consumes them in disposable mode and
re-exports a normalized pin every downstream job reads.
- Admission guards unchanged in semantics: forks still cannot run without
a disposable allocation; disposable runs still never touch production
feeds, termux, or the commit-builds page.
- Upstream trusted-controller path (channel_build inputs) byte-identical.
- Fixtures updated for run-id leases; two dead two-dispatch tests and the
fixture's allocation-probe plumbing removed; smoke matrix test skips
banana's new admission job (no toolchain by design).
Fork CI now requires a disposable channel allocation (R2_DISPOSABLE_RUN
guard in desktop-bundled-release.yml), but 'release.py --build-commit REV
--publish' still fired the old direct dispatch, so every fork commit build
died at admission. cmd_build_commit now detects a non-upstream repository
(case-insensitive NousResearch/hermes-agent compare), dispatches the
allocation workflow with disposable_channel/build_commit/bundle_env baked
in (all --bundle-env/--bundle-unset values travel inside the immutable
request; the follow-up never re-passes them), polls the allocation run to
completion (15s interval, 15min budget), extracts the printed follow-up
dispatch from the run logs (channel_disposable's single-line JSON
'command'), validates its shape (gh workflow run of this workflow against
the same repository), and auto-dispatches it with the local maintainer's
gh login — falling back to a clear run-summary pointer when log recovery
fails. Upstream behavior is unchanged. Also fixes a pre-existing TypeError
that masked check_output failures whose CalledProcessError has stderr=None.
The desktop payloads shipped the builder's full uv cache (2.0GB mac,
1.1GB win-x64, 800MB win-arm64) so a mutable-venv rebuild could run
offline. But the bundle contract never needed offline rebuilds — it
needs rebuilds that never invoke a compiler: packages with no
downloadable wheel (sdist-only or platform gaps) would otherwise
demand Xcode CLT/MSVC/Rust on the user's machine.
prune_uv_cache_to_built keeps only what the builder compiled itself,
detected from the cache's own records (sdists-v9 entries, cached
wheels not listed in uv.lock, archive-v0 buckets by dist-info) and
drops the ~245 packages whose wheels PyPI re-serves in seconds.
Validated live on a clean host per target:
- macos arm64 (iris, no uv): 2.0GB -> 3MB; sync online, 0 builds,
pilk/psutil/alibabacloud-tea installed from the slim
- win11 arm64 (promise, no uv/MSVC): 800MB -> 48MB; sync online,
0 dependency builds, all 12 CI-built packages (cryptography,
httptools, brotlicffi, dependency-injector, ruamel-yaml-clib,
obstore, davey, firecrawl-anydoc, pilk, psutil, alibabacloud set)
import clean. The kept set matched the CI build log's Built list
exactly (12/12, no false positives).
Index-dir detection is name-agnostic (pypi/ vs custom --index-url
hashed dir), proven by the rewritten bundle test: sdist-only survives
the slim, downloadable wheels are dropped, the build machine's cache
is untouched, and an online rebuild compiles nothing.
The e2e-screen-record action installed ffmpeg through three different
OS package managers (apt, brew, winget). winget is the flaky leg on
windows-11-arm and serves an x64-gyan build that runs emulated on ARM;
choco's community package wraps the same gyan x64 zip, so a choco
fallback would not fix either problem. PM already ships a locked,
sha256-pinned ffmpeg (martin-riedl posix, BtbN win32 including a NATIVE
winarm64 build), so the recording action now verifies ffmpeg on PATH
instead of installing it, and the pinned binary rides the same
tools-cache as node/python.
- setup-pm learns a `packages` input (extra PM tools beyond the
toolchain roots, e.g. ffmpeg) threaded through setup_toolchain.py's
prepare/install/archive-inputs phases; the tools-cache key gains an
extra-packages fragment so existing keys stay byte-identical.
- The six chat-driver jobs pass `packages: ffmpeg` to setup-pm.
- e2e-screen-record drops the apt/brew/winget install steps, the
ffmpeg actions/cache steps and the save-cache input; Xvfb (headless
linux) and the macOS replayd-approval hack stay.
- The "Install locked chat driver dependencies" step installs only the
tests-js workspace with --omit=dev instead of the whole apps/desktop
tree: the drivers need @playwright/test, zod (previously a phantom
hoisted from @assistant-ui/react), js-yaml and semver only. 64 pure
packages in ~2s vs ~1800 including electron-builder and native
builds; the tree-shaken node_modules runs the real driver modules
(verified by importing desktop-chat-smoke.ts and update-window-chat
end to end).
- tests pin the new contracts; tests/install/README.md updated.
Share the real composer, provider-witness and completed-reply check across
post-build bundle smoke and desktop-bearing install/update checkpoints.
Keep native automatic-relaunch proof separate from post-update chat.
Download receipt-bound artifacts without release credentials and install
DMG, ZIP, MSIX and universal MSIXBUNDLE on each native architecture.
Split Windows assembly from feed publication; publish tested bytes only.
Bind candidate smoke results into the manifest used by stable promotion.
Verify historical/source provenance without assuming a version IPC commit,
strip CI identity from source build children, and use the actual Electron
PID rather than Playwright's Windows launcher wrapper.
Validation: real Linux Electron chat and sequential OLD/NEW source smoke
with preserved history; 145 targeted Python tests and 14 JS tests passed;
TypeScript, shell/PowerShell parsing and workflow checks passed.
Native macOS/Windows deployment and historical upgrades need Actions proof.
Dependency acquisition during packaging left native wheels and packager
inputs outside the pre-build cache save. Compose PM and existing providers
into a preparation phase, then require builds to consume admitted inputs.
Share native preparation with PM Bundle. Keep path-bound environments and
signing outputs separate from reusable caches. Use read-only cache tokens
for commit builds and preserve the one-command local build path.
Verify pinned tools through PM, probe PTYs under the prepared Electron,
and supply dmgbuild through a build-only PM package. Resolve bundled tool
stores from their payload manifest so relocation preserves discovery.
Validation: focused Python and JS tests, checkJs, Ruff, Windows checks,
anti-slop, cache relocation, and network-denied Linux AppImage builds.
Relocated runtime smoke passed with NixOS host libraries supplied.
Native Windows/macOS signing and live GitHub cache behavior remain untested.
Keep the minimal sandbox command's existing app/data identity while sharing
its parser, seeding, execution and cleanup with the standard sandbox.
Remove the unused SDK adapter copy and protocol shadow from the environment
facade, plus unconsumed checkout-updater dependencies and their dead helper.
Validation: 72 focused Python tests across sandbox and adapter/backend
paths, 5 checkout-updater tests, Electron typecheck and ESLint passed.
Ruff, touched-file Windows checks and the TS ratchet pass. Sandbox tests
also passed against the original scripts before consolidation.
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
Keep desktop rollback in staged publication instead of restoring backups
inside a candidate directory that will be discarded.
Use the npm tag endpoint, let the downloader own archive verification,
and ensure CI toolchain roots without repeating dependency verification.
Remove an unused resolver input and replace overlapping tests with
fault-proven lifecycle coverage.
Verified with the scoped Python gate and before-pack tests. Native
Windows/macOS update execution and the full suite were not run.
Preserve upstream fixes without restoring retired dependency installers.
Run configured-feature checks in the selected build interpreter. Reuse a
supported base Python during bootstrap, and preserve durable backup media.
Refresh the dependency lock through PM. Keep the frozen historical import
surface unchanged. Adapt incoming native tests to the platform markers.
Verification: the incoming 86-file pass found two fixture mismatches;
both passed after correction. Targeted PM/update/compatibility checks,
Electron and renderer typechecks, and desktop tests passed.
Native Windows/macOS update journeys and the full suite remain unrun.
Run historical updater completion in a fresh interpreter so cached imports
cannot revive retired dependency installers. Share Git and ZIP completion,
carry receipt and recovery state, and preserve child exit status.
Route plugin admission, binary acquisition, desktop launch and build paths
through PM. Replace redundant helpers and tests with real worker, package,
publication and launch checks. Keep the shipped compatibility surface fixed.
Targeted Python and desktop checks pass. Native update journeys and fresh
production image qualification remain pending. This is a checkpoint before
those acceptance runs.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Give native payload dependency builds two hours without changing the normal install timeout. Allow three hours for the standalone PM bundle job so setup, cache saves, and smoke tests fit around the build.
Verified seven focused tests, Ruff, and actionlint. Full CI builds were not rerun.
Resolve the default cache before isolating HOME so payload builds use the directory that CI restores and saves. Remove the standalone bundle workflow dependency on an unset cache variable.
Verified offline wheel reuse in isolated children, 20 focused tests, Ruff, and actionlint. Full native release builds and the separate source-build timeout remain unverified.
Commit builds have no update channel, even when installed through dpkg. Read the installed stamp before checking the refusal message and require exit code 2 for both artifact kinds.
Verified real CLI refusal paths, negative controls, and related tests: 15 passed. Ruff and type checks passed. Full deb validation was not rerun.
Every scheduled skills-index.yml run since 2026-07-20 was cancelled at the 15-minute
job timeout, so the live skills-index.json has been frozen at that date and every
`hermes skills search` fell through to live GitHub API calls (~500 inspect requests per
cold search, against a 60/hr unauthenticated budget). The freshness watchdog has been
appending to #66616 four times a day since.
Root cause: enrich_owners() walks every ClawHub skill's detail endpoint (~2s each) to
fetch an owner handle for the "View source" link. The catalog grew from ~50k to 78k
skills, so even at 30 workers that phase alone runs over an hour; nothing bounded it.
- enrich_owners() gains budget_seconds: on expiry it stops and ships the remainder
without an owner (the link is a nicety; the index is not).
- build_skills_index.py passes an 8-minute budget.
- skills-index.yml build job timeout 15 -> 50 min to cover the measured critical path
(clawhub walk ~14 min || github taps ~8 min, skills.sh resolve ~6 min, enrichment 8 min).
Reuse a matching ARM64 Visual Studio environment instead of growing its paths on each initialization. Preserve Rust and Cargo homes before isolating bundle state. Hand the private Termux assembly directory to its container user and restore host ownership afterward.
Verified targeted tests, real ARM64 SDK/OpenSSL execution, Rust compilation, and container facts publication. Full release packages were not rebuilt. The Windows x64 source-build timeout and the unchanged Termux linkage fixture failure remain unresolved.