`release.py release` gains two flags. They can be used together.
--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.
--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.
The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.
The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.
A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:
- The next version was derived from it, so the next cut would reuse the
version. It now takes the newer of the R2 head and the newest
published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
their publication pass, so a bundle-less release would re-advance
every 15 minutes. The head is now the newer of the R2 head and the
published release whose final tag binds the Docker stable alias
digest.
`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.
Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:
- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]
Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
tools/lazy_deps.py was not deleted; it survives as an old-updater stub
that raises or stops for relaunch. The job display name and the
checker/test prose claimed it was deleted, which sends readers looking
for a missing file. The display name is not a required status check
(main requires only "All required checks pass", which keys on job ids),
so rename it to "No production imports of the tools.lazy_deps stub".
Stable Release Publication ran every 15 minutes (96 runs a day, each
checking out full history, setting up node and buildx, logging into
Docker Hub, and taking the release-signing environment) only because the
sequencer held a failed run for a 15-minute backoff that the failure
event could never satisfy, so the cron was what actually retried.
Drop the backoff: the reconcile pass started by a failed Stable Release
reruns its failed jobs right away. MAX_ATTEMPTS burning, oldest-first
retry ordering, the attempt-entry check, and the needs_retarget repair
stay. The schedule trigger goes; workflow_run and workflow_dispatch
remain the recovery paths.
The shared stable-release concurrency group cannot deadlock: the rerun
waits as pending behind this job, and the sequencer only confirms the
new attempt is queued before it exits and frees the group.
Upstream's e2e-desktop-core job provisioned setup-node + uv sync and pointed the
packaged smoke at a checkout .venv. Provision through setup-pm, hand its selected
interpreter to Desktop via HERMES_DESKTOP_PYTHON, and read the packager contract
from electron-builder.config.cjs, where the build config now lives.
The workflow edit trips the review-label gate. The new test is marked
windows_only, so tests-os.yml's '-m windows_only' job on windows-latest
already collects it.
The 32-file Linux lane runs many more child processes than CPU cores. Hosted children remained alive without readiness output and exhausted per-child deadlines across unrelated suites; four concurrent files complete the focused tenancy, terminal model and compaction cases locally without weakening their assertions.
Same runner class as the os-tests Windows row. Seven files run in
parallel, each spawning process trees; measured on the fixed suite the
test step drops from 153-171 s to 124 s and the job from ~5 min to 3.1.
The serve tree-kill test snapshotted process_tree(first.pid) and then
checked that same tree after `taskkill /T /F` - the snapshot is exactly
what taskkill /T kills, so the check could not fail. Windows never
re-parents: a grandchild spawned detached (cmd /c start /b ...) keeps a
dangling ppid and is invisible to both Process.children() and
taskkill /T. Reviewer sabotage (web_server spawns two detached sleepers
before READY) passed while both orphans survived.
Survivors are now every live process created since the spawn whose
HERMES_HOME, cwd or argv points into the test's scratch profile
(owned_processes), with a positive control that the scan sees the
backend itself before the kill. Same check for gateway stop. Cleanup
kills everything the profile owns.
Also: state-db guard matches the holder PID as a whole number, and the
e2e-windows job ends with a scan that fails on any python.exe /
hermes.exe left running.
Adds tests/e2e/core/windows (windows_only + integration; 18 tests, 5 strict-xfail KNOWN
entries for #120504#121150#121015#121114#120205) and an e2e-windows job in
tests-os.yml running it on windows-latest, one pytest process per file, no retries.
route already means "which legs run"; a leg name is the most specific
route. Presets keep their meaning (all, the linux trio, windows-desktop,
macos-desktop, the bundled three); any other value selects legs by name,
so a leg name, a fragment of one, or the job name GitHub shows
("<leg> / e2e") runs just those legs. A route that selects nothing fails
instead of producing a green empty run.
The generator is now the one interpreter of route for source legs: the
linux/windows/macos jobs run when it selected legs for them, replacing
three hand-kept preset lists. Dispatch route becomes a string input (a
choice cannot take a leg name) and reaches the gen step through env, never
interpolated into the script.
A push-triggered canary rebuilt the whole desktop matrix on every main
push. A daily schedule caps automatic canaries at one per 24h; release.py
already exits clean when HEAD carries the last canary tag, so a quiet day
ships nothing. workflow_dispatch fires one on demand, gated to the default
branch so a feature-branch dispatch cannot tag unmerged code.
Every leg installed a release tag and updated to HEAD, so nothing ran
HEAD's installer on an empty machine (where all four 2026-09-23
install.ps1 breaks lived) and nothing exercised the updater we ship
today -- tag legs run the OLD build's updater handing off to HEAD.
The generator appends a HEAD -> NEXT start after the sampled tags, so the
column runs wherever the update legs run (dispatch and stable-release).
NEXT is a reserved update ref: the drivers mint a child of the install
commit that adds one marker file, written to the object store only, which
the local bare clone carries into serve.git.
On Windows the HEAD leg takes every git.exe dir off PATH and installs no
remote get-url shim: Get-PinnedGit returns any git on PATH, so either one
skipped pinned-git staging. launch-from-spec's HEAD observer now uses the
driver's real git so it cannot poll '' forever on that leg.
The published image had no Xvnc/Xfce because nothing set the Dockerfile's
HERMES_BOT_DESKTOP argument, and a hosted instance (unprivileged, no sudo,
sealed /opt/hermes) cannot install at run time. The image layer is the only
delivery path.
- docker.yml: variant axis [slim, desktop]. :latest / :main / :v* stay the
image they are today; :latest-desktop / :main-desktop / :v*-desktop carry
the packages plus Playwright's headed Chromium. Slim owns the build cache
scope; one manifest per variant so a desktop publish failure never skips
slim's :latest.
- Dockerfile / stage2-hook.sh: XDG_RUNTIME_DIR=/tmp/hermes-runtime seeded
0700 as hermes (containers have no logind; the $HOME/.cache fallback was
the shared /opt/data volume), refused when foreign-owned; deterministic
Chromium discovery exporting the headless shell for ordinary browsing.
- bot_desktop: memory gate reads the cgroup working set (usage minus
inactive_file) so it cannot tighten over uptime and refuse to restart a
screen idle-stop just stopped; installable() gives three distinct dead-end
messages instead of a sudo line nobody there can run; env_for_agent
replaces a headless-shell pin so agent and dock share one Chromium.
Squash of IAvecilla/hermes-agent:bot-desktop-cloud-image (#112381, 13
commits), which GitHub auto-closed when its base branch merged as #108914.
Review fixes from pefontana (cache scope, per-variant merge, red browser
test) are included.
Co-authored-by: pefontana <pefontana@users.noreply.github.com>
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
The terminal job only needs the Ink TUI, not every workspace (desktop/electron).
The e2e-upgrade bwrap check used different flags from _bwrap_usable (no --proc,
no --die-with-parent), so it could pass while the suite fell back to running the
real updater unsandboxed; it now asserts _helpers.BWRAP_OK itself.
The six provider keys sat in job-level env, so every step saw them: uv sync
(and any sdist build backend it runs), setup-uv, checkout and the retry
action. The job now carries only secrets.X != '' booleans for the gate, and
the values are set on 'Run live canaries' alone. actionlint clean.
tests/e2e/core/terminal drives the real `hermes --tui` over a PTY, so the e2e
job now installs the Node workspaces and builds ui-tui, and
HERMES_E2E_REQUIRE_TUI=1 makes a missing build fail instead of skip.
tests/e2e/core/upgrade runs a real N-1 -> HEAD `hermes update`: it needs full
history + tags, bubblewrap (every updater runs sandboxed so it can never reach
a real gateway or systemd), the warm uv cache, and up to ~15 min for one file.
It gets its own 60-minute job instead of stretching the e2e job.
Runs tests/e2e/core/live (`-m live`) nightly, on `v*` tags and on
workflow_dispatch (optional -k filter). Secrets-gated on LIVE_*_API_KEY
repo secrets (each case skips without its key; the job no-ops when none
are configured), main-repo only, one run per ref (never cancels a release
gate), 25-minute timeout. Uses direct pytest because scripts/run_tests.sh
starts from `env -i` so no credential can reach a test. Publishes a
usage/cost table to the step summary and uploads junit + usage JSONL.
(cherry picked from commit 75c5656ff3905512bf93c1fd887bbd5e85396f29)
A stalled stream burned the 600 s per-test timeout until the 30-min job
timeout cancelled the lane, and the failure()-only upload then saved no
traces. The CLI --reporter flag also dropped the config's html report.
The onboarding spec is a smoke test; #120005 needs a non-default profile.
The legacy visual Playwright lane stays disabled; the core suite gets its own
reusable workflow (retries 0, one worker, failure-only artifacts) gated on the
same python_prod/frontend classification and counted by All required checks
pass. The legacy config ignores e2e/core so the specs never run twice.
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
The e2e job already discovers tests/e2e/core/delivery/ (C12 messaging
exactly-once, C13 cron virtual-clock soak) through
`run_tests.sh --include-integration tests/e2e`; both need the 900 s
per-file budget under load (C12 150-590 s on a loaded 20-core box).
Cell 5 lives under tests/conformance/ and runs in the unit job.
Conflicts:
- scripts/releases/stamping.py, tests/scripts/test_version_stamping.py:
took ethie/pm-clean. The release branch's side was only its base's copy
of "stamping a payload snapshot skips the bootstrap-installer check"
(8411fdb333, same patch-id as 8d34601f47 here); the install-stamp
refactor c13ea774e6 supersedes the rest.
- tests/ci/test_stable_release_graph.py: kept pm-clean's release-epoch
contract (no HERMES_RELEASE_EPOCH on termux-deb, version on docker and
nix only) and the release branch's per-group receipt wiring.
Semantic conflict: the dispatch log step (6e64e961d8) read
inputs.termux_only, which the jobs input replaced. It reads JOBS now, and a
dispatch that selects only some groups has no release.py replay, as a
termux-only one had none before.
agent/bedrock_adapter.py calls lazy_deps.ensure("provider.bedrock") at
import time. The HERMES_DISABLE_LAZY_INSTALLS kill-switch was only set by
a per-test fixture, so collecting any test module that imports the
adapter ran a real `uv pip install boto3` into the shared CI venv (the
unit job never synced the bedrock extra). test_bedrock_adapter.py raced
it: when the install had not landed yet, its botocore tests skipped and
test_call_converse_replays_thinking_botocore_accepts failed with
"No module named 'botocore'" (FLAKY on this PR's second CI run).
- tests/conftest.py sets the kill-switch at import, before collection.
- tests.yml syncs --extra bedrock with the other lazy-install extras the
suite exercises, so the botocore tests keep running, deterministically.
- The unguarded botocore test importorskips like its siblings.
- Invariant: test_lazy_deps.py asserts the switch is set at collection
(red on origin/main's conftest, green here).
The tests.yml pinned uv 0.9.28, which resolves `uv python install 3.11`
to CPython 3.11.14 linking SQLite 3.50.4. That SQLite has the WAL-reset
bug, so Hermes deliberately falls back to DELETE journal mode and the
unit job never exercised WAL: ~2,600 tests that reach a WAL SessionDB on
a current SQLite ran DELETE, and the 38 requires_wal tests were skipped.
Bump the pin to uv 0.12.13 (resolves 3.11.16 / SQLite 3.53.1) in both
jobs, and add a step that fails the job when the venv's SQLite is
WAL-reset vulnerable, so a later pin change cannot silently revert it.
Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.
- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
raises=AssertionError and RpcError subclasses it, so a gateway crash counted
as the expected failure. Every invariant is now asserted normally; only the
known leftovers (surviving tool tree and the tool_call it leaves without a
result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
the xfail accepts. The DB check used to sit behind the orphan assert and
never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
first prefix break, before usage and integrity. Messages/system prompt,
usage and integrity are asserted first; the tools-array drift is checked
last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
the fake provider's handler thread, where an assert only dropped the
connection. Script records those failures and the test re-raises them after
every turn; steer must land and the interrupted turn must report
interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
Readiness now requires our child's pid from authenticated /health/detailed
and retries on a fresh port when the child reports it in use. The fixture
guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
constructor failed: messages_fts"; the delete arm's busy tolerance keys on
the result code. A failed episode's roles are stopped so the shared chamber
and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
it). The passive update check made a GitHub round-trip from every test
surface, and on a blobless clone whose objects lag upstream its
`git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
call counts and the agent's stale-kill log, so a cross-turn breaker trip
can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
retried into green; own uv cache entry (cache-suffix: e2e).
CI's pinned uv only knows CPython 3.11.14, whose bundled SQLite has the WAL-reset
bug, so Hermes ran state.db in DELETE mode and every torture-chamber episode
skipped (green over zero coverage). The chaos cleanup also called os.pidfd_open,
which that build lacks, turning two strict xfails into errors.
The core E2E suites spawn real processes (serve, gateway, tui_gateway,
MCP servers, concurrent SQLite writers). One subprocess per file keeps
them isolated and parallel instead of one sequential pytest.
installer-tests.yml predates nothing it still owned. Its pytest step
(test_source_launcher_stages.py) is platforms("windows") and already runs in
both tests-os Windows lanes, so every installer PR ran it twice. The
`installer` lane never gated anything on its own either: every path that set
it also sets `python`, which gates tests-os.
The two standalone scripts/tests/*.ps1 suites become one platforms("windows")
pytest file parametrized over Windows PowerShell 5.1 and pwsh 7, so
list_os_marked_tests picks them up with everything else. The `installer` lane
goes away from the classifier, detect-changes, ci.yaml and the
all-checks-pass gate; the classifier contract now pins that install.ps1 and
its suites turn `python` on.
transitions splits into one job per receipt (darwin-arm64, darwin-x64,
win32-bundle), and each packaged install job waits only on its own. The
Mac install arms no longer wait for the other arch or for the smokes.
candidate-manifest moves into stable-release.yml and waits for every
candidate call, so it still runs after every smoke (decision 23). The
smoke results it records are the calls' own results, mapped to the smoke
job names the final manifest requires. publish-bundles and complete read
its digest again, which the per-group split had left unset.
read_manifest resolves its opener per call instead of binding
urllib.request.urlopen as an import-time default, so the process trust
setup applies. The receipt fixtures gain the runner's RUNNER_TEMP and the
baseline's macOS identity.
Review findings against the pm-clean installers, each reproduced first:
- install.sh: `curl | bash` aborted before main under `set -u` (empty
BASH_SOURCE). The entry guard falls back to $0.
- install.sh: setup/gateway read stdin, which under `curl | bash` is the
script itself. They open /dev/tty when a terminal can be opened, and
otherwise skip with guidance.
- Both: any uv on PATH was trusted. uv 0.6.17 has no `python install
--no-bin`. A PATH uv now has to run and be at least the pinned version,
otherwise the pin is staged.
- install.sh: the staged uv went under ~/.hermes/tools even with a custom
--hermes-home. It now goes to pm's store_root() default,
$HERMES_HOME/tools.
- Both: when a stash failed, the script logged "overwritten below" and ran
`reset --hard` anyway. Local work is now parked before checkout, and a
stash failure stops the install.
- Both: reruns ignored an explicit HERMES_REPO_URL. It now repoints origin.
- Both: --commit had no ancestor guard. The pin must be on the installed
branch.
- install.sh: the blobless fallback was `--depth 1 --single-branch`, so a
non-tip --commit could not check out. It now keeps full history with
blobs fetched on demand.
- Both: ported main's recovery for a commit-less .git (moved aside, #40998)
and for an unmerged index (reset -q before the stash, #4735). Stashing
before checkout makes both reachable.
- install.ps1: on Windows PowerShell 5.1, `native 2>$null` / `2>&1` under
Stop turns stderr into a terminating NativeCommandError, verified on a
Win11 host. Every native call now goes through Invoke-Native, which
relaxes the preference only for that call.
- install.ps1: clone publishes from a staging dir, with retries and a
blobless fallback, and refuses a non-empty destination (mirrors
install.sh). UV_NO_CONFIG and `--no-registry` are restored. pwsh 7
HttpRequestException falls back to the mirror, except TLS trust failures.
tar resolves from System32. A literal CR/LF in the desktop failure
message is removed.
Deletes six tests that regex-extracted main's legacy install.sh functions
(node, browsers, PATH block, lockfile churn; pm runs `npm ci` whenever a
lockfile exists). The two behaviours still relevant are covered by new
behavioural tests.
cryptography ships no win_arm64 wheel, so every Windows ARM64 venv sync
compiles it from the sdist and needs MSVC, Clang, Rust and static OpenSSL.
Only setup-hermes.ps1 (and so activate.ps1) prepared that environment,
between a `pm install --tools-only` and the real sync. install.ps1,
`hermes update` and repair ran the same sync without it and failed in
openssl-sys.
PM owns the sync, so PM prepares it. pm/native_build.py holds the adapter
(moved from scripts/build/windows_deps.py) plus source_build_environment(),
which prepares only on win32-arm64 when the synced project carries the
provider script. A payload has prebuilt dependencies and needs no compiler.
VenvPackage.apply and build_environment pass the result to uv children
only. It carries the bridged pip index settings, which managed_environment
applies only to the ambient environment. The state root stays the store
parent, so existing vcpkg/OpenSSL builds are reused.
setup-hermes.ps1 collapses to one `pm install`: the tools-only split existed
only for this preparation, and pm install already puts its tools on PATH
before the venv sync (pm/cli.py activate check).
Not yet verified live on Windows ARM64.
The products stage (source_completion) writes install-stamp.json, and
this lane deliberately runs only prerequisites/repository/config/complete,
so the checkout never has one. The lane now says so with
--no-source-stamp; full-install callers (windows-e2e.ps1) stay strict.
Replayed locally: the four stages leave no stamp, the old invocation
fails exactly like CI, the lane invocation verifies.
Upstream carries only CalVer tags, which version_from_tag rejects on
purpose, so every PR image build died in "Write install stamp" with
"no reachable release tag". The runtime already reads a tagless source
checkout as base "unknown" plus its commit; the image now records the
same: the workflow admits GITHUB_SHA via --commit, adds version args only
when a release is reachable, and write_install_stamp accepts a missing
base version when the caller supplied the commit (a local tree still
stays unstamped).