`release.py release` gains two flags. They can be used together.
--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.
--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.
The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.
The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.
A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:
- The next version was derived from it, so the next cut would reuse the
version. It now takes the newer of the R2 head and the newest
published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
their publication pass, so a bundle-less release would re-advance
every 15 minutes. The head is now the newer of the R2 head and the
published release whose final tag binds the Docker stable alias
digest.
`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.
Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:
- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]
Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
Stable Release Publication ran every 15 minutes (96 runs a day, each
checking out full history, setting up node and buildx, logging into
Docker Hub, and taking the release-signing environment) only because the
sequencer held a failed run for a 15-minute backoff that the failure
event could never satisfy, so the cron was what actually retried.
Drop the backoff: the reconcile pass started by a failed Stable Release
reruns its failed jobs right away. MAX_ATTEMPTS burning, oldest-first
retry ordering, the attempt-entry check, and the needs_retarget repair
stay. The schedule trigger goes; workflow_run and workflow_dispatch
remain the recovery paths.
The shared stable-release concurrency group cannot deadlock: the rerun
waits as pending behind this job, and the sequencer only confirms the
new attempt is queued before it exits and frees the group.
Conflicts:
- scripts/releases/stamping.py, tests/scripts/test_version_stamping.py:
took ethie/pm-clean. The release branch's side was only its base's copy
of "stamping a payload snapshot skips the bootstrap-installer check"
(8411fdb333, same patch-id as 8d34601f47 here); the install-stamp
refactor c13ea774e6 supersedes the rest.
- tests/ci/test_stable_release_graph.py: kept pm-clean's release-epoch
contract (no HERMES_RELEASE_EPOCH on termux-deb, version on docker and
nix only) and the release branch's per-group receipt wiring.
Semantic conflict: the dispatch log step (6e64e961d8) read
inputs.termux_only, which the jobs input replaced. It reads JOBS now, and a
dispatch that selects only some groups has no release.py replay, as a
termux-only one had none before.
installer-tests.yml predates nothing it still owned. Its pytest step
(test_source_launcher_stages.py) is platforms("windows") and already runs in
both tests-os Windows lanes, so every installer PR ran it twice. The
`installer` lane never gated anything on its own either: every path that set
it also sets `python`, which gates tests-os.
The two standalone scripts/tests/*.ps1 suites become one platforms("windows")
pytest file parametrized over Windows PowerShell 5.1 and pwsh 7, so
list_os_marked_tests picks them up with everything else. The `installer` lane
goes away from the classifier, detect-changes, ci.yaml and the
all-checks-pass gate; the classifier contract now pins that install.ps1 and
its suites turn `python` on.
transitions splits into one job per receipt (darwin-arm64, darwin-x64,
win32-bundle), and each packaged install job waits only on its own. The
Mac install arms no longer wait for the other arch or for the smokes.
candidate-manifest moves into stable-release.yml and waits for every
candidate call, so it still runs after every smoke (decision 23). The
smoke results it records are the calls' own results, mapped to the smoke
job names the final manifest requires. publish-bundles and complete read
its digest again, which the per-group split had left unset.
read_manifest resolves its opener per call instead of binding
urllib.request.urlopen as an import-time default, so the process trust
setup applies. The receipt fixtures gain the runner's RUNNER_TEMP and the
baseline's macOS identity.
Stable now calls the desktop bundle workflow once per build group, and
each Mac arch and the Windows bundle assembly stage a <group>-receipt.json
into its attempt archive before the group's smoke runs. The receipts let
the next commit start each install arm from its own group's bytes.
publish-bundles and complete temporarily lose their candidate-manifest
digest source; the next commit moves that job into stable-release.yml and
wires the digest back.
The stable and latest aliases still move at publication. The image tag is
the attempt ref, so an early push never points a client at an unreleased
attempt.
termux_only is replaced by jobs=termux. smoke-win32-universal is removed: the per-arch MSIX smokes cover each arch, and stable's install arms install the msixbundle on both arches.
validate_receipt no longer requires smoke results: receipts are staged right after the bytes and before the smokes run (decision 11), so only the final manifest (validate_candidates) still requires every SMOKE_JOBS result.
The macOS x64 dmg chat in run 35629258153 passed. The job went red
only when actions/upload-artifact got a 403 on FinalizeArtifact. That
one red matrix cell made smoke-darwin a failure, so publish-channel
skipped and the channel head did not move.
Stop screen recording and Upload smoke diagnostics run after the chat
and only collect evidence. continue-on-error keeps a failed stop or a
failed upload from failing the job. The install and chat steps still
fail the job, and publication still requires every smoke job to succeed.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
desktop preparation parses its claim ref and carries the attempt ref as
archive_tag beside the plain payload version. Stable admission accepts
only attempt refs and returns them; accepted candidates, bootstrap and
advance_stable read the attempt-scoped archive, and docker manifests may
carry the attempt-ref image tag while stable/latest aliases stay put
until publish.
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)
Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Admission parses the attempt ref for the version and the attempt, and the
claim metadata must name the same attempt. release has written "attempt"
into every claim since the attempt refs landed; without this, admission
would refuse each of them for an unexpected key.
The attempt stays out of the workflow outputs: emit writes an explicit key
list and no job reads it.
The admit job used the workflow read token. gh release view answers
"release not found" for a draft that exists when the token cannot read
drafts. Give that job contents: write. The other jobs keep their own
permissions, and the workflow default stays read.
Run all release identity cases in one native PowerShell session instead of
charging a cold start and a scheduler-sensitive deadline to every parameter.
Keep the workflow scripts and native exit handling real, with the canonical
per-file runner as the deadlock guard.
Native Windows collection called POSIX-only APIs, and the lane omitted its
Telegram dependency. Several live tests also encoded obsolete runtime identity,
release-channel, and process-launch contracts.
Refresh the tests against the current production seams. Give hosted PowerShell
probes enough time under contention and install the Telegram extra in the OS
lane.
Validated on native Windows 3.14.7:
- 12 passed and 3 skipped across the six files that failed in run 35672259145
Validated on Linux:
- 116 passed and 12 skipped across the same six files
- Ruff, actionlint, YAML parsing, and Windows marker discovery passed
Windows VERSIONINFO and the MSIX package version were two derivations of
one fact, and the sideload one counted minutes since the last stable, which
overflows a 16-bit field after 45 days. Both now come from
storePackageVersionAt's year.hourOfYear.secondOfHour.0, so a later build
always sorts above an earlier one and the cap has nothing left to cap.
scripts/ci/list_os_marked_tests.py matched the literal lane word inside a
platforms() string, so 80 files gated platforms("posix") never reached the
macOS lane ("posix = linux or macOS" was false for CI), and "any" files
reached none. The selector now resolves specs the way the conftest gate
does (posix ⊇ linux+macos, any ⊇ all, "not X" admits the rest) and only
looks inside mark.platforms(...) calls, dropping a false positive whose
only "macos" was inside a generated-file string.
scripts/run_tests_parallel.py still grepped the retired linux_only/
macos_only/windows_only names, so its "N files SKIPPED on this host" note
had been silent since the migration. It now shares the selector's
resolver and names the spec and the lane(s) it runs on.
Restore the platforms("windows") mark that
test_suppress_platform_ver_console_stubs_syscmd_ver lost in the
windows_only migration (its docstring still declared it); it passed
vacuously on Linux and was deselected on the Windows lane.
batch-sign-binaries.mjs and electron-builder.config.cjs::windowsSigning
treat a missing AZURE_SIGN_* set as "sign nothing, warn" — right for a
fork or a local build, but a release-signing lane with the vars
unprovisioned would have published UNSIGNED installers with only a
console.warn in the log. The macOS leg already refuses to build without
CSC/Apple credentials under the publishing gate; the Windows legs now
do the same for the Azure vars, before any payload is built.
archive-inputs (push to main on pm/lock.json), the nightly canary R2
prune and termux-verify's bionic runtime job all called
r2.credentials(), which exits 2 on a missing CLOUDFLARE_R2_* env. On a
repo without the release-signing environment provisioned, merging this
branch would turn main red on the first push and the nightly red daily.
A job-level `if` cannot read `secrets`, so the gates use the documented
shapes: single-step consumers expose the secret through the job env and
test `env.CLOUDFLARE_R2_ACCOUNT_ID` in the step `if`; the multi-step
termux job is gated by a job that outputs a provisioned flag. A stable
release candidate (inputs.release) still runs and fails loudly.
termux-verify also triggered on a personal branch (ethie/cli-bundles)
and a test glob that no longer exists; it now runs on main pushes and
PRs touching the Termux paths.
The 6-target cold->warm smoke existed to prove setup-pm still cold-boots
when its code changes without a lock bump, and to reach linux-arm64,
darwin-x64 and win32-arm64. Both are now covered without a dedicated
workflow:
* The tools cache key hashes pm/**, the action itself and
scripts/ci/setup_toolchain.py, not only pm/lock.json. A provisioning
change misses the cache on every lane that uses the action, so the
cold path runs where the tests already are.
* tests-os gains a windows-11-arm leg running the same windows-marked
files (tests/pm carries ten of them). It needs the ARM64 build deps
because several extras build from sdist, and fewer workers on the
4-core runner.
The Windows SDK adapter test is the one thing left that no other lane
ran natively; it keeps its two Windows runners under
windows-bundle-sdk.yml, path-triggered on the signing scripts. The
run-scoped cache cleanup workflow and its script only served the smoke
and go with it.
Four origin/main merges brought back `linux_only` / `macos_only` /
`windows_only` marks in 41 test files, along with the pre-platforms()
versions of scripts/ci/list_os_marked_tests.py and check_os_marker_fakes.py.
Because the legacy names are no longer registered, pytest treated them as
unknown marks — a warning — so every Windows- or macOS-only test RAN on
Linux (test_local_runtime_recovery.py tripped the live-system kill guard).
Rewrite the marks, restore the platforms()-aware CI scripts (keeping main's
os.walk fix for vanishing __pycache__ dirs), drop the stale _BASELINE entries,
and make the conftest reject the retired marks outright so the next merge
cannot resurrect them silently.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
A channel dispatch sends both `channel` and `build_commit` (the commit is
provenance, not a mode), so the tag/commit staging steps — gated on
`inputs.build_commit != ''` — ran inside a channel build. There a pinned
request forces HERMES_PAYLOAD_TAG and HERMES_BUILD_COMMIT empty, so the step
fell through to tag mode and handoff refused the empty tag:
ValueError: Invalid release handoff identity (scripts/releases/handoff.py:24)
Channel packages are already staged by the pinned-request step, so gate both
per-platform staging steps on the channel signal the rest of the workflow
uses. The commit-only status page had the same flaw one job later: it would
fetch commit-namespace receipts a channel build never wrote and publish a
false "not built" matrix under releases/commit/<sha>/. The channel flow
renders its own page in publish-channel.
A self-need or an unknown dependency is rejected by the platform parser, so
the test only pins what GitHub does not fail closed on: a result read from a
job the reader never declared. That is the shape the desktop release hit —
`needs.build-win32-release.result` with only `build-win32` in the list.
Dropping archive-inputs from the result jobs replaced the whole `needs`
list with the job's own name: `build-win32` needed `build-win32`, and
`build-darwin` needed `build-darwin`. GitHub rejects a workflow with a
cyclic job dependency outright, so the dispatch died at parse time — and
the result jobs read `needs.build-win32-release.result` /
`needs.build-win32-commit.result` (the darwin pair likewise), which are
only reachable through a declared `needs` entry.
Name the trust-branch legs instead, and pin the contract: no job may
need itself, and every `needs.<job>` read in an `if:` or expression must
be a declared dependency.
The archive-inputs reusable-workflow job existed only to serialize one
command behind validate. Run scripts.ci.archive_inputs directly inside
validate with the job's own gate shape (skipped for disposable and
termux-only runs), so the archiving never gates or skips the native
build legs and the whole dispatch loses one job-level hop.
The termux-main pool deletes a package's previous archive when it rebuilds, so
the runtime-lib pin table and the bionic lock rows rot without warning. The last
rotation broke a build on eight rows at once, and the stager's concurrent
downloads only surfaced whichever 404 won the race.
pm now owns the pin table it repairs: scripts/termux/runtime_libs.json moves to
pm/termux_runtime_libs.json, so pins live in pm/ and scripts consume them — the
direction scripts/ci/archive_inputs.py already reads pm/lock.json in.
`hermes pm update --termux` repins exactly the rows whose archive the pool has
replaced, hashing each replacement against the index SHA256 before writing
url/version/hash together. `--check` reports without writing and exits 1, so a
retired pin can fail a cheap preflight instead of a payload build.
It is a repair, not an update: an alive pin is never moved, because a repin can
land a rebuilt library under a moved soname and the table is the payload's
recursive DT_NEEDED closure. A pin whose package the pool has dropped outright
is reported and left alone. `--termux` runs alone — names/--target/--uv/--npm
are ignored, since repairing foreign-target pins is not a version resolution.
Verified: `pm update --termux --check` against the live pool reports 89 rows
served; a table deliberately pinned to the retired libiconv 1.18-1 repins to
1.19 with the pool's hash through the real network path; 19 new tests; the
tests/pm, tests/ci and tests/scripts suites have the same failure set as the
base commit (91 pre-existing Windows environment failures, none new).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
With the macos_only marker the win32 skipif is redundant (the marker already
skips every non-darwin host) and the is_macos() fake in _wire contradicts the
repo rule that host-specific behaviour is tested on that host, not by making
the interpreter believe it is elsewhere. Drop both.
_get_service_pids(all_profiles=True) also runs a bare `launchctl list` prefix
scan, so on a macOS box with a live ai.hermes.gateway* fleet the scoping
assertions picked up real PIDs (the one failure the reporter could clear by
booting the gateway out). Stub subprocess.run in _wire so the tests assert on
routing, not on the developer machine.
The cherry-picked tests/ci change-detector (asserting one specific file is in
the macos_only list) is dropped: the selector test suite already covers the
mechanism, and the marker is now the file's declaration.