installer-tests.yml predates nothing it still owned. Its pytest step
(test_source_launcher_stages.py) is platforms("windows") and already runs in
both tests-os Windows lanes, so every installer PR ran it twice. The
`installer` lane never gated anything on its own either: every path that set
it also sets `python`, which gates tests-os.
The two standalone scripts/tests/*.ps1 suites become one platforms("windows")
pytest file parametrized over Windows PowerShell 5.1 and pwsh 7, so
list_os_marked_tests picks them up with everything else. The `installer` lane
goes away from the classifier, detect-changes, ci.yaml and the
all-checks-pass gate; the classifier contract now pins that install.ps1 and
its suites turn `python` on.
The 6-target cold->warm smoke existed to prove setup-pm still cold-boots
when its code changes without a lock bump, and to reach linux-arm64,
darwin-x64 and win32-arm64. Both are now covered without a dedicated
workflow:
* The tools cache key hashes pm/**, the action itself and
scripts/ci/setup_toolchain.py, not only pm/lock.json. A provisioning
change misses the cache on every lane that uses the action, so the
cold path runs where the tests already are.
* tests-os gains a windows-11-arm leg running the same windows-marked
files (tests/pm carries ten of them). It needs the ARM64 build deps
because several extras build from sdist, and fewer workers on the
4-core runner.
The Windows SDK adapter test is the one thing left that no other lane
ran natively; it keeps its two Windows runners under
windows-bundle-sdk.yml, path-triggered on the signing scripts. The
run-scoped cache cleanup workflow and its script only served the smoke
and go with it.
The runner maps a hyphenated input to INPUT_LOCK-SOURCE, not INPUT_LOCK_SOURCE.
The registration step read the underscored name, saw nothing, and threw
"invalid lockSource" from every job that composed setup-pm, so the whole CI
run failed before any consumer ran. The unit test used the same wrong
spelling, which is why it stayed green; it now drives the real mapping.
The CI cache barely ever restored anything useful, for four independent
reasons found in live run logs and the fork's cache store:
- `uv cache prune --ci` (setup-pm/prune) discards downloaded wheels before
the save, so the snapshot carried only source-built wheels — the next
run "restored" it and still cold-downloaded everything. Upstream main's
logs showed the end state: a 209KB stub cache exact-hitting forever.
- Tool-only jobs (icons-freshness etc.) auto-saved a 1.5KB empty uv cache
under the production exact key; caches are immutable, so the stub won
forever and blocked real saves.
- The npm cache key had no restore-keys, so one lockfile bump missed the
exact key and every npm job went cold.
- `2efa4ff94f` deleted the electron-builder toolchain save but left the
assemble job restoring `eb2-` — a fossil nobody regenerates; assembly
cold-downloads winCodeSign/ATS/dotnet on every release.
Changes:
- pm/cache_lock.py: move prune_uv_cache_to_lock out of
scripts/bundles/native.py (which re-exports it); the lock-exactness
contract now serves both the bundle ship gate and CI caches.
- pm.build_env learns --exact-lock --lock-source: prune cache entries the
project uv.lock cannot resolve, keeping lock-required downloaded wheels
(unlike --ci). Refuses --ci/--prune-cache combinations.
- setup-pm: python-cache auto-save now requires extras (no stubs from
tool-only jobs); key drops the prune flag and bumps to v3 — pruned and
unpruned saves share one namespace since both are lock-exact; pre-save
pruning switched from --ci to --exact-lock; npm cache gains a
lockfile-agnostic restore prefix.
- save-pm-cache: same exact-lock prune before explicit saves.
- desktop-bundled-release: build legs (cache-mode: write) restore+save the
electron-builder toolchain under eb3- keyed on the locked builder
version; assemble restores the same namespace; the dead default-cache
resolution step is removed (assembly resolves no electron artifacts).
- cleanup_pm_toolchain_caches.py: match v2 and v3 smoke keys.
Validation: tests/scripts/test_bundle_native.py 11/11 (incl. both
lock-prune gates), tests/pm failures identical before/after the diff,
tests/scripts/test_bundle_payload.py 5/5, tests/ci cleanup 3/3,
tests-js setup-pm-post 1/1 and the three setup-pm-cache contract tests
updated and green (6 failures in that file pre-date this diff and are
drift between the workflow and its stale assertions); pm.build_env
--exact-lock E2E against a real uv cache copy pruned 3 stale entries and
kept the rest.
The e2e-screen-record action installed ffmpeg through three different
OS package managers (apt, brew, winget). winget is the flaky leg on
windows-11-arm and serves an x64-gyan build that runs emulated on ARM;
choco's community package wraps the same gyan x64 zip, so a choco
fallback would not fix either problem. PM already ships a locked,
sha256-pinned ffmpeg (martin-riedl posix, BtbN win32 including a NATIVE
winarm64 build), so the recording action now verifies ffmpeg on PATH
instead of installing it, and the pinned binary rides the same
tools-cache as node/python.
- setup-pm learns a `packages` input (extra PM tools beyond the
toolchain roots, e.g. ffmpeg) threaded through setup_toolchain.py's
prepare/install/archive-inputs phases; the tools-cache key gains an
extra-packages fragment so existing keys stay byte-identical.
- The six chat-driver jobs pass `packages: ffmpeg` to setup-pm.
- e2e-screen-record drops the apt/brew/winget install steps, the
ffmpeg actions/cache steps and the save-cache input; Xvfb (headless
linux) and the macOS replayd-approval hack stay.
- The "Install locked chat driver dependencies" step installs only the
tests-js workspace with --omit=dev instead of the whole apps/desktop
tree: the drivers need @playwright/test, zod (previously a phantom
hoisted from @assistant-ui/react), js-yaml and semver only. 64 pure
packages in ~2s vs ~1800 including electron-builder and native
builds; the tree-shaken node_modules runs the real driver modules
(verified by importing desktop-chat-smoke.ts and update-window-chat
end to end).
- tests pin the new contracts; tests/install/README.md updated.
Dependency acquisition during packaging left native wheels and packager
inputs outside the pre-build cache save. Compose PM and existing providers
into a preparation phase, then require builds to consume admitted inputs.
Share native preparation with PM Bundle. Keep path-bound environments and
signing outputs separate from reusable caches. Use read-only cache tokens
for commit builds and preserve the one-command local build path.
Verify pinned tools through PM, probe PTYs under the prepared Electron,
and supply dmgbuild through a build-only PM package. Resolve bundled tool
stores from their payload manifest so relocation preserves discovery.
Validation: focused Python and JS tests, checkJs, Ruff, Windows checks,
anti-slop, cache relocation, and network-denied Linux AppImage builds.
Relocated runtime smoke passed with NixOS host libraries supplied.
Native Windows/macOS signing and live GitHub cache behavior remain untested.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Merge ethie/shared-product-builders with the CI dependency cache and native Windows setup work. Preserve UTF-8 diagnostics in the shared Python environment runner. Pass a persistent cache through isolated native staging and PM-runtime construction. Reuse one Windows prerequisite installer from source setup, native adapters, and CI, preserving Rust homes across HOME isolation.
Verified 85 targeted Python tests (5 host skips), 18 JavaScript tests, workflow validation, and scoped lint/typecheck. On native Windows ARM64, five prerequisite contracts passed and the actual shared provider reused OpenSSL, compiled its header with MSVC, and retained Rust under isolated HOME. Full signed distribution builds and live Actions cache transfer remain CI verification.
Termux removes old package files, so a pinned URL and hash do not keep
build inputs available. Preserve the exact bytes without changing pins.
Archive every PM HTTP artifact and the Termux runtime inputs by SHA256.
CI reads R2 first. Only a missing object permits an upstream download,
hash verification, immutable upload, and verified readback. Seed the
actual toolchain and payload stores before their consumers run.
Use the public archive as a pinned fallback in PM, bootstrap installers,
and Nix fetchers. Keep network retries bounded and report attempted URLs.
Keep publication credentials in protected CI jobs, not installed clients.
Verification:
- 283 targeted tests passed; five POSIX tests skipped on Windows.
- All 87 preserved Termux packages passed local archive miss/hit checks.
- Native ARM64 ripgrep installed through the mirror and ran successfully.
- Wheel import, workflow lint, Python lint, shell syntax, and pins passed.
Live R2 publication, POSIX tests, and Nix builds remain for native CI.
The real-byte archive checks used loopback HTTP, not the live bucket.
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.
Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.
Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
Keep downloads bound to their remote representation and publish through
atomic destination-local staging. Serialize shared partial ownership.
Keep explicit CA trust scoped to provider probes. Preserve checkpoint
history and edited files, validate all profile inputs before dependency
publication, and separate data removal from installed runtime ownership.
Exclude machine-specific PM state from portable transfers. Keep plugin
files and nested skill tools intact. Preserve native test isolation.
Focused native Windows receipts cover the individual repairs and their
integration. This commit does not claim a full-suite or release build.
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
Preserve the PM runtime-repair module boundary. Port the incoming stderr-streaming fix without restoring the deleted managed_uv downloader. Targeted runtime/progress tests and root JS checks passed; desktop typechecks passed.
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.
The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.
e2e-screen-record comment: hhttps -> https.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).
- windows.ps1: readiness handshake after BeginInvoke — one /progress
round-trip must succeed (≤15s) before the server is returned; on failure
tear the listener down and continue without UI. The URL now means
"serving", not "bound". Also fixes the browser opening to a page that never
loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
spawn the real PowerShell script; the Windows-only job now runs them only
when scripts/desktop-update/**, the Electron updater launcher, conftest,
pyproject, or those tests change (push/dispatch fail open). A PR that
never touched that surface cannot be failed by its process timing.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.
Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.
Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.
The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
The composite action .github/actions/e2e-screen-record owns setup and
lifecycle on all three OSes: ffmpeg via apt/brew-verify/winget+cache,
capture via x11grab/gdigrab/avfoundation, mkv at 15fps stopped by 'q'
on live stdin with kill fallback. Linux runners have no display, so
start brings up a dedicated Xvfb :99 and exports DISPLAY - one display
serves both the recorder and any app a later step launches.
Recording moves out of the GUI driver into workflow infrastructure -
that is what makes it uniform - and a missing ffmpeg or a zero-frame
file now FAILS the leg instead of skipping silently: the graceful-skip
path is how the windows leg shipped no recording.mkv while green.
Lifecycle proven locally: start against lavfi testsrc, q-stop, ffprobe
duration check (record-start.sh/record-stop.sh under nix ffmpeg).
scripts/tests/ has held three PowerShell suites that no workflow ever
invoked -- there is no Windows runner in CI, so they have been inert since
they landed. A regression test nothing executes is worse than none: it
reads as coverage.
Adds a windows-latest job, gated on a new `installer` lane so it only fires
for PRs touching install.ps1 or its tests. The 8.3 suite runs under both
pwsh 7 and Windows PowerShell 5.1, since install.ps1 arrives via `irm | iex`
into whichever shell the user already has and 5.1 is what ships with Windows.
Only the 8.3 suite is wired up. The other two fail on main today for
unrelated reasons; they can join once they are fixed.
After the test-suite prune, the Python slices (~2.3m each) are no longer
CI's critical path — Desktop E2E (5.2m, the longest job) and the Docker
build are, and both run on every python-lane PR even when the diff never
leaves tests/. Neither consumes the test suite: Playwright drives the
built app + hermes serve backend, and the image copies installed code.
New python_prod lane = python minus tests-only diffs. e2e-desktop and
docker gate on it; every pytest/lint lane keeps gating on python.
Fail-open contract preserved: .github/ changes and empty diffs set
python_prod=true, and runner infrastructure (scripts/run_tests.sh,
run_tests_parallel.py) is deliberately NOT tests-only since a bad
runner edit can mask real failures.
Replay over the last 231 main commits: 39 (17%) would skip both jobs,
cutting their critical path from ~8m to ~3m. E2E-verified through the
real script entrypoint (tests-only/prod/mixed/fail-open) + 83 tests/ci
green.
* fix(ci): publish inline E2E evidence
Upload bounded screenshot evidence from E2E, then publish validated images
from a trusted workflow_run job to commit-pinned branches in the evidence repo.
Wait briefly for the live CI review comment marker before publishing, so
GitHub's read-after-write delay cannot leave an orphaned evidence branch.
* fix(ci): isolate privileged credentials from PR jobs
Keep App private keys and Docker Hub credentials out of PR-controlled
workflows. Use protected environments for trusted publishing and a public
repository variable for the App client ID.
* fix(ci): attach E2E evidence with restricted bot session
Replace the App-backed evidence repository publisher with gh-image uploads
from a dedicated bot session in the gh-image environment.
* fix(ci): publish validated E2E evidence from forks
Let the trusted default-branch publisher handle bounded, validated evidence
artifacts from fork PR CI without checking out or executing fork code.
* ci: surface E2E screenshots in review comment
* ci: mark completed review commits in past tense
* ci: surface approved sensitive-file reviews
* ci: link sensitive files to reviewed changes
* ci: stage desktop E2E visual evidence
Track screenshots newly introduced against main and package visual diffs for a trusted publisher.
* fix(ci): pass E2E evidence output paths
Supply the manifest and staging-directory arguments required by the screenshot status helper.
* fix(ci): download the OSV SARIF artifact
Match the artifact name and result filename emitted by the pinned upstream reusable workflow.
- scripts/validate_plugin_catalog.py: standalone stdlib+pyyaml structural
validator for plugin-catalog entries and removed.yaml (no hermes install
needed; runtime twin of hermes_cli/plugin_catalog.py). --json support,
unknown top-level keys warn instead of failing for forward compat.
- .github/actions/plugin-validate: composite action plugin authors drop
into their own repo's CI — installs hermes-agent from a chosen ref and
runs 'hermes plugins validate <path>'.
- .github/workflows/plugin-catalog-ci.yml: admission gate on PRs touching
plugin-catalog/** — structural job plus pinned-source job that clones
each changed entry's repo, hard-fails on unreachable pinned sha
(supply-chain gate), and validates the plugin at that exact commit.
Composite actions cannot access the secrets context — the runner's
template engine rejects secrets.* references at load time with
'Unrecognized named-value: secrets'.
Move APP_ID and APP_PRIVATE_KEY from direct secrets.* references inside
the composite action to inputs passed by each calling workflow. The
fallback logic (GITHUB_TOKEN when APP_ID is empty, for fork PRs) stays
in the composite action's check step.
Replace the long-lived fine-grained PAT (AUTOFIX_BOT_PAT) with short-lived
(1-hour) installation access tokens minted via a new get-app-token composite
action wrapping actions/create-github-app-token@v3.2.0.
The PAT was used in 13 spots across 8 workflow files for gh CLI / GitHub API
calls. The per-repo GITHUB_TOKEN (1,000 req/hr) was getting rate-limited when
multiple workflows fire concurrently (deploy-site, skills-index, ci-timings,
supply-chain-audit, js-autofix). App installation tokens get 5,000 req/hr
per installation and are scoped to the App's permissions, not a user account.
New composite action: .github/actions/get-app-token/
- Wraps actions/create-github-app-token@bcd2ba49 (v3.2.0, SHA-pinned)
- Reads APP_ID + APP_PRIVATE_KEY repo secrets
- Outputs a 1hr installation token via steps.app-token.outputs.token
Requires two new repo secrets (set after creating the GitHub App):
- APP_ID: the App's numeric ID
- APP_PRIVATE_KEY: the PEM private key
App installation permissions needed:
contents: write (js-autofix push, pypi release upload)
pull-requests: write (js-autofix PR create/merge, supply-chain comment)
issues: write (skills-index-freshness issue creation)
actions: write (skills-index workflow trigger)
workflows: write (skills-index triggers deploy-site.yml)
The AUTOFIX_BOT_PAT secret can be deleted once CI passes on this PR.
The comment in js-autofix.yml noting that PAT pushes trigger downstream
workflows is updated — App tokens have the same property (they are not
GITHUB_TOKEN), so the concurrency-cancel loop logic is unchanged.
A PR more than 100 commits ahead of its merge-base paginates the compare
endpoint; pages after the first carry files: null, so the bare .files[]
made jq die with 'cannot iterate over: null' on every retry and forced
the classifier's fail-open path (all lanes on, ci_review gate demanding
a label with zero CI-sensitive files in the diff). .files[]? keeps the
page-one file list and ignores the null tail.
#66373 swapped GITHUB_TOKEN -> AUTOFIX_BOT_PAT across the workflows. That
PAT is empty on fork PRs (forks get no repo secrets), which broke every
fork PR two ways:
1. detect-changes classified with the empty PAT -> the compare API failed
all 3 retries -> the classifier failed open and force-enabled the
ci_review lane on EVERY fork PR.
2. The ci-reviewed / mcp-catalog-reviewed label gates then read labels with
the same empty PAT via a hard-failing retry step -> the job failed with
no recovery a fork contributor could perform (they can't self-add the
label; re-running can't fix it).
Restores the pre-#66373 fork-safe behavior without reverting the commit's
real improvements (job timeouts, per-file flake retry, network-install
retries):
- detect-changes + ci.yml: token falls back to the built-in read-only
github.token when AUTOFIX_BOT_PAT is empty. On main it uses the PAT
(authoritative); on forks it uses github.token, which can read the
public compare endpoint. (An input `default:` only applies on omission,
not on an empty passed value — hence the explicit `|| github.token`.)
- lint ci-review + supply-chain mcp-catalog gates: restore the inline
`gh pr view ... || true` label read with the github.token fallback,
dropping the hard-failing retry "Fetch PR labels" step. Graceful
degrade to "label absent" on an API blip, same as before #66373.
Same-repo enforcement is unchanged (byte-identical logic; the PAT is still
used there). Fork PRs classify correctly and the gates read labels via the
read-only token exactly as they did before the regression.
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory
The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.
New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.
- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
merged with the directory at import time (directory wins). All
existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
conflicting reassignments (incl. against the legacy map), validates
email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
a legacy entry; failure message prints the exact add_contributor
command. Also auto-resolves bare <login>@users.noreply.github.com
emails is intentionally NOT added (kept id+login form only, matching
previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
merge precedence, CLI idempotency/conflict/validation, subprocess E2E.
* feat(ci): one-shot per-file flake retry in the parallel test runner
A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.
- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
deterministic failure still exits 1; retries=0 restores old behavior.
This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.
* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s
These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.
* fix(ci): job timeouts everywhere + retries on all network installs
Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
that lacked it: pip installs (deploy-site, skills-index), npm ci
(deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
test deps). Deterministic build steps (npm run build) deliberately
NOT retried — split into separate steps so a real build failure fails
fast instead of retrying 3x.
* docs(agents): document the file-retry flake policy
* fix(ci): curl retries on deploy hook + skills-index probe
* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile
From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
run_id-suffixed keys — the cache never matched once, so LPT slicing
always ran blind and unbalanced slices pushed heavy files toward the
per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
'gh pr view || true' turned an API blip into 'label absent' → false
BLOCKING failure. Now 3x retry, and API failure is reported as an API
failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
curl --retry 3 (ADD cannot retry; checksums still enforced); npm
--fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
get continue-on-error so an artifact-service blip can't fail a green
test slice.
* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list
- test_tui_gateway_server.py: session.create / non-eager session.resume
arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
test and fires into the NEXT test's _make_agent mock, racily
corrupting captured state (the recurring session_resume shard
failures). Replaced the per-test whack-a-mole stub with a module-wide
autouse fixture; the 3 worker-lifecycle tests that genuinely need the
deferred build opt back in via @pytest.mark.real_agent_prewarm (new
marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
live PROVIDER_REGISTRY instead of a hand-list that had drifted
(missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
tests failed on any machine with HF_TOKEN exported. E2E-verified with
HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.
* test: de-flake 30 timing-sensitive test files for loaded CI runners
Root-cause fixes from the flake audit (session-DB mining + repo sweep):
Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
unbounded blocking read (parent wedge now fails THIS test with a clear
message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
(the 1s partition window mid-interpreter-startup is how a child PID
escaped the live-system guard in CI)
Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
voice_cli_integration, docker_environment, session_store_lock_io,
planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
(joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
compression fork-lock TTL 1s->3s (12 refresh chances per lease);
compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)
* fix(tests): repair indentation from de-flake batch edit
* fix(tests): harden env isolation and replace remaining sleep-sync races
The full 42k-test run and complete npm check surfaced three more classes:
- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
leaked into Python/TUI tests. Pin the default Honcho host in the
hermetic fixture, isolate the one fallback test from ~/.honcho, and
blank SSH_* around terminalSetup tests. This flipped 20 false failures
back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
time.sleep globally, then busy-polled with that same mocked sleep. Under
full-suite load the poller could starve the writer. Each test now waits
on an Event emitted by the exact flush/retry transition; 30/30 passed
under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
not fire before its assertion. A loaded runner descheduled the test for
>500ms and both chunks arrived. Producer controls now gate second-chunk
and completion transitions explicitly.
Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.
* refactor(ci): use gh bot pat, better retries
refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.
Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference
Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.
ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth
Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.
19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call
---------
Co-authored-by: ethernet <arilotter@gmail.com>
git diff on a lockfile is unreadable: npm reorders entries, rewrites
integrity hashes, and moves packages between nesting levels, so a
one-line package.json bump produces a thousand-line textual diff.
scripts/ci/lockfile_diff.py instead parses the `packages` map out of
both versions of every tracked package-lock.json (via `git show`),
reduces each to {install path: version}, and set-diffs the maps —
reorder/hash churn vanishes, leaving only actual version movement
(added / removed / updated, with nested dedup copies tracked
separately).
The lockfile-diff workflow posts the result as a Markdown table in a
PR comment gated behind a hidden marker: subsequent pushes PATCH the
existing comment instead of stacking new ones, and a push that reverts
all lockfile changes updates the comment to say so. Advisory only —
never fails on findings; fork PRs (read-only token) degrade to a
warning.
Wired through the ci.yml orchestrator with a new npm_lock lane in
classify_changes.py (fails open on .github/ changes per the existing
contract).
ci: centralize path-gating behind single orchestrator + all-checks-pass
gate
Replace the scattered per-workflow detect-changes pattern with a single
ci.yml orchestrator that runs the classifier once, then conditionally
calls sub-workflows via workflow_call based on lane outputs. A final
all-checks-pass job (if: always()) aggregates all results so branch
protection only needs to require one check.
Changes:
- New .github/workflows/ci.yml orchestrator (detect + conditional calls
+ all-checks-pass gate)
- Extend classify_changes.py with scan/deps/mcp_catalog lanes, absorbing
supply-chain-audit's internal changes job
- Update detect-changes/action.yml to expose the new lane outputs
- Convert all 10 PR-gated sub-workflows to workflow_call-only triggers,
removing their push/pull_request triggers and per-step detect-changes
guards (gating now happens at the orchestrator level)
- lint.yml + supply-chain-audit.yml receive event_name as a
workflow_call
input to replace github.event_name (which is "workflow_call" inside
called workflows)
- supply-chain-audit.yml: remove internal changes job + *-gate jobs
(orchestrator handles gating, booleans arrive as inputs)
- contributor-check.yml: remove internal filter step
- Update test_classify_changes.py for 6-lane output + new supply-chain
test cases
`npm ci` / `uv sync` / toolchain header fetches occasionally die on
transient network blips — e.g. node-pty's node-gyp fetching Node headers
(an undici assert) during the typecheck job's `npm ci`, which killed the job
before `tsc` ever ran. "Re-run and it goes green" is exactly what CI should
do itself.
- New reusable `.github/actions/retry` composite action wraps a command and
retries on failure (3x / 10s, command passed via env so it can't inject).
Applied to every PR-path network install: npm ci (typecheck, desktop
build, docs site), uv sync (tests, e2e), uv tool install (lint),
pip install (docs site).
- typecheck now runs `npm ci --ignore-scripts`: `tsc` needs only sources +
type defs, so skipping install scripts drops node-pty's native rebuild
(whose header fetch was the flake) and is faster. Validated locally — tsc
passes for ui-tui, apps/shared, and apps/desktop with scripts skipped.
- ripgrep download uses `curl --retry`.
Docker (main-only) and the release/windows workflows are intentionally left
for a follow-up.