Commit Graph

39 Commits

Author SHA1 Message Date
ethernet
dc11e3b3bc feat(release): add --skip-bundles and --skip-tests to stable releases
`release.py release` gains two flags. They can be used together.

--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.

--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.

The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.

The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.

A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:

- The next version was derived from it, so the next cut would reuse the
  version. It now takes the newer of the R2 head and the newest
  published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
  tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
  their publication pass, so a bundle-less release would re-advance
  every 15 minutes. The head is now the newer of the R2 head and the
  published release whose final tag binds the Docker stable alias
  digest.

`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.

Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:

- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]

Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
2026-09-24 13:31:33 -04:00
ethernet
66d54ebf51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/docker.yml
#	Dockerfile
#	apps/desktop/src/app/settings/about-settings.tsx
#	docker/stage2-hook.sh
2026-09-24 02:23:44 -04:00
IAvecilla
686c34d3f6 feat(bot_desktop): ship Bot Screen on hosted images (-desktop tags)
The published image had no Xvnc/Xfce because nothing set the Dockerfile's
HERMES_BOT_DESKTOP argument, and a hosted instance (unprivileged, no sudo,
sealed /opt/hermes) cannot install at run time. The image layer is the only
delivery path.

- docker.yml: variant axis [slim, desktop]. :latest / :main / :v* stay the
  image they are today; :latest-desktop / :main-desktop / :v*-desktop carry
  the packages plus Playwright's headed Chromium. Slim owns the build cache
  scope; one manifest per variant so a desktop publish failure never skips
  slim's :latest.
- Dockerfile / stage2-hook.sh: XDG_RUNTIME_DIR=/tmp/hermes-runtime seeded
  0700 as hermes (containers have no logind; the $HOME/.cache fallback was
  the shared /opt/data volume), refused when foreign-owned; deterministic
  Chromium discovery exporting the headless shell for ordinary browsing.
- bot_desktop: memory gate reads the cgroup working set (usage minus
  inactive_file) so it cannot tighten over uptime and refuse to restart a
  screen idle-stop just stopped; installable() gives three distinct dead-end
  messages instead of a sudo line nobody there can run; env_for_agent
  replaces a headless-shell pin so agent and dock share one Chromium.

Squash of IAvecilla/hermes-agent:bot-desktop-cloud-image (#112381, 13
commits), which GitHub auto-closed when its base branch merged as #108914.
Review fixes from pefontana (cache scope, per-variant merge, red browser
test) are included.

Co-authored-by: pefontana <pefontana@users.noreply.github.com>
2026-09-23 19:54:47 -07:00
ethernet
babbec1c4c feat(pm): isolate developer test environment from runtime extras 2026-09-23 18:13:38 -04:00
ethernet
86f42bf3b6 fix(docker): stamp a commit-only identity when no release tag is reachable
Upstream carries only CalVer tags, which version_from_tag rejects on
purpose, so every PR image build died in "Write install stamp" with
"no reachable release tag". The runtime already reads a tagless source
checkout as base "unknown" plus its commit; the image now records the
same: the workflow admits GITHUB_SHA via --commit, adds version args only
when a release is reachable, and write_install_stamp accepts a missing
base version when the caller supplied the commit (a local tree still
stays unstamped).
2026-09-23 15:44:18 -04:00
ethernet
c13ea774e6 refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.

Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).

Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
2026-09-23 11:41:01 -04:00
ethernet
7e9d4239c6 fix(release): preserve reachable version identity in CI 2026-09-22 12:00:06 -04:00
ethernet
dcd678d6fd fix(release): close publication custody gaps 2026-09-22 02:19:38 -04:00
ethernet
95e03fd701 fix(release): preserve stable ordering through publication 2026-09-21 23:36:19 -04:00
ethernet
dd72571178 feat(release): finalize stable claims into versioned builds 2026-09-21 22:39:16 -04:00
ethernet
d29da5fe20 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
2026-09-19 13:45:07 -04:00
ethernet
482f0be706 fix(ci): docker test job removes the image install stamp before setup-pm
write_install_stamp.py marks the checkout distribution=docker for the image
build; left in place, PM on the runner looks for the image's packaged runtime
and fails with "packaged PM runtime is missing".
2026-09-19 02:01:48 -04:00
ethernet
b676997d2d ci: provision locked Python and Node toolchains through PM 2026-09-08 12:45:15 -04:00
ethernet
7bb62782cc merge: integrate ethie/py314 into ethie/pm-clean
Bring in the Python 3.14 runtime pins and wake-engine changes while
preserving the staged stable-release gate and review fixes.

The merge has no conflicts. Targeted tests on the existing Python 3.11
dev environment passed: 130 passed, 8 skipped. The lock check passed
with Python 3.14.7. Workflow lint and shell syntax checks also passed.
Full Python 3.14 runtime and native release acceptance remain for CI.
2026-09-07 15:12:55 -04:00
ethernet
b0ab0162b0 feat(release): gate stable promotion through the full release pipeline
Run the entire CI workflow before Docker build and tests. Require Nix,
native payload smoke tests, install/update E2E and signed-package upgrade
acceptance before publishing. Keep Desktop Playwright E2E deferred.

Archive tested Docker images and signed bundle candidates with provenance
and hashes. Publishers consume those exact artifacts without rebuilding.
Advance stable channels only after all required publications succeed.
Keep canaries on their separate path and reject direct stable-builder
publication that bypasses the gate.

Move shared release transport, manifests and gates to Python. Keep native
Electron adapters in JS and share feed/MIME facts as JSON. Replace the
R2/feed JS implementation and move its protocol tests to Python.

Verified targeted Python and JS tests, real loopback transport and CLI
execution, temporary Git admission, workflow graph lint, and typechecks.
No live stable release was run. Native signing, package upgrades and real
registry/Store promotion still need their release-run receipts. Separate
services cannot promote atomically. A promotion failure keeps the run red.
2026-09-07 14:40:10 -04:00
ethernet
cd0f97f833 feat(python): pin bundled runtime to 3.14 everywhere (pm, termux lane, CI, installers)
pm python node: 3.14.7+20260901 (freshest python-build-standalone 3.14
build) for the 6 desktop targets; the bionic row moves from the third-party
TUR python3.11 deb to the official termux-main python_3.14.6-1 deb (which
lags PBS by one patch — pinned manually, documented). All 7 digests fetched
from the live sources (PBS release API + termux-main Packages index).
pm/packages.py: main_bin_rel python3.14, deb_package python, bionic fetch
constant, latest_versions guards bionic (no PBS build exists).

termux lane: PYTHON_ABI cp311->cp314, python3.11->python3.14 paths,
libpython3.11.so->3.14, TARGET_ENV 3.11.15->3.14.6 AND sys_platform
linux->android (CPython 3.13+ reports 'android', docs-verified) — linux-
gated markers no longer admit the termux target. runtime_libs.json needs no
change: every python 3.14.6-1 Depends is already staged.

CI: python-version/--python 3.11->3.14 across all 11 workflows incl. the
uv lockfile-check lane. Installers derive the minor from the lock already;
fallbacks bumped. Sandbox images nikolaik/python-nodejs:python3.11-nodejs20
-> python3.14-nodejs22 (tag exists). runtime_repair fall-forward cap now
tracks the <3.15 requires-python window. Docs/README python version claims
updated.
2026-09-07 14:19:18 -04:00
ethernet
49945b1402 refactor(tests): one runner everywhere — xdist removed, per-file isolation on every host
pytest-xdist is gone: run_tests.sh no longer dispatches on host, the
Windows xdist (--dist loadfile) arm is deleted, and pytest-xdist is
dropped from the dev extra (uv lock removes it + execnet). The
per-file subprocess runner (run_tests_parallel.py) is THE runner on
every host — the shape the CI linux lane already used.

- tests-os.yml: the macOS/Windows lanes now run scripts/run_tests.sh
  like every other lane, passing the marker-narrowed file list via
  --files and -m after --. The runner natively tolerates per-file
  empty collections (exit 5, platform-gated file) and fails the run
  when NOTHING collected — replacing the hand-rolled exit-5 branch.
- tests/gateway/conftest.py: the hasattr(config, 'workerinput')
  controller-guard was xdist-only dead code; the file-locked cache
  already handled the per-file model. Removed.
- tests/tools/test_browser_supervisor.py: the port came from xdist's
  worker_id (fixed 9225 under per-file isolation — concurrent files
  collide). Binds an ephemeral port instead (bind 0, read back).
- ~35 comment sites named xdist as the isolation mechanism; they now
  describe the shared-process hazard they actually guard against
  (bare pytest runs, same-file ordering) without naming a runner that
  no longer exists.
2026-09-04 19:01:41 -04:00
ethernet
e97de8d8c3 Merge branch 'ethie/windows-tests' into ethie/pm-clean
# Conflicts:
#	.gitignore
#	agent/deadline.py
#	pyproject.toml
#	scripts/run_tests.sh
#	tests/agent/lsp/test_install_and_lint_fixes.py
#	tests/cron/test_file_permissions.py
#	tests/cron/test_media_delivery_parity.py
#	tests/cron/test_script_claim_heartbeat.py
#	tests/hermes_cli/test_dep_ensure.py
#	tests/hermes_cli/test_gateway_wsl.py
#	tests/hermes_cli/test_linux_desktop_entry.py
#	tests/hermes_cli/test_npm_engine.py
#	tests/hermes_cli/test_runtime_repair.py
#	tests/hermes_cli/test_tui_npm_install.py
#	tests/test_hermes_constants.py
#	tests/test_install_autostash_conflict_recovery.py
#	tests/test_install_macos_launcher.py
#	tests/test_install_sh_acp_launcher.py
#	tests/test_install_sh_bootstrap_marker.py
#	tests/test_install_sh_node_deps_failure.py
#	tests/test_install_sh_symlink_stomp.py
#	tests/test_windows_subprocess_no_window_flags.py
#	tests/tools/test_approval_timeout_overflow.py
#	tests/tools/test_browser_homebrew_paths.py
#	tests/tools/test_find_shell.py
#	tests/tools/test_lazy_deps.py
#	tests/tools/test_lazy_deps_durable_target.py
#	tools/approval.py
#	uv.lock
2026-09-04 01:06:06 -04:00
ethernet
c35a6d65c9 merge: upstream/main into ethie/pm-clean
Bring in 597 upstream commits while preserving the branch's intentional
divergence (pm store, MSIX desktop, Termux removal).

Resolution notes:
- local_runtime/local-models cluster (30 both-added files): took upstream's
  evolved version — our side was a stale feat-merge snapshot with zero
  post-merge commits, and our non-conflicted importers were verified against
  upstream's exports.
- install scripts (install.ps1/install.sh/setup-hermes.sh): kept our staged
  pm-store bootstrappers (upstream still ships the old monolithic installer).
- Deleted-by-us files (node-bootstrap.sh, install_ps1 tests, termux.md):
  kept deleted — the pm store replaced that machinery.
- runtime_repair.py: kept ours (pm-based), restored upstream's managed_uv.py
  which surviving upstream files still import.
- Core files (run_agent, update_cmd, main.py, hermes_constants, estop,
  aux_client, browser_tool, desktop entry, config_defaults): per-file merges
  combining both sides' features (upstream fleet-wide estop, PID-identity
  daemon kill, editable-install guard; our utf-8-sig sweep, stable channel,
  PYTHONPATH desktop entry).
- Desktop/i18n: took upstream's flag-gated local-models components; merged
  both sides' i18n keys; merged run-electron-builder.mjs and notifications
  tests keeping both sides' tests.
- pyproject.toml: upstream's expanded exclude-newer list + our
  google-cloud-pubsub entry; uv.lock regenerated from the merged pyproject.
2026-09-01 20:22:04 -04:00
joaomarcos
c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
ethernet
47f4ab3a17 feat(desktop): bundle, publish, and update the desktop app as MSIX
Wire the desktop app onto the pm store for real distribution:
- MSIX bundle: electron-builder config, appx assets, manifest, copilot
  key + deep-link routing, App Installer + Windows Store variant
  (sign only the msix; inner binaries covered by the package block map)
- Rust CLI shim (apps/desktop/shim) — bundled builds run from the store
  python + shim, never the venv; payload symlinks relativized so the
  relocatable venv survives relocation
- Cloudflare R2 release pipeline: publish binaries + update feeds,
  nightly channels/tags, stamp-first version resolution
- Update system: gate, uninstall steward, boot bootstrap, release
  channels, update receipts
- install.ps1 reduced to a 361-line stage-protocol bootstrapper (heavy
  deps are pm's job); darwin updater + update-channel mirror ripped
- doctor: main's re-landed TCC anchor kept, termux branches removed

Rebuilt from ethie/pm onto the pm-store stack. 22 hot files hand-merged;
uv.lock + package-lock.json keep main's newer dep tree; test_engines
reads the pm/lock.json pin; lazy_deps.py deleted (all 222 importers
migrated to pm in the foundation commit).
2026-08-31 18:00:48 -04:00
ethernet
ef4a82eed0 refactor(tests): xdist is the standard runner; drop the per-file subprocess machinery
scripts/run_tests.sh now runs pytest-xdist -n <N> --dist loadfile as
the single canonical path on every OS (Linux and Windows CI lanes both
use it). The per-file subprocess model (run_tests_parallel.py, and the
interim run_xdist.sh experiment) is deleted along with its two
self-tests: persistent xdist workers pay the interpreter+import wall
once per worker instead of once per file (~0.5-1.5s x ~3400 files was
a ~6-minute floor on Windows), and --dist loadfile pins a file's tests
to ONE worker, bounding state pollution to co-scheduled files — which
is exactly the class of flake we are now committing to fix properly.

The serial process-killer quarantine phase is dropped too. It existed
to keep process-tree-sweep tests from killing sibling xdist workers;
the durable fix belongs in those tests (sweeps must target their own
children, not enumerate every python process), and keeping a
divergent two-phase path would hide that work.

Kept from the old wrapper: hermetic env -i scrubbing, Windows
location-var forwarding, venv probing, bytecode pre-compile,
-m 'not integration', and the HERMES_TEST_IMAGE docker-knob
allowlist. HERMES_TEST_FILE_TIMEOUT/FILE_RETRIES/SLICE go away with
the runner they parameterized; -j/HERMES_TEST_WORKERS now map to xdist
-n (Linux CI pins 96, Windows 32, default auto).

Docs updated to match: AGENTS.md (runner contract, flake policy,
isolation section), CONTRIBUTING.md, tests/conftest.py comments,
classify_changes docstring + its lane expectation (a .sh runner no
longer trips the supply-chain scan lane), comfyui README,
hermes-agent contributor guide, debugpy skill and its website doc.
2026-08-31 17:52:32 -04:00
ethernet
9a81c291b9 fix(registries): reads resolve None to the active home scope; calm windows CI
normalize_scope unified the two scope contracts that live on the same
registries, which was wrong: WRITE paths (register/snapshot/restore)
treat None as the process-global layer, but READ paths (get_provider/
list_providers/registry_generation) treat None as 'resolve to the
active home's scope' — hermes_home_key(None) is the active default home.
Routing reads through normalize_scope hid every scoped registration
from ambient-home lookups: 52 linux failures across plugin discovery,
secret-source profile isolation, web-search/browser backend selection
(the CI run on e66a627aa5).

- read sites go back to hermes_home_key(scope); write/slot sites keep
  normalize_scope (None preserved); the docstring now spells out both
  contracts so the unification bug can't be reintroduced blind
- run_tests_parallel.py default worker count: cpu_count*2 -> cpu_count.
  Oversubscription made the 32-core windows runner run 64 pytest
  subprocesses; the suite's own linux measurement showed one worker per
  core is the flat optimum, and windows CI showed the cost — 160 files
  past the 300s per-file cap (they pass in seconds on an idle box), each
  burn a kill + retry cycle, and the churn dominated the lane's wall
  clock. docker.yml's explicit nproc pin now just documents its own
  intent instead of fighting a *2 default
- windows CI lane: HERMES_TEST_FILE_TIMEOUT=900 (via a matrix field)
  so contended-but-passing files stop getting killed at 300s
2026-08-31 17:52:32 -04:00
ethernet
10f99bc15e ci: run the work lanes on larger runners and merge the split jobs
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.

The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.

Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.

The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.

JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.

The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.

The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.

`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.

node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.

The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.

The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.

`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.

The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.

Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
  the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
  `ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
  returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
  parallel units and for a plain `npm run check`, in both directions. Against
  the 13-leg matrix the count is 13 to 11, and the whole difference is the
  three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
  the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
2026-08-22 02:25:12 -04:00
ethernet
1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
ethernet
7aecab56db ci: move the review comment and the image build out of the CI run
The CI run stayed in progress until its last job ended. Two advisory jobs
set that time: the review-comment poller (40 minutes) and the Docker image
build (45 minutes). Neither job was required to merge.

GitHub refuses `gh run rerun` on a run that is in progress. Thus a reviewer
who added the `ci-reviewed` label had to wait for the two slow jobs, and
label-rerun.yml carried a 2100-second wait loop for this reason. The fast
required jobs were ready long before.

Each slow job now runs in its own workflow:

- docker.yml owns its `pull_request` trigger and does its own change
  detection. The new `detect` job runs the same composite action with the
  same condition that ci.yml applied, so a tests-only PR still skips the
  build. The `workflow_call` trigger is gone.
- ci-review-comment.yml starts on `workflow_run` when CI starts. It reads
  the workflow and the scripts from the default branch, which is the trust
  boundary that the old job got from its `ref: default_branch` checkout.

The poller reads job results through the API, so it can report on a run
that it does not belong to. `WATCH_WORKFLOWS` names sibling workflows for
the same commit, and `select_watched_runs` keeps the newest run for each
name. Thus the comment still shows the Docker results. The list is
newline-separated, because a workflow name can contain a comma.

The poller always exits 0 now. It reports on the CI run from a different
run, so a failed CI job is not a failure of the poller. The CI run has its
own gate for that.

Also correct a parse error in label-rerun.yml. STATUS came from the already
truncated RUN_ID, so its value was the run id and never "completed". Thus
the wait branch always ran.

ci.yml no longer needs `packages: write`, because the image build has left.
2026-08-09 15:40:04 -04:00
ethernet
177002838b ci: retry uv python install 2026-08-03 11:26:56 -04:00
Kshitij Kapoor
acfd376d66 ci(docker): retry buildx setup on transient Docker Hub failures
The Docker Build, Test, and Publish workflow fails when
docker/setup-buildx-action can't pull the moby/buildkit:buildx-stable-1
image from Docker Hub. The failure happens during builder bootstrap at
the auth token exchange — a transient network blip (connection reset,
read timeout, rate limiting) that self-resolves on re-run.

Recent failure (run 30449230291, merge job):
  read tcp 10.1.0.171:45666->104.18.43.178:443: read: connection reset by peer

This has hit us before and will again — it's the same class of
transient Docker Hub flake that the merge job already retries for
imagetools create. But buildx setup had no retry, so a single network
hiccup killed the entire job (build, publish, or merge) even though
nothing was wrong with the code or the image.

Fix: wrap each of the 3 buildx setup steps (build, publish, merge jobs)
with continue-on-error + a conditional retry step. The maintained action
is preserved as-is — we just give it a second attempt if the first
fails. The action generates a unique builder name per invocation, so the
retry never collides with the failed first attempt. The second attempt
has no continue-on-error, so genuine persistent failures still fail the
job.

The docker/setup-buildx-action maintainer has explicitly said retry
belongs at the workflow level, not inside the action [1], and other
repos use this same continue-on-error pattern for this exact issue [2].

[1] docker/setup-buildx-action#510
[2] joshjhall/containers#688, ethpandaops/eth-client-docker-image-builder#391
2026-07-30 21:12:59 +05:30
Teknium
43333acdda ci: pin uv version in setup-uv to eliminate per-job manifest fetch
Unpinned, astral-sh/setup-uv resolves 'latest' by fetching
https://raw.githubusercontent.com/astral-sh/versions/.../uv.ndjson on
EVERY job. A transient failure of that fetch fails the whole job before
any test runs (2026-07-28: tests slice 5/8 died 12s in with
'##[error]fetch failed' on PR #73514). Pinning version makes setup-uv
download the binary directly — one less external hop per job across all
7 call sites (tests, lint x2, docker, e2e-desktop, lockfile-check).
2026-07-28 17:59:23 -07:00
ethernet
3f9944bad9 fix(ci): run trusted Docker publish directly (#69803) 2026-07-22 23:36:45 -04:00
ethernet
957ea640de fix(ci): publish inline E2E evidence (#69699)
* fix(ci): publish inline E2E evidence

Upload bounded screenshot evidence from E2E, then publish validated images
from a trusted workflow_run job to commit-pinned branches in the evidence repo.

Wait briefly for the live CI review comment marker before publishing, so
GitHub's read-after-write delay cannot leave an orphaned evidence branch.

* fix(ci): isolate privileged credentials from PR jobs

Keep App private keys and Docker Hub credentials out of PR-controlled
workflows. Use protected environments for trusted publishing and a public
repository variable for the App client ID.

* fix(ci): attach E2E evidence with restricted bot session

Replace the App-backed evidence repository publisher with gh-image uploads
from a dedicated bot session in the gh-image environment.

* fix(ci): publish validated E2E evidence from forks

Let the trusted default-branch publisher handle bounded, validated evidence
artifacts from fork PR CI without checking out or executing fork code.
2026-07-23 02:15:23 +00:00
Teknium
597615ade4 fix(ci): make tests, workflows, and attribution reliable under load (#66373)
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory

The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.

New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.

- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
  merged with the directory at import time (directory wins). All
  existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
  conflicting reassignments (incl. against the legacy map), validates
  email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
  a legacy entry; failure message prints the exact add_contributor
  command. Also auto-resolves bare <login>@users.noreply.github.com
  emails is intentionally NOT added (kept id+login form only, matching
  previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
  merge precedence, CLI idempotency/conflict/validation, subprocess E2E.

* feat(ci): one-shot per-file flake retry in the parallel test runner

A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.

- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
  deterministic failure still exits 1; retries=0 restores old behavior.

This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.

* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s

These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.

* fix(ci): job timeouts everywhere + retries on all network installs

Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
  burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
  that lacked it: pip installs (deploy-site, skills-index), npm ci
  (deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
  test deps). Deterministic build steps (npm run build) deliberately
  NOT retried — split into separate steps so a real build failure fails
  fast instead of retrying 3x.

* docs(agents): document the file-retry flake policy

* fix(ci): curl retries on deploy hook + skills-index probe

* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile

From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
  run_id-suffixed keys — the cache never matched once, so LPT slicing
  always ran blind and unbalanced slices pushed heavy files toward the
  per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
  'gh pr view || true' turned an API blip into 'label absent' → false
  BLOCKING failure. Now 3x retry, and API failure is reported as an API
  failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
  silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
  so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
  consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
  curl --retry 3 (ADD cannot retry; checksums still enforced); npm
  --fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
  get continue-on-error so an artifact-service blip can't fail a green
  test slice.

* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list

- test_tui_gateway_server.py: session.create / non-eager session.resume
  arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
  test and fires into the NEXT test's _make_agent mock, racily
  corrupting captured state (the recurring session_resume shard
  failures). Replaced the per-test whack-a-mole stub with a module-wide
  autouse fixture; the 3 worker-lifecycle tests that genuinely need the
  deferred build opt back in via @pytest.mark.real_agent_prewarm (new
  marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
  live PROVIDER_REGISTRY instead of a hand-list that had drifted
  (missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
  tests failed on any machine with HF_TOKEN exported. E2E-verified with
  HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.

* test: de-flake 30 timing-sensitive test files for loaded CI runners

Root-cause fixes from the flake audit (session-DB mining + repo sweep):

Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
  sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
  replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
  unbounded blocking read (parent wedge now fails THIS test with a clear
  message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
  (the 1s partition window mid-interpreter-startup is how a child PID
  escaped the live-system guard in CI)

Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
  mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
  mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
  voice_cli_integration, docker_environment, session_store_lock_io,
  planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
  (joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
  10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
  setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
  5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
  iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
  0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
  compression fork-lock TTL 1s->3s (12 refresh chances per lease);
  compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)

* fix(tests): repair indentation from de-flake batch edit

* fix(tests): harden env isolation and replace remaining sleep-sync races

The full 42k-test run and complete npm check surfaced three more classes:

- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
  leaked into Python/TUI tests. Pin the default Honcho host in the
  hermetic fixture, isolate the one fallback test from ~/.honcho, and
  blank SSH_* around terminalSetup tests. This flipped 20 false failures
  back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
  time.sleep globally, then busy-polled with that same mocked sleep. Under
  full-suite load the poller could starve the writer. Each test now waits
  on an Event emitted by the exact flush/retry transition; 30/30 passed
  under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
  not fire before its assertion. A loaded runner descheduled the test for
  >500ms and both chunks arrived. Producer controls now gate second-chunk
  and completion transitions explicitly.

Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.

* refactor(ci): use gh bot pat, better retries

refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.

Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference

Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.

ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth

Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.

19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
  comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call

---------

Co-authored-by: ethernet <arilotter@gmail.com>
2026-07-17 20:55:24 +00:00
emozilla
2abe11a7fe security(ci): pass untrusted refs through env, not run: interpolation
lint.yml inlined github.head_ref (the fork PR branch name, attacker-
controlled) into the diff-summary run: block. GitHub expands ${{ }} into
the script text before bash tokenizes it, so a branch like x$(id) runs on
the lint runner. The pull_request trigger keeps the token read-only, but
the sink still allows CI resource abuse and cache/artifact tampering, and
would become RCE-with-secrets under pull_request_target.

Route head_ref through an env var (env values are not subject to expression
injection) and reference "$HEAD_REF". Apply the same to the two docker.yml
sites that interpolate github.event.release.tag_name.

Fixes GHSA-jpw6-c7jr-c56v, GHSA-2843-hjmf-7x96.
Credit: @technotion, @youngstar-eth.
2026-07-03 12:44:30 -04:00
ethernet
cca8b4ef4e fix(ci): unify amd64/arm64 docker pipelines 2026-06-29 19:17:04 -07:00
ethernet
2fa66950e8 change(ci): upload-artifact from v4 -> v7 2026-06-26 19:15:18 -07:00
ethernet
4b0a2040e7 change(ci): use run_tests in docker 2026-06-26 19:15:18 -07:00
ethernet
18f7ad49ab change(ci): update all UV installs 2026-06-26 19:15:18 -07:00
ethernet
f0cb049217 change(ci): migrate docker smoketests to real tests 2026-06-26 19:15:18 -07:00
ethernet
fb1dd1bf91 change(ci): docker-publish.yml -> docker.yml 2026-06-26 19:15:18 -07:00