Commit Graph

55 Commits

Author SHA1 Message Date
ethernet
babbec1c4c feat(pm): isolate developer test environment from runtime extras 2026-09-23 18:13:38 -04:00
ethernet
13ebc163a1 ci: fold the PowerShell installer job into the tests-os Windows lanes
installer-tests.yml predates nothing it still owned. Its pytest step
(test_source_launcher_stages.py) is platforms("windows") and already runs in
both tests-os Windows lanes, so every installer PR ran it twice. The
`installer` lane never gated anything on its own either: every path that set
it also sets `python`, which gates tests-os.

The two standalone scripts/tests/*.ps1 suites become one platforms("windows")
pytest file parametrized over Windows PowerShell 5.1 and pwsh 7, so
list_os_marked_tests picks them up with everything else. The `installer` lane
goes away from the classifier, detect-changes, ci.yaml and the
all-checks-pass gate; the classifier contract now pins that install.ps1 and
its suites turn `python` on.
2026-09-23 17:19:29 -04:00
ethernet
3a6e61b189 ci: drop the PM toolchain smoke matrix; key the tools cache on the PM code
The 6-target cold->warm smoke existed to prove setup-pm still cold-boots
when its code changes without a lock bump, and to reach linux-arm64,
darwin-x64 and win32-arm64. Both are now covered without a dedicated
workflow:

* The tools cache key hashes pm/**, the action itself and
  scripts/ci/setup_toolchain.py, not only pm/lock.json. A provisioning
  change misses the cache on every lane that uses the action, so the
  cold path runs where the tests already are.
* tests-os gains a windows-11-arm leg running the same windows-marked
  files (tests/pm carries ten of them). It needs the ARM64 build deps
  because several extras build from sdist, and fewer workers on the
  4-core runner.

The Windows SDK adapter test is the one thing left that no other lane
ran natively; it keeps its two Windows runners under
windows-bundle-sdk.yml, path-triggered on the signing scripts. The
run-scoped cache cleanup workflow and its script only served the smoke
and go with it.
2026-09-21 12:07:53 -04:00
ethernet
db3ac3ea00 fix(ci): the PM cache prune reads its lock-source input under the runner's real name
The runner maps a hyphenated input to INPUT_LOCK-SOURCE, not INPUT_LOCK_SOURCE.
The registration step read the underscored name, saw nothing, and threw
"invalid lockSource" from every job that composed setup-pm, so the whole CI
run failed before any consumer ran. The unit test used the same wrong
spelling, which is why it stayed green; it now drives the real mapping.
2026-09-19 01:38:09 -04:00
ethernet
cf0bc5d5fa perf(ci): restore a uv cache that actually warms; give the electron toolchain a producer again
The CI cache barely ever restored anything useful, for four independent
reasons found in live run logs and the fork's cache store:

- `uv cache prune --ci` (setup-pm/prune) discards downloaded wheels before
  the save, so the snapshot carried only source-built wheels — the next
  run "restored" it and still cold-downloaded everything. Upstream main's
  logs showed the end state: a 209KB stub cache exact-hitting forever.
- Tool-only jobs (icons-freshness etc.) auto-saved a 1.5KB empty uv cache
  under the production exact key; caches are immutable, so the stub won
  forever and blocked real saves.
- The npm cache key had no restore-keys, so one lockfile bump missed the
  exact key and every npm job went cold.
- `2efa4ff94f` deleted the electron-builder toolchain save but left the
  assemble job restoring `eb2-` — a fossil nobody regenerates; assembly
  cold-downloads winCodeSign/ATS/dotnet on every release.

Changes:

- pm/cache_lock.py: move prune_uv_cache_to_lock out of
  scripts/bundles/native.py (which re-exports it); the lock-exactness
  contract now serves both the bundle ship gate and CI caches.
- pm.build_env learns --exact-lock --lock-source: prune cache entries the
  project uv.lock cannot resolve, keeping lock-required downloaded wheels
  (unlike --ci). Refuses --ci/--prune-cache combinations.
- setup-pm: python-cache auto-save now requires extras (no stubs from
  tool-only jobs); key drops the prune flag and bumps to v3 — pruned and
  unpruned saves share one namespace since both are lock-exact; pre-save
  pruning switched from --ci to --exact-lock; npm cache gains a
  lockfile-agnostic restore prefix.
- save-pm-cache: same exact-lock prune before explicit saves.
- desktop-bundled-release: build legs (cache-mode: write) restore+save the
  electron-builder toolchain under eb3- keyed on the locked builder
  version; assemble restores the same namespace; the dead default-cache
  resolution step is removed (assembly resolves no electron artifacts).
- cleanup_pm_toolchain_caches.py: match v2 and v3 smoke keys.

Validation: tests/scripts/test_bundle_native.py 11/11 (incl. both
lock-prune gates), tests/pm failures identical before/after the diff,
tests/scripts/test_bundle_payload.py 5/5, tests/ci cleanup 3/3,
tests-js setup-pm-post 1/1 and the three setup-pm-cache contract tests
updated and green (6 failures in that file pre-date this diff and are
drift between the workflow and its stale assertions); pm.build_env
--exact-lock E2E against a real uv cache copy pruned 3 stale entries and
kept the rest.
2026-09-15 12:56:39 -04:00
ethernet
5f97cd3fba feat(ci): pm-locked ffmpeg + minimal chat driver deps for native smoke legs
The e2e-screen-record action installed ffmpeg through three different
OS package managers (apt, brew, winget). winget is the flaky leg on
windows-11-arm and serves an x64-gyan build that runs emulated on ARM;
choco's community package wraps the same gyan x64 zip, so a choco
fallback would not fix either problem. PM already ships a locked,
sha256-pinned ffmpeg (martin-riedl posix, BtbN win32 including a NATIVE
winarm64 build), so the recording action now verifies ffmpeg on PATH
instead of installing it, and the pinned binary rides the same
tools-cache as node/python.

- setup-pm learns a `packages` input (extra PM tools beyond the
  toolchain roots, e.g. ffmpeg) threaded through setup_toolchain.py's
  prepare/install/archive-inputs phases; the tools-cache key gains an
  extra-packages fragment so existing keys stay byte-identical.
- The six chat-driver jobs pass `packages: ffmpeg` to setup-pm.
- e2e-screen-record drops the apt/brew/winget install steps, the
  ffmpeg actions/cache steps and the save-cache input; Xvfb (headless
  linux) and the macOS replayd-approval hack stay.
- The "Install locked chat driver dependencies" step installs only the
  tests-js workspace with --omit=dev instead of the whole apps/desktop
  tree: the drivers need @playwright/test, zod (previously a phantom
  hoisted from @assistant-ui/react), js-yaml and semver only. 64 pure
  packages in ~2s vs ~1800 including electron-builder and native
  builds; the tree-shaken node_modules runs the real driver modules
  (verified by importing desktop-chat-smoke.ts and update-window-chat
  end to end).
- tests pin the new contracts; tests/install/README.md updated.
2026-09-14 13:35:13 -04:00
ethernet
59b8eeb5c5 fix caching and record 2026-09-14 11:58:30 -04:00
ethernet
2efa4ff94f refactor(desktop): prepare dependencies before saving build caches
Dependency acquisition during packaging left native wheels and packager
inputs outside the pre-build cache save. Compose PM and existing providers
into a preparation phase, then require builds to consume admitted inputs.

Share native preparation with PM Bundle. Keep path-bound environments and
signing outputs separate from reusable caches. Use read-only cache tokens
for commit builds and preserve the one-command local build path.

Verify pinned tools through PM, probe PTYs under the prepared Electron,
and supply dmgbuild through a build-only PM package. Resolve bundled tool
stores from their payload manifest so relocation preserves discovery.

Validation: focused Python and JS tests, checkJs, Ruff, Windows checks,
anti-slop, cache relocation, and network-denied Linux AppImage builds.
Relocated runtime smoke passed with NixOS host libraries supplied.
Native Windows/macOS signing and live GitHub cache behavior remain untested.
2026-09-13 14:28:31 -04:00
ethernet
9a42d60f25 fix(build): handle large release uploads and trim packaging waste 2026-09-13 11:04:03 -04:00
ethernet
5e4a2a3d24 refactor(pm): remove legacy dependency and launch managers
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.

Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.

Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.

Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.

Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
2026-09-12 14:57:38 -04:00
ethernet
a154b89b9f Route build and CI Python preparation through PM operations 2026-09-11 18:14:43 -04:00
ethernet
fea2858c99 merge: unify shared product builders, caches, and Windows prerequisites
Merge ethie/shared-product-builders with the CI dependency cache and native Windows setup work. Preserve UTF-8 diagnostics in the shared Python environment runner. Pass a persistent cache through isolated native staging and PM-runtime construction. Reuse one Windows prerequisite installer from source setup, native adapters, and CI, preserving Rust homes across HOME isolation.

Verified 85 targeted Python tests (5 host skips), 18 JavaScript tests, workflow validation, and scoped lint/typecheck. On native Windows ARM64, five prerequisite contracts passed and the actual shared provider reused OpenSSL, compiled its header with MSVC, and retained Rust under isolated HOME. Full signed distribution builds and live Actions cache transfer remain CI verification.
2026-09-11 13:45:05 -04:00
ethernet
4cc2b7bab5 fix(ci): remove legacy uv cache compatibility 2026-09-11 13:32:32 -04:00
ethernet
493ae9daa3 fix(ci): reuse uv wheels and save caches after bundle failures 2026-09-11 13:28:15 -04:00
ethernet
c8a9682505 fix(pm): preserve pinned binary inputs in R2
Termux removes old package files, so a pinned URL and hash do not keep
build inputs available. Preserve the exact bytes without changing pins.

Archive every PM HTTP artifact and the Termux runtime inputs by SHA256.
CI reads R2 first. Only a missing object permits an upstream download,
hash verification, immutable upload, and verified readback. Seed the
actual toolchain and payload stores before their consumers run.

Use the public archive as a pinned fallback in PM, bootstrap installers,
and Nix fetchers. Keep network retries bounded and report attempted URLs.
Keep publication credentials in protected CI jobs, not installed clients.

Verification:
- 283 targeted tests passed; five POSIX tests skipped on Windows.
- All 87 preserved Termux packages passed local archive miss/hit checks.
- Native ARM64 ripgrep installed through the mirror and ran successfully.
- Wheel import, workflow lint, Python lint, shell syntax, and pins passed.

Live R2 publication, POSIX tests, and Nix builds remain for native CI.
The real-byte archive checks used loopback HTTP, not the live bucket.
2026-09-10 18:17:08 -04:00
ethernet
e4cc7f09d9 merge: integrate upstream catalog with PM publication
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.

Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.

Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
2026-09-09 16:49:27 -04:00
ethernet
1c8fae6180 fix(pm): preserve runtime and user state across failure paths
Keep downloads bound to their remote representation and publish through
atomic destination-local staging. Serialize shared partial ownership.

Keep explicit CA trust scoped to provider probes. Preserve checkpoint
history and edited files, validate all profile inputs before dependency
publication, and separate data removal from installed runtime ownership.

Exclude machine-specific PM state from portable transfers. Keep plugin
files and nested skill tools intact. Preserve native test isolation.

Focused native Windows receipts cover the individual repairs and their
integration. This commit does not claim a full-suite or release build.
2026-09-09 15:17:08 -04:00
Teknium
d47adec28f Merge origin/main into feat/plugin-catalog
Python plugin CLI/loader/web/tui files taken from main wholesale; the
catalog layer is re-ported onto main's decomposed shapes in the
following commits. plugin_index.py removed (catalog is the sole
discovery system).
2026-09-09 04:15:27 -07:00
ethernet
d36562ac9f fix(release): provision Windows bundle tools on cache misses 2026-09-08 14:32:30 -04:00
ethernet
b676997d2d ci: provision locked Python and Node toolchains through PM 2026-09-08 12:45:15 -04:00
ethernet
37650f810c merge: integrate desktop update acceptance tests
Preserve the PM runtime-repair module boundary. Port the incoming stderr-streaming fix without restoring the deleted managed_uv downloader. Targeted runtime/progress tests and root JS checks passed; desktop typechecks passed.
2026-09-06 21:59:45 -04:00
ethernet
642579db60 Merge remote-tracking branch 'upstream/main' into ethie/pm-clean
# Conflicts:
#	.github/actions/detect-changes/action.yml
#	.github/workflows/ci.yaml
#	.github/workflows/tests-os.yml
#	agent/prompt_builder.py
#	agent/ssl_verify.py
#	agent/subdirectory_hints.py
#	apps/desktop/electron/main.ts
#	apps/desktop/electron/preload.ts
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/global.d.ts
#	apps/desktop/src/i18n/ar.ts
#	apps/desktop/src/store/updates.ts
#	cron/suggestions.py
#	gateway/channel_directory.py
#	hermes_cli/config.py
#	hermes_cli/doctor.py
#	hermes_cli/linux_desktop_entry.py
#	hermes_cli/main.py
#	hermes_cli/web_routers/profiles.py
#	hermes_constants.py
#	plugins/platforms/photon/adapter.py
#	scripts/ci/classify_changes.py
#	scripts/install.ps1
#	tests/agent/test_relay_runtime_plugins.py
#	tests/ci/test_classify_changes.py
#	tests/hermes_cli/test_gui_command.py
#	tests/hermes_cli/test_linux_desktop_entry.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/state/test_fts_runtime_rebuild.py
#	tests/tools/test_lazy_deps.py
#	tests/tools/test_macos_protected_search.py
#	tools/browser_tool.py
#	tools/file_operations.py
#	tools/lazy_deps.py
#	tools/mcp_tool.py
#	tools/working_diff.py
#	uv.lock
2026-09-04 03:18:16 -04:00
yoniebans
49a5c440c5 fix(install-e2e): review follow-ups — harness needle, per-matrix cap wording, typo
The harness asserted the driver still throws 'not implemented yet' for
the desktop-installer@latest update route; that arm is implemented now
(Invoke-PhaseInstallGui -Mode "update"), so the check failed against
its own tree. It asserts the implemented contract instead.

The 256-job cap wording in the workflow comment and README now states
the scope GitHub applies it at: each per-OS matrix separately, not the
combined leg count. At the 10-tag bound the largest matrix is windows
at 180.

e2e-screen-record comment: hhttps -> https.
2026-09-03 10:15:53 +02:00
yoniebans
558d76c401 Merge upstream main (afc3d9d34c): refresh before review
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
2026-09-02 17:53:49 +02:00
Teknium
bfbb34bbec fix(update): Windows progress server hands out its URL only once it is serving
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).

- windows.ps1: readiness handshake after BeginInvoke — one /progress
  round-trip must succeed (≤15s) before the server is returned; on failure
  tear the listener down and continue without UI. The URL now means
  "serving", not "bound". Also fixes the browser opening to a page that never
  loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
  consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
  spawn the real PowerShell script; the Windows-only job now runs them only
  when scripts/desktop-update/**, the Electron updater launcher, conftest,
  pyproject, or those tests change (push/dispatch fail open). A PR that
  never touched that surface cannot be failed by its process timing.
2026-09-02 00:44:19 -07:00
yoniebans
d19038e76b Merge upstream main (b81383ec21) into the install-e2e suite branch
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
  the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
  one-line compat forwarder) and the new implementation already drains
  both pipes asynchronously with bounded abandonment, which supersedes
  this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
  import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
  addition inside the consolidated file; rewired the five upstream specs
  still importing './mock-server' to the consolidated path (export sets
  verified identical) and dropped the superseded apps/desktop/e2e copy.
2026-09-01 19:33:13 +02:00
ethernet
3d12e86ef1 feat(pm): unified package manager — pm store foundation
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.

Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
2026-08-31 18:00:48 -04:00
Teknium
7e17ff0ab3 Merge origin/main into feat/plugin-catalog — reconcile with landed index/manifest-v2 tracks 2026-08-27 21:29:42 -07:00
Brooklyn Nicholson
5caea5e501 ci: run cargo test for the bootstrap installer
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.

Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
2026-08-20 11:40:49 -05:00
ethernet
00c3872882 feat(ci): add nix flake check as unrequired job
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
2026-08-18 20:42:06 -04:00
ethernet
1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
brooklyn!
9eab7a4473 fix(ci): stop running uv lock --check on PRs that can't touch the lockfile (#84675) 2026-08-12 13:30:01 -05:00
ethernet
6fd7f05de4 fix macos screen record 2026-08-12 10:40:57 -04:00
ethernet
cdaf0cf091 ci(install-e2e): one screen-recording mechanism on every runner, Xvfb for headless linux
The composite action .github/actions/e2e-screen-record owns setup and
lifecycle on all three OSes: ffmpeg via apt/brew-verify/winget+cache,
capture via x11grab/gdigrab/avfoundation, mkv at 15fps stopped by 'q'
on live stdin with kill fallback. Linux runners have no display, so
start brings up a dedicated Xvfb :99 and exports DISPLAY - one display
serves both the recorder and any app a later step launches.

Recording moves out of the GUI driver into workflow infrastructure -
that is what makes it uniform - and a missing ffmpeg or a zero-frame
file now FAILS the leg instead of skipping silently: the graceful-skip
path is how the windows leg shipped no recording.mkv while green.

Lifecycle proven locally: start against lavfi testsrc, q-stop, ffprobe
duration check (record-start.sh/record-stop.sh under nix ffmpeg).
2026-08-12 04:15:24 -04:00
Brooklyn Nicholson
34833303f5 ci(install): actually run the PowerShell installer tests
scripts/tests/ has held three PowerShell suites that no workflow ever
invoked -- there is no Windows runner in CI, so they have been inert since
they landed. A regression test nothing executes is worse than none: it
reads as coverage.

Adds a windows-latest job, gated on a new `installer` lane so it only fires
for PRs touching install.ps1 or its tests. The 8.3 suite runs under both
pwsh 7 and Windows PowerShell 5.1, since install.ps1 arrives via `irm | iex`
into whichever shell the user already has and 5.1 is what ships with Windows.

Only the 8.3 suite is wired up. The other two fail on main today for
unrelated reasons; they can join once they are fixed.
2026-08-04 15:34:32 -06:00
Teknium
cc0af6b9e8 ci: skip Desktop E2E + Docker build on tests-only PRs (python_prod lane)
After the test-suite prune, the Python slices (~2.3m each) are no longer
CI's critical path — Desktop E2E (5.2m, the longest job) and the Docker
build are, and both run on every python-lane PR even when the diff never
leaves tests/. Neither consumes the test suite: Playwright drives the
built app + hermes serve backend, and the image copies installed code.

New python_prod lane = python minus tests-only diffs. e2e-desktop and
docker gate on it; every pytest/lint lane keeps gating on python.
Fail-open contract preserved: .github/ changes and empty diffs set
python_prod=true, and runner infrastructure (scripts/run_tests.sh,
run_tests_parallel.py) is deliberately NOT tests-only since a bad
runner edit can mask real failures.

Replay over the last 231 main commits: 39 (17%) would skip both jobs,
cutting their critical path from ~8m to ~3m. E2E-verified through the
real script entrypoint (tests-only/prod/mixed/fail-open) + 83 tests/ci
green.
2026-08-01 14:59:14 -07:00
ethernet
957ea640de fix(ci): publish inline E2E evidence (#69699)
* fix(ci): publish inline E2E evidence

Upload bounded screenshot evidence from E2E, then publish validated images
from a trusted workflow_run job to commit-pinned branches in the evidence repo.

Wait briefly for the live CI review comment marker before publishing, so
GitHub's read-after-write delay cannot leave an orphaned evidence branch.

* fix(ci): isolate privileged credentials from PR jobs

Keep App private keys and Docker Hub credentials out of PR-controlled
workflows. Use protected environments for trusted publishing and a public
repository variable for the App client ID.

* fix(ci): attach E2E evidence with restricted bot session

Replace the App-backed evidence repository publisher with gh-image uploads
from a dedicated bot session in the gh-image environment.

* fix(ci): publish validated E2E evidence from forks

Let the trusted default-branch publisher handle bounded, validated evidence
artifacts from fork PR CI without checking out or executing fork code.
2026-07-23 02:15:23 +00:00
ethernet
433673067e ci: surface E2E screenshots in review comment (#69631)
* ci: surface E2E screenshots in review comment

* ci: mark completed review commits in past tense

* ci: surface approved sensitive-file reviews

* ci: link sensitive files to reviewed changes

* ci: stage desktop E2E visual evidence

Track screenshots newly introduced against main and package visual diffs for a trusted publisher.

* fix(ci): pass E2E evidence output paths

Supply the manifest and staging-directory arguments required by the screenshot status helper.

* fix(ci): download the OSV SARIF artifact

Match the artifact name and result filename emitted by the pinned upstream reusable workflow.
2026-07-22 23:19:53 +00:00
Teknium
8dd07bd517 feat(ci): plugin-validate reusable action + plugin-catalog admission gate
- scripts/validate_plugin_catalog.py: standalone stdlib+pyyaml structural
  validator for plugin-catalog entries and removed.yaml (no hermes install
  needed; runtime twin of hermes_cli/plugin_catalog.py). --json support,
  unknown top-level keys warn instead of failing for forward compat.
- .github/actions/plugin-validate: composite action plugin authors drop
  into their own repo's CI — installs hermes-agent from a chosen ref and
  runs 'hermes plugins validate <path>'.
- .github/workflows/plugin-catalog-ci.yml: admission gate on PRs touching
  plugin-catalog/** — structural job plus pinned-source job that clones
  each changed entry's repo, hard-fails on unreachable pinned sha
  (supply-chain gate), and validates the plugin at that exact commit.
2026-07-22 07:22:43 -07:00
ethernet
1f76bdc5b2 fix(ci): pass App secrets as inputs to composite action
Composite actions cannot access the secrets context — the runner's
template engine rejects secrets.* references at load time with
'Unrecognized named-value: secrets'.

Move APP_ID and APP_PRIVATE_KEY from direct secrets.* references inside
the composite action to inputs passed by each calling workflow. The
fallback logic (GITHUB_TOKEN when APP_ID is empty, for fork PRs) stays
in the composite action's check step.
2026-07-20 17:27:30 -04:00
ethernet
7a69b82ad4 ci: migrate AUTOFIX_BOT_PAT to GitHub App token
Replace the long-lived fine-grained PAT (AUTOFIX_BOT_PAT) with short-lived
(1-hour) installation access tokens minted via a new get-app-token composite
action wrapping actions/create-github-app-token@v3.2.0.

The PAT was used in 13 spots across 8 workflow files for gh CLI / GitHub API
calls. The per-repo GITHUB_TOKEN (1,000 req/hr) was getting rate-limited when
multiple workflows fire concurrently (deploy-site, skills-index, ci-timings,
supply-chain-audit, js-autofix). App installation tokens get 5,000 req/hr
per installation and are scoped to the App's permissions, not a user account.

New composite action: .github/actions/get-app-token/
  - Wraps actions/create-github-app-token@bcd2ba49 (v3.2.0, SHA-pinned)
  - Reads APP_ID + APP_PRIVATE_KEY repo secrets
  - Outputs a 1hr installation token via steps.app-token.outputs.token

Requires two new repo secrets (set after creating the GitHub App):
  - APP_ID: the App's numeric ID
  - APP_PRIVATE_KEY: the PEM private key

App installation permissions needed:
  contents: write    (js-autofix push, pypi release upload)
  pull-requests: write (js-autofix PR create/merge, supply-chain comment)
  issues: write       (skills-index-freshness issue creation)
  actions: write     (skills-index workflow trigger)
  workflows: write   (skills-index triggers deploy-site.yml)

The AUTOFIX_BOT_PAT secret can be deleted once CI passes on this PR.
The comment in js-autofix.yml noting that PAT pushes trigger downstream
workflows is updated — App tokens have the same property (they are not
GITHUB_TOKEN), so the concurrency-cancel loop logic is unchanged.
2026-07-20 16:48:25 -04:00
Siddharth Balyan
21f971bbab fix(ci): null-safe files iteration in the paginated compare (#66867)
A PR more than 100 commits ahead of its merge-base paginates the compare
endpoint; pages after the first carry files: null, so the bare .files[]
made jq die with 'cannot iterate over: null' on every retry and forced
the classifier's fail-open path (all lanes on, ci_review gate demanding
a label with zero CI-sensitive files in the diff). .files[]? keeps the
page-one file list and ignores the null tail.
2026-07-18 19:17:09 +05:30
Teknium
1e01a4bbe7 fix(ci): restore fork-safe token fallback on PR gates broken by #66373 (#66577)
#66373 swapped GITHUB_TOKEN -> AUTOFIX_BOT_PAT across the workflows. That
PAT is empty on fork PRs (forks get no repo secrets), which broke every
fork PR two ways:

1. detect-changes classified with the empty PAT -> the compare API failed
   all 3 retries -> the classifier failed open and force-enabled the
   ci_review lane on EVERY fork PR.
2. The ci-reviewed / mcp-catalog-reviewed label gates then read labels with
   the same empty PAT via a hard-failing retry step -> the job failed with
   no recovery a fork contributor could perform (they can't self-add the
   label; re-running can't fix it).

Restores the pre-#66373 fork-safe behavior without reverting the commit's
real improvements (job timeouts, per-file flake retry, network-install
retries):

- detect-changes + ci.yml: token falls back to the built-in read-only
  github.token when AUTOFIX_BOT_PAT is empty. On main it uses the PAT
  (authoritative); on forks it uses github.token, which can read the
  public compare endpoint. (An input `default:` only applies on omission,
  not on an empty passed value — hence the explicit `|| github.token`.)
- lint ci-review + supply-chain mcp-catalog gates: restore the inline
  `gh pr view ... || true` label read with the github.token fallback,
  dropping the hard-failing retry "Fetch PR labels" step. Graceful
  degrade to "label absent" on an API blip, same as before #66373.

Same-repo enforcement is unchanged (byte-identical logic; the PAT is still
used there). Fork PRs classify correctly and the gates read labels via the
read-only token exactly as they did before the regression.
2026-07-17 15:52:59 -07:00
Teknium
597615ade4 fix(ci): make tests, workflows, and attribution reliable under load (#66373)
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory

The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.

New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.

- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
  merged with the directory at import time (directory wins). All
  existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
  conflicting reassignments (incl. against the legacy map), validates
  email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
  a legacy entry; failure message prints the exact add_contributor
  command. Also auto-resolves bare <login>@users.noreply.github.com
  emails is intentionally NOT added (kept id+login form only, matching
  previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
  merge precedence, CLI idempotency/conflict/validation, subprocess E2E.

* feat(ci): one-shot per-file flake retry in the parallel test runner

A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.

- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
  deterministic failure still exits 1; retries=0 restores old behavior.

This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.

* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s

These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.

* fix(ci): job timeouts everywhere + retries on all network installs

Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
  burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
  that lacked it: pip installs (deploy-site, skills-index), npm ci
  (deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
  test deps). Deterministic build steps (npm run build) deliberately
  NOT retried — split into separate steps so a real build failure fails
  fast instead of retrying 3x.

* docs(agents): document the file-retry flake policy

* fix(ci): curl retries on deploy hook + skills-index probe

* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile

From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
  run_id-suffixed keys — the cache never matched once, so LPT slicing
  always ran blind and unbalanced slices pushed heavy files toward the
  per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
  'gh pr view || true' turned an API blip into 'label absent' → false
  BLOCKING failure. Now 3x retry, and API failure is reported as an API
  failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
  silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
  so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
  consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
  curl --retry 3 (ADD cannot retry; checksums still enforced); npm
  --fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
  get continue-on-error so an artifact-service blip can't fail a green
  test slice.

* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list

- test_tui_gateway_server.py: session.create / non-eager session.resume
  arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
  test and fires into the NEXT test's _make_agent mock, racily
  corrupting captured state (the recurring session_resume shard
  failures). Replaced the per-test whack-a-mole stub with a module-wide
  autouse fixture; the 3 worker-lifecycle tests that genuinely need the
  deferred build opt back in via @pytest.mark.real_agent_prewarm (new
  marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
  live PROVIDER_REGISTRY instead of a hand-list that had drifted
  (missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
  tests failed on any machine with HF_TOKEN exported. E2E-verified with
  HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.

* test: de-flake 30 timing-sensitive test files for loaded CI runners

Root-cause fixes from the flake audit (session-DB mining + repo sweep):

Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
  sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
  replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
  unbounded blocking read (parent wedge now fails THIS test with a clear
  message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
  (the 1s partition window mid-interpreter-startup is how a child PID
  escaped the live-system guard in CI)

Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
  mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
  mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
  voice_cli_integration, docker_environment, session_store_lock_io,
  planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
  (joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
  10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
  setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
  5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
  iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
  0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
  compression fork-lock TTL 1s->3s (12 refresh chances per lease);
  compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)

* fix(tests): repair indentation from de-flake batch edit

* fix(tests): harden env isolation and replace remaining sleep-sync races

The full 42k-test run and complete npm check surfaced three more classes:

- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
  leaked into Python/TUI tests. Pin the default Honcho host in the
  hermetic fixture, isolate the one fallback test from ~/.honcho, and
  blank SSH_* around terminalSetup tests. This flipped 20 false failures
  back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
  time.sleep globally, then busy-polled with that same mocked sleep. Under
  full-suite load the poller could starve the writer. Each test now waits
  on an Event emitted by the exact flush/retry transition; 30/30 passed
  under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
  not fire before its assertion. A loaded runner descheduled the test for
  >500ms and both chunks arrived. Producer controls now gate second-chunk
  and completion transitions explicitly.

Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.

* refactor(ci): use gh bot pat, better retries

refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.

Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference

Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.

ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth

Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.

19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
  comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call

---------

Co-authored-by: ethernet <arilotter@gmail.com>
2026-07-17 20:55:24 +00:00
ethernet
f8ddf4fd86 feat(ci): semantic package-lock.json diff as an upserted PR comment (#65206)
git diff on a lockfile is unreadable: npm reorders entries, rewrites
integrity hashes, and moves packages between nesting levels, so a
one-line package.json bump produces a thousand-line textual diff.

scripts/ci/lockfile_diff.py instead parses the `packages` map out of
both versions of every tracked package-lock.json (via `git show`),
reduces each to {install path: version}, and set-diffs the maps —
reorder/hash churn vanishes, leaving only actual version movement
(added / removed / updated, with nested dedup copies tracked
separately).

The lockfile-diff workflow posts the result as a Markdown table in a
PR comment gated behind a hidden marker: subsequent pushes PATCH the
existing comment instead of stacking new ones, and a push that reverts
all lockfile changes updates the comment to say so. Advisory only —
never fails on findings; fork PRs (read-only token) degrade to a
warning.

Wired through the ci.yml orchestrator with a new npm_lock lane in
classify_changes.py (fails open on .github/ changes per the existing
contract).
2026-07-16 03:18:15 +00:00
ethernet
ef7aabd3d1 ci: add ci-reviewed label gate for CI-sensitive files 2026-07-16 01:42:02 +05:30
ethernet
7b3f3047ab feat(ci): run JS tests in CI, add npm run check in ws root 2026-07-13 17:22:17 -04:00
ethernet
f0cb049217 change(ci): migrate docker smoketests to real tests 2026-06-26 19:15:18 -07:00
ethernet
05c896cf52 ci: refactor paths & clones
ci: centralize path-gating behind single orchestrator + all-checks-pass
gate

Replace the scattered per-workflow detect-changes pattern with a single
ci.yml orchestrator that runs the classifier once, then conditionally
calls sub-workflows via workflow_call based on lane outputs. A final
all-checks-pass job (if: always()) aggregates all results so branch
protection only needs to require one check.

Changes:
- New .github/workflows/ci.yml orchestrator (detect + conditional calls
  + all-checks-pass gate)
- Extend classify_changes.py with scan/deps/mcp_catalog lanes, absorbing
  supply-chain-audit's internal changes job
- Update detect-changes/action.yml to expose the new lane outputs
- Convert all 10 PR-gated sub-workflows to workflow_call-only triggers,
  removing their push/pull_request triggers and per-step detect-changes
  guards (gating now happens at the orchestrator level)
- lint.yml + supply-chain-audit.yml receive event_name as a
workflow_call
  input to replace github.event_name (which is "workflow_call" inside
  called workflows)
- supply-chain-audit.yml: remove internal changes job + *-gate jobs
  (orchestrator handles gating, booleans arrive as inputs)
- contributor-check.yml: remove internal filter step
- Update test_classify_changes.py for 6-lane output + new supply-chain
  test cases
2026-06-23 09:30:50 -07:00
Brooklyn Nicholson
56b4ef74a6 ci: make dependency installs resilient to transient flakes
`npm ci` / `uv sync` / toolchain header fetches occasionally die on
transient network blips — e.g. node-pty's node-gyp fetching Node headers
(an undici assert) during the typecheck job's `npm ci`, which killed the job
before `tsc` ever ran. "Re-run and it goes green" is exactly what CI should
do itself.

- New reusable `.github/actions/retry` composite action wraps a command and
  retries on failure (3x / 10s, command passed via env so it can't inject).
  Applied to every PR-path network install: npm ci (typecheck, desktop
  build, docs site), uv sync (tests, e2e), uv tool install (lint),
  pip install (docs site).
- typecheck now runs `npm ci --ignore-scripts`: `tsc` needs only sources +
  type defs, so skipping install scripts drops node-pty's native rebuild
  (whose header fetch was the flake) and is faster. Validated locally — tsc
  passes for ui-tui, apps/shared, and apps/desktop with scripts skipped.
- ripgrep download uses `curl --retry`.

Docker (main-only) and the release/windows workflows are intentionally left
for a follow-up.
2026-06-23 09:30:50 -07:00