Commit Graph

252 Commits

Author SHA1 Message Date
Teknium
bfbb34bbec fix(update): Windows progress server hands out its URL only once it is serving
`Start-UiServer` printed the -SelfTestUi URL (and opened the browser window)
as soon as the TcpListener was bound, but the runspace that answers /progress
starts asynchronously — BeginInvoke returns before the pipeline is open and
the script block is JIT'd, which is seconds on a loaded runner. The kernel
accepted connections into the backlog during that gap and nobody answered
them. The self-test hit it three times (#90371 and two follow-ups each
widened a timeout instead of removing the race) and it just failed an
unrelated hermes_state.py PR (run 33591547099, two 5s stale-backlog
timeouts = red).

- windows.ps1: readiness handshake after BeginInvoke — one /progress
  round-trip must succeed (≤15s) before the server is returned; on failure
  tear the listener down and continue without UI. The URL now means
  "serving", not "bound". Also fixes the browser opening to a page that never
  loads on a slow machine.
- test: 1s per-attempt probe timeout so a single dead backlog socket cannot
  consume half the readiness budget.
- CI: new `desktop_updater` classifier lane. tests/test_desktop_update_windows_*.py
  spawn the real PowerShell script; the Windows-only job now runs them only
  when scripts/desktop-update/**, the Electron updater launcher, conftest,
  pyproject, or those tests change (push/dispatch fail open). A PR that
  never touched that surface cannot be failed by its process timing.
2026-09-02 00:44:19 -07:00
Teknium
ff7745fb0a ci: every ci.yaml run 0-jobbed since 24f5a60ed1 — e2e-desktop disable as bare if: false
The `${{ false && (...) }}` if-expression on the reusable-workflow
e2e-desktop job made GitHub's workflow parser fail at startup (annotation:
"An unexpected error has occurred"), so every ci.yaml run on main and
every PR dispatched 0 jobs. Discriminator on this branch: pre-24f5a60
blob → 26 jobs; paren-relocated && form → 0; bare if: false → 25.
Lane stays disabled; comment documents the re-enable line.
2026-09-01 21:18:24 -07:00
ethernet
24f5a60ed1 fix: re-disable e2e 2026-09-01 18:08:14 -04:00
ethernet
375ce8eee5 ci: block tracked paths that collide case-insensitively
Linux is case-sensitive; Windows and macOS are not. Two tracked paths
differing only by case (README.md vs readme.md, src/Foo.py vs SRC/foo.py)
land fine on Linux and silently break every clone on a case-insensitive
host — the filesystem holds one, so checkout fails or whichever wins
clobbers the other. Git won't stop the pair from landing; it only warns
at checkout time on a case-insensitive FS. This is the enforcement point.

Adds scripts/check-case-collisions.py (index scan keyed on casefolded
full paths) + an unconditional workflow_call job wired into ci.yaml and
the all-checks-pass gate — unconditional because a collision can ship in
any kind of PR (docs, JS, config), not just Python, so gating on a
language lane would be the same passive-rule trap the infographic check
closes. Tests in tests/scripts/test_case_collision_check.py build
collisions via git update-index --cacheinfo so they run on
case-insensitive filesystems too.
2026-09-01 15:37:57 -04:00
teknium1
9ce95929a2 ci: re-enable the Desktop E2E lane — harness root-fixed by #99671
The lane was disabled Aug 2 2026 (#76627) because the mock-backend
Electron window never got a title after the Aug 1 engines/npm churn
(#76499/#76562/#76575), failing every PR identically. #99671 fixed the
root cause: per-platform/layout Electron binary resolution in the e2e
harness (apps/desktop/e2e/electron-binary.ts). The suite is green again
on Node 26 + npm 12 — delete the temporary `false &&` guard and update
the stale comment block.

Fixes #76627
2026-09-01 12:04:29 -07:00
Teknium
5a8e8a6b87 fix(terminal): strict Linux-only gating for background-executor systemd scopes (#70716 follow-up)
Cross-platform hardening of @toprakeker's systemd cgroup isolation
(PR #71378, landed via #81264):

- Gate every scope-path branch on a new _IS_LINUX constant instead of
  'not _IS_WINDOWS', so macOS (and any other POSIX platform) provably
  never touches systemd code — no probe subprocess, no scope argv,
  byte-identical legacy spawn.
- Unit tests: darwin no-op guarantee (no probe exec, no scope argv build,
  legacy argv byte-identical, no unit recorded) and probe-returns-False
  off Linux.
- New live Windows E2E (tests/tools/test_process_registry_windows_live.py,
  wired into the on-demand windows-venv-e2e lane): real spawn_local on
  windows-latest asserting jobs run exactly as before — spawned, output
  captured, exit code correct, systemd path never reached even under
  faked gateway identity.

Refs #70716, #71378.
2026-09-01 02:32:53 -07:00
joaomarcos
c26f75baab fix(security): keep profile exports out of source and image contexts
Route automatic profile exports to a managed store instead of the current checkout, and enforce a CI/Docker boundary that rejects archive files before they can be published.
2026-09-01 01:00:23 -07:00
Teknium
18c9048070 ci(wine2e): include TestWindowsRuntimeSelfLock in the on-demand windows-latest lane 2026-08-31 12:50:29 -07:00
Teknium
7790c8f4c4 fix(telegram): bound stale-client cleanup and add Windows CLOSE-WAIT live probes (#87057)
Follow-ups on top of the salvaged commits from PR #87111 (@HexLab98) and
PR #87265 (@JoaoMarcos44):

- keep main's #92991 stall watchdog (150s progress-based) as the single
  steady-state liveness probe instead of adding a second overlapping one
- orphaned-client aclose() cleanup uses the wall-clock thread deadline and
  is tracked in _background_tasks so a wedged close can neither hang nor
  leak one task per reconnect attempt (from #87265's review findings)
- merge #87265's no-keepalive getUpdates pool (max_keepalive_connections=0)
  with #87111's TCP-keepalive socket options on all transports
- add tests/gateway/test_telegram_closewait_windows_live.py: live probes
  against a real half-closing HTTP server, skipif non-win32, wired into
  the on-demand windows-venv-e2e lane (wine2e/**)
2026-08-31 12:28:49 -07:00
Teknium
7acc65a399 test(windows): live E2E for the git trampoline self-heal on the wine2e lane
Real windows-latest coverage for the #88136 salvage: probes drive the
actual _git_is_trampoline/_locate_real_git/_ensure_non_trampoline_git
helpers against the runner's genuine Git-for-Windows install plus a real
fork-bomb-guard trampoline stand-in. Wired into the on-demand
windows-venv-e2e lane (wine2e/** pushes only).
2026-08-31 12:21:46 -07:00
Teknium
90e916efc9 fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):

- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
  false-positives), and an explicit start-time expectation is now honored
  on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
  retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
  caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
  scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
  require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
  unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
  ProcessRegistry._terminate_host_pid (previously unverified), and the
  session-close path now runs the same daemon identity verification as
  the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
  probes (real spawned processes, real psutil ancestry) wired into the
  on-demand windows-latest wine2e lane.

Fixes #98814
Fixes #89614
2026-08-31 10:41:54 -07:00
Gille
3e1629009a fix(install): reject incompatible system npm 2026-08-28 05:05:36 -07:00
Teknium
547f4c952d chore: remove single-use Windows live-E2E proof workflow (run 33157529194 green) 2026-08-28 02:59:17 -07:00
Teknium
2a4278209f ci: single-use Windows live-E2E proof workflow [proof do-not-merge] 2026-08-28 02:59:17 -07:00
Teknium
824f7e081a chore: remove Windows real-profile PROOF workflow + live test
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
2026-08-26 19:25:33 -07:00
Teknium
e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium
00d5632249 chore: remove Windows real-profile PROOF workflow + live test
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium
9e9e1b2245 feat(browser): consented auto-close of a running browser for Windows real-profile [proof workflow do-not-merge]
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.

- close_browser_holding_profile: graceful terminate → kill → poll until the
  cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
  unrelated same-name process on a different dir).
- Config key + docs admonition.

Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.

Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
2026-08-26 19:25:33 -07:00
Teknium
b73f78714a chore: remove Windows real-profile PROOF workflow + live tests
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium
52772de402 test(ci): run Windows diagnostic first + hard-bound the live test [do-not-merge]
Prior run hung 24min in the product-path test (snapshot_real_profile against a
locked profile blocks on Windows — itself a finding). Diagnostic now runs FIRST
(each strategy internally bounded, reports fast), live test second under a
faulthandler 150s dump-and-die so a hang can't burn the job. Job timeout 12min.
2026-08-26 19:25:33 -07:00
Teknium
92b95f16ce test(ci): proof workflow uses per-sha concurrency, no cancel-in-progress [do-not-merge]
Previous runs were auto-cancelling each other (ref-scoped group + cancel-in-progress). Per-sha group lets each proof run finish so the diagnostic actually reports.
2026-08-26 19:25:33 -07:00
Teknium
bd504bee6d test(browser): PROOF diag — probe which read strategy beats Chrome's Windows lock [do-not-merge]
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
2026-08-26 19:25:33 -07:00
Teknium
2ecb18c4a7 test(browser): PROOF — Windows live E2E for locked-DB real-profile copy [do-not-merge]
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.

This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
2026-08-26 19:25:33 -07:00
Teknium
cb45a813c3 chore: drop the on-demand Windows Rust lane from the PR — wine2e workflows run from the pushed branch and never need to live on main 2026-08-25 21:59:19 -07:00
Teknium
bc4ebcf3b0 test(bootstrap): pin the exit-2 self-marker heal contract + on-demand windows-latest Rust lane
The heal decision is extracted into should_heal_self_marker_refusal()
so the contract is testable: heal ONLY on exit 2 + a marker naming this
process. Five tests pin it — self-owned heals, foreign owner (real live
sibling process) never heals, missing/garbage marker never heals,
non-exit-2 never heals, and the full acquire -> refuse -> drop-claim ->
retry-precondition lifecycle with a real UpdateMarkerGuard.

windows-rust-e2e.yml mirrors the wine2e pattern: fires only on
wine2e-rust/** pushes, runs the crate's cargo test --lib on
windows-latest (the shipping platform). The permanent Linux lane stays
authoritative for the unix-gated pipe-drain fixtures.
2026-08-25 21:59:19 -07:00
ethernet
0012dd1e0c perf(ci): set python test workers to one for each core, from measurement
`run_tests.sh` defaults to twice the core count, and the value this branch
started with came from a rule of thumb of 1.5x cores plus a measurement on a
16-core machine. A sweep on the real runner disagrees with both.

Run 32549672063 on the 96-core runner (EPYC 7763, 377GB) timed the whole suite
at six worker counts, two repetitions for each. A warmup run came first, and
retries were off:

    workers   x cores   rep 1   rep 2   mean
       48       0.5x     138s    139s   138s
       96       1.0x     127s    126s   126s   <- fastest
      144       1.5x     130s    134s   132s
      192       2.0x     132s    133s   132s
      240       2.5x     140s    139s   140s
      288       3.0x     143s    142s   142s

One worker for each core wins. Both repetitions agree on the order.

The shape is the more useful result. The range is 126s to 142s across a 6x
range of worker counts. The suite has sufficient concurrency at this machine
size, so nothing above the core count buys anything. The remaining time
belongs to the slowest individual files and to the setup. A future gain must
come from those, and not from this number.

The sweep ran from a temporary workflow that this branch does not keep.
2026-08-22 02:25:12 -04:00
ethernet
10f99bc15e ci: run the work lanes on larger runners and merge the split jobs
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.

The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.

Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.

The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.

JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.

The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.

The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.

`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.

node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.

The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.

The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.

`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.

The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.

Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
  the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
  `ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
  returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
  parallel units and for a plain `npm run check`, in both directions. Against
  the 13-leg matrix the count is 13 to 11, and the whole difference is the
  three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
  the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
2026-08-22 02:25:12 -04:00
Teknium
aefcf4d10a ci(windows-venv-e2e): drop --timeout (pytest-timeout not in dev-only sync) 2026-08-21 19:11:55 -07:00
Teknium
f9aed7d7f6 test(windows): on-demand live venv-holder E2E lane + probe suite (#91277)
On-demand workflow (fires only on wine2e/** pushes, never on PRs/main)
that runs a live venv-holder E2E on windows-latest: real spawned
processes with Hermes argv shapes, real detection/classification/
message code against the live process table. Tests pin CORRECT behavior
for the cluster issues (#90778 mislabeling, #78089 long-path exemption,
#87594 ancestor-exclusion, #81774 serve premise), so unfixed bugs fail
on the runner — empirical premise-check before the consolidation fix.
2026-08-21 19:11:55 -07:00
Brooklyn Nicholson
4d2c546a53 fix(ci): drop --locked, the installer crate has no tracked lockfile
apps/bootstrap-installer/.gitignore excludes src-tauri/Cargo.lock — a
create-tauri-app scaffold default nobody revisited. With nothing tracked,
`--locked` fails outright ("cannot create the lock file ... because
--locked was passed") and the cache key hashed an absent file.

Keyed on Cargo.toml instead. The underlying gap — a signed installer that
re-resolves its whole dependency graph on every build, in a repo whose
pinning policy is otherwise strict — is noted in the workflow and left
for its own change rather than widening this one.
2026-08-20 12:03:00 -05:00
Brooklyn Nicholson
dd1e5b723d fix(ci): declare the rust lane on the detect job, and test that wiring
The lane shipped dead. `classify_changes.py` emitted `rust`, the composite
action re-exported it, and ci.yaml's `rust-tests` job gated on
`needs.detect.outputs.rust` — but the `detect` job never declared that
output, so the expression was the empty string and the job reported
"skipping" on the very PR that added it. GitHub does not error on a
reference to an output a job never declared, so nothing went red.

Adds the missing line plus the invariant that catches the whole class:
every `needs.detect.outputs.X` referenced by a job's `if` must be
declared by `detect`. Verified it fails with the line removed.

The related check — every lane reaching the composite action — is
separate on purpose: nix.yml and docker.yml own their triggers and
re-export different subsets, so `docker` and `nix` are legitimately not
ci.yaml detect outputs.
2026-08-20 11:46:30 -05:00
Brooklyn Nicholson
5caea5e501 ci: run cargo test for the bootstrap installer
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.

Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
2026-08-20 11:40:49 -05:00
Teknium
dc77f2c87f Revert "ci: force workflow re-parse after zero-job dispatch window"
This reverts commit ab173e26d2.
2026-08-19 02:15:02 -07:00
Teknium
43e67f1f0e ci: move orchestrator to ci.yaml — ci.yml workflow identity is wedged (0-job startup_failure on every dispatch; identical content dispatches fine under a new path, proven by probe PR #89894) 2026-08-19 02:13:28 -07:00
Teknium
ab173e26d2 ci: force workflow re-parse after zero-job dispatch window 2026-08-19 02:07:57 -07:00
Teknium
9162ea6db1 Revert "chore(ci): touch ci.yml to bust poisoned workflow-parse cache"
This reverts commit 0f73adb74f.
2026-08-19 01:22:51 -07:00
Teknium
0f73adb74f chore(ci): touch ci.yml to bust poisoned workflow-parse cache 2026-08-19 01:22:49 -07:00
ethernet
00c3872882 feat(ci): add nix flake check as unrequired job
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
2026-08-18 20:42:06 -04:00
ethernet
1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
Teknium
f30ea0eb9a Revert "chore: nudge ci.yml blob to bust poisoned workflow-parse cache"
This reverts commit 824409fbe511ba6c872cff2c972ce62a82ee5c4a.
2026-08-17 20:39:02 -07:00
Teknium
f6667e7bf9 chore: nudge ci.yml blob to bust poisoned workflow-parse cache 2026-08-17 20:39:02 -07:00
Teknium
7e77d9bd71 Revert "ci: cache-bust workflow re-parse"
This reverts commit b363038c1d39805062f484806f8acf80ffc4d770.
2026-08-17 02:40:59 -07:00
Teknium
077a7d536a ci: cache-bust workflow re-parse 2026-08-17 02:40:59 -07:00
Teknium
1ed94d2452 Revert "ci: touch ci.yml to force workflow re-parse"
This reverts commit a7ba91e2e58d337c899eef2760dcc67f44d208c5.
2026-08-17 02:11:27 -07:00
Teknium
e9e3291e71 ci: touch ci.yml to force workflow re-parse 2026-08-17 02:11:27 -07:00
kshitij
de0abc0617 ci(js-tests): drop dead electron download-cache path, skip redundant npm upgrade
Review findings on the caching commit:

- ~/.cache/electron was dead weight: with npm ci skipped on an exact
  cache hit, the download cache is never read (electron's unpacked
  binary lives in node_modules/electron/dist, inside the cached tree);
  it only inflated every saved archive by ~110MB.
- 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs
  (~5-15s each); now a no-op when the bundled npm is already 12.x,
  which also keeps the installed major aligned with the npm12
  cache-key tag.

yaml + actionlint pass.
2026-08-14 12:14:03 -04:00
kshitij
f56a9a1185 ci(js-tests): cache the installed node_modules tree, not just the npm tarball cache
Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.

Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:

- key includes runner.os + node26 + npm12 so a toolchain bump never
  reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
  skipped means nothing would repair it), so anything but an exact
  lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
  check jobs (with-scripts tree + ~/.cache/electron), which differ in
  postinstall artifacts

Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
2026-08-14 12:14:03 -04:00
copilot-swe-agent[bot]
0b50c8e48f ci: wrap uv python install in retry action on OS test lane
Co-authored-by: OutThisLife <770929+OutThisLife@users.noreply.github.com>
2026-08-13 13:18:49 -05:00
brooklyn!
9eab7a4473 fix(ci): stop running uv lock --check on PRs that can't touch the lockfile (#84675) 2026-08-12 13:30:01 -05:00
kshitij
f51aa6a9b5 fix(ci): repair red main — busy-mode test + missing checkout in skills-index workflows
Three separate reds on main. Two are fixed here; the third needs no code.

1. tests/gateway/test_multiplex_busy_input_mode.py (blocks every merge)

Fails "Python tests / Run tests slice 5/12" and therefore "All required
checks pass". Semantic merge conflict between two PRs merged ~1h apart:

  a31be480 fix(gateway): respect routed profile busy modes             (added the test)
  c8f235a1 feat(gateway): allow selective multiplex profile serving    (added the gate)

c8f235a1 taught _profile_name_for_source to reject a route whose target
profile is not in the served set (profiles_to_serve). Each PR was green on
its own base; neither ran against the other's merge result.

The test asserts a route to profile "research" resolves to that profile's
busy mode, but never patches profiles_to_serve — so it reads the runner's
REAL on-disk profiles. "research" is not among them, the route is rejected
before the busy-mode snapshot is consulted, and the assertion gets the
gateway default:

  WARNING gateway.run: Rejecting profile route 'research-chat':
                       target profile 'research' is not served
  AssertionError: assert 'interrupt' == 'steer'

Patch profiles_to_serve for the assertion — the same seam every sibling
test in tests/gateway/test_profile_resolution.py already patches
(test_route_inside_allowlist_resolves, test_route_outside_allowlist_rejects).

This also removes an ambient-state dependency: the test previously passed
or failed based on which profiles happened to exist on the machine running
it. Verified passing under an empty HERMES_HOME.

Test-only. The serving gate from c8f235a1 is correct and left intact.

2. Skills-index workflows: local action used without actions/checkout

check-freshness has failed on all 12 of its last 12 scheduled runs:

  ##[error]Can't find 'action.yml', 'action.yaml' or 'Dockerfile' under
  '.../.github/actions/get-app-token'. Did you forget to run
  actions/checkout before running your local action?

./.github/actions/get-app-token is a LOCAL composite action and cannot
resolve without the repo on disk. skills-index-freshness.yml had no
checkout step at all. The step is gated on `status != 'ok'`, so the
watchdog broke exactly when it was supposed to file its issue — the live
index is currently 521.4h stale (limit 26h) and nobody was told.

An audit of all workflows for this bug class found one more instance:
skills-index.yml's `trigger-deploy` job, which re-triggers the docs deploy
so a refreshed index reaches the live site. Its sibling `build-index` job
checks out; this one did not. That is plausibly why the index went stale
in the first place. Both are fixed; the audit now reports zero remaining
jobs that use a local action without a prior checkout.

Pinned to the same actions/checkout SHA used by the other 35 call sites.

3. "Publish inline E2E evidence" — no fix needed

Failed once at 13:33Z on a transient TLS error reaching api.github.com
("certificate is not valid for any names") while installing a gh
extension. The last 25 runs of that workflow are 25/25 success. Infra
blip, not a code defect.
2026-08-11 21:26:28 +05:30