Commit Graph

1830 Commits

Author SHA1 Message Date
ethernet
cd667debfa fix(install-e2e): windows transcripts were empty; player gets #zip= hash + one player per run
Windows transcripts were ZERO bytes: ts-prefix.ps1 formatted with
{0:D2}, but Floor() returns a double and the D specifier is
integer-only - it threw per line, and under the driver's relaxed EAP
every line errored into the void. {0:00} fixes it (custom numeric
format works on doubles). Reproduced the exact pipeline locally
(empty file + Format specifier invalid), verified the fix produces
prefixed merged stdout+stderr with exit code intact. That is also
why the log timeline never auto-synced: there was nothing in the
files to sync.

The GitHub artifact URL 307s to /suites/... server-side and strips
the ?zip= query param. The player now reads the zip URL from a
#zip= HASH param (client-side, survives the redirect) with ?zip=
as fallback; the hash path was verified in a real browser against a
real leg zip (auto-fetch + boot).

Per ethie's design, one player artifact for the whole run: new
leg-player job uploads playback.html (archive:false) before the
matrix legs, the report job needs it, and each ran cell gets TWO
links - 📼 to the player with #zip=<that leg's logs zip> and ⬇️ to
the raw zip. Per-leg player uploads removed from all three run
workflows.
2026-08-13 22:51:44 -04:00
ethernet
08d9f27773 fix(install-e2e): results chart 📼 links against real artifact names
Debugged against run 31635036702 (real jobs + artifacts replayed
through the renderer). Two naming realities the renderer ignored:

1. upload-artifact with archive:false IGNORES the name: input and
   registers the artifact under the FILE's basename - every leg's
   player artifact is called 'playback.html' (all the same blob). The
   renderer now uses any 'playback.html' artifact for the player half
   of the 📼 link instead of install-e2e-player-<leg_id>.
2. The posix arms append -<sha> to the logs artifact name at upload
   (install-e2e-logs-<leg_id>-<sha>), while windows does not. Match
   by prefix instead of exact name.

Also fixed the pass/fail summary counters, which compared cells with
=== against '&#x2705;' - the appended reel link made every ran cell
count as neither passed nor failed (the summary said 0 passed while
the table was full of checks).

Replayed against the real run: 31 passed, 25 failed, 149 skipped,
24 📼 links on ran cells (both outcomes), none on skips.
2026-08-13 14:41:04 -04:00
ethernet
e138cb555d feat(install-e2e): hook the leg player into the results table
Each leg uploads playback.html as a single-file artifact (archive:
false) before the driver runs, so it exists even on failure. The
results chart now links every leg that RAN (pass or fail, not skip)
to its player with ?zip= pointing at that leg's logs artifact.

Leg<->artifact mapping: the generator mints a leg_id per matrix entry
(sanitized matrix name, exported legId()), every run workflow names
its artifacts install-e2e-{player,logs}-<leg-id>, and the report job
feeds the run's artifact name->id list to the results renderer, which
rebuilds the leg id from the parsed job name. GitHub does not link
jobs to artifacts, so the deterministic name is the join key.

Empirical finding: GitHub artifact downloads are auth-gated (the
download URL 307s to /suites/... which is 404 anonymous), so a
locally-opened player page cannot fetch the zip cross-origin. The
player now degrades gracefully: ?zip= fetch failure renders a real
download link for the zip (a normal click carries the user's session)
plus a drag-and-drop / file-picker path, and no-param opens as a pure
drop target. Verified in a real browser against a real artifact URL.

Verified: generator emits leg_id, results renderer emits
✅/❌ [📼](...?zip=...) only on ran cells, npm run check PASS,
install tests 36/36, strict tsc PASS, actionlint x4 PASS.
2026-08-12 15:50:46 -04:00
ethernet
4578b5d0a6 ci(install-e2e): result chart says WHY a cell skipped
Skipped cells split into their reason: pre-desktop (a desktop-surface
method against a tag that predates apps/desktop) vs TODO (declared,
no driver arm yet). The report job passes pick-releases' annotated
tags into --format results; the shared methodNeedsDesktop() is the
same predicate the plan chart uses, so plan and results agree about
what pre-desktop means. Without --tags the renderer keeps the flat
skip label (backward compatible).

Verified against run 31579084845 real job list: 82 legs, 44 skips
labeled correctly.
2026-08-12 10:40:57 -04:00
ethernet
2b39b885d6 test(install-e2e): macos desktop-installer arm - the published dmg, driven for real
macos gains the desktop-installer@latest install method: the website's
Hermes-Setup.dmg (verified live), mounted with hdiutil and its app
binary run DIRECTLY - an open-launched app inherits none of the git
redirect env, so direct exec is what keeps the isolation honest while
staying the same binary and first-launch flow.

install-e2e-macos-run.yml takes the windows shape: one workflow, one
inner job per driver arm, native skips. Arm 1 delegates script installs
to the shared OS-agnostic run workflow; arm 2 stages, installs from the
dmg, and drives both app-update methods through launch-from-spec.mjs -
open-app-update launches the installed .app (the double-click surface,
env via Playwright), hermes-desktop-app-update captures the product's
own hermes desktop spawn. Both end on sha asserts, never version
strings.
2026-08-12 04:36:03 -04:00
ethernet
adf7d55f4b ci(install-e2e): windows composes install x update - one driver, one job
windows-desktop-gui-e2e.ps1 and windows-installer-script-e2e.ps1 fold
into tests/install/windows-e2e.ps1 with orthogonal -InstallMethod and
-Route axes: the install phase dispatches on one, the update phase on
the other, and shared workroot state carries how OLD landed - so any
implemented update method can follow any implemented install method.
Implementing a new pair is now a driver function plus a gate edit,
never a new job.

The run workflow collapses to ONE inner job whose if: is the
implemented-pairs table. Newly cheap pairs go live with the merge:
  desktop-installer@latest -> hermes-update / installer-script /
    installer-script+desktop / hermes-desktop-app-update
  installer-script(+desktop) -> hermes-desktop-app-update
  installer-script+desktop -> open-app-update (the -IncludeDesktop
    install registers real Start Menu / Desktop shortcuts)
Only desktop-installer@latest as an UPDATE method stays a declared
TODO. scripts/windows_e2e_harness.ps1 executes the parse/parameter/
dispatch checks under pwsh before any Windows runner spins up.
2026-08-12 04:30:12 -04:00
ethernet
0f903e14a3 test(install-e2e): hermes-desktop-app-update goes live on the script driver
Playwright must own the spawn (it needs the inspection pipe), but
hermes desktop is not just build+launch - stamp checks, integrity
gates, sandbox fixups, and a constructed child environment. So the
driver intercepts the product's own launch: a sitecustomize.py on
PYTHONPATH (opt-in via HERMES_E2E_CAPTURE_LAUNCH) wraps subprocess.run,
captures argv/cwd/env at the spawn site, and fakes success instead of
spawning; launch-from-spec.mjs then _electron.launch-es exactly that
spec and clicks Settings -> About -> Update now. Completion is product
state, not a Playwright event: the handoff result file or the checkout
reaching the expected sha (source installs write no result file).

Ships with the driver, so it works unchanged on every sampled OLD ref
- no product flag, no pre-flag fallback split. Both launch shapes are
matched (npm exec electron / packaged exe under apps/desktop/release);
npm BUILD calls pass through untouched. Exit 0 without a capture fails
the leg: a version that never reached its launch must not pass.

Probe-the-probe: scripts/launch_capture_probe.sh runs control rows
(no opt-in, non-launch argv) and both treatment shapes - all green
locally. Gate flips on the shared run workflow for linux/macos;
windows adopts the same path with the driver restructuring.
2026-08-12 04:20:56 -04:00
ethernet
239523414e test(install-e2e): installer-script+desktop is its own install and update method
The one-liner with its desktop stage opted in (--include-desktop /
-IncludeDesktop) is a real install kind, distinct on both sides:
on windows the stage builds Hermes.exe AND registers Start Menu /
Desktop shortcuts - a second path to a hand-launchable app - while
on linux/macos it builds into the checkout and registers no OS
entry point.

Declared on every OS and driven by both script drivers: the drivers
pass the flag through (hard failure if the ref predates it - the
tag-has-desktop gate already skips pre-desktop tags upstream) and
assert the built app exists under apps/desktop/release afterwards.
The run-workflow gates run +desktop pairs only on desktop-bearing
tags; app-update pairs from +desktop installs stay declared TODOs.
2026-08-12 04:01:16 -04:00
ethernet
1af2093663 test(install-e2e): split app-update into open-app-update + hermes-desktop-app-update
The desktop app has two launch paths, so app-update becomes two
methods. open-app-update starts the app from the OS entry point the
desktop installer created (the installed exe / the .app), so it exists
only where a desktop installer does. hermes-desktop-app-update starts
the app via hermes desktop, which every install method provides on
every OS that ships the desktop app - on linux it is the only app
surface, since no desktop installer or packaged artifact exists there.

Both variants are desktop-surface methods on every OS, so the
tag_has_desktop annotation moves from windows-only to every matrix
entry, install-e2e-run.yml grows the input, and the plan chart marks
pre-desktop cells on all OSes.

The windows GUI arm's implemented pair renames to open-app-update;
every other new combination is a declared TODO that natively skips.
2026-08-12 03:47:27 -04:00
ethernet
3cd6636471 fix(install-e2e): stray paren broke the generator - every node invocation died 2026-08-11 23:24:00 -04:00
ethernet
8e6f6d863c ci(install-e2e): result chart on the run summary - conclusions per combination x tag
generate-e2e-matrix.mjs grows --format results: reads the run's own
job list as NDJSON {name, conclusion} on stdin (per-leg conclusions
are NOT reachable through needs - a matrix job collapses to one
aggregate result) and re-renders the plan chart with each cell's
outcome. Legs are recognized by the exact name shape buildMatrices
mints, so unrelated jobs fall out; duplicate leg names (one windows
job per driver arm, only one runs) merge by significance - real
outcomes beat skips, failures beat successes. A final report job
(if: always, needs all three OS jobs) appends the chart to its step
summary via gh api with the default token.

Verified against two real runs: 31536931863 renders 11 passed / 0
failed / 54 skipped all-green; 31557865241 (the pre-EAP-fix run)
renders its 4 real failures + cancellations over the sibling arm's
skips.
2026-08-11 23:19:45 -04:00
ethernet
db969ce696 ci(install-e2e): result chart on the run summary - conclusions per combination x tag
generate-e2e-matrix.mjs grows --format results: reads the run's own
job list as NDJSON {name, conclusion} on stdin (per-leg conclusions
are NOT reachable through needs - a matrix job collapses to one
aggregate result) and re-renders the plan chart with each cell's
outcome. Legs are recognized by the exact name shape buildMatrices
mints, so unrelated jobs fall out; duplicate leg names (one windows
job per driver arm, only one runs) merge by significance - real
outcomes beat skips, failures beat successes. A final report job
(if: always, needs all three OS jobs) appends the chart to its step
summary via gh api with the default token.

Verified against two real runs: 31536931863 renders 11 passed / 0
failed / 54 skipped all-green; 31557865241 (the pre-EAP-fix run)
renders its 4 real failures + cancellations over the sibling arm's
skips.
2026-08-11 23:16:09 -04:00
ethernet
80a0a198dd ci(install-e2e): plan chart on the run summary - combination x tag markdown table
generate-e2e-matrix.mjs grows --format markdown: one row per
{os, install -> update} combination, one column per starting tag,
appended to GITHUB_STEP_SUMMARY by the expand job. Cells mark
dispatched legs; run-vs-grey stays the run workflows' call, so the
only special cell is pre-desktop (the one annotation the plan owns).
JSON mode unchanged.
2026-08-11 23:06:42 -04:00
ethernet
ea4cd375f8 ci(install-e2e): retire the bubblewrap sandbox - git redirect everywhere, macos legs live
The fake Internet (bubblewrap + slirp4netns + MITM proxy +
upload-pack shim, 883 lines across dev-sandbox.sh, stage2-run.sh,
proxy.py, ssh-shim.sh, openssl.cnf, install-update-e2e.sh) existed to
isolate install.sh's network. The GIT_CONFIG_GLOBAL insteadOf redirect
the windows driver introduced does the same job with a gitconfig file
and works on any OS, so:

* install-e2e-run.yml now runs tests/install/installer-script-e2e.sh
  directly on the bare runner - no sandbox deps, no userns sysctls -
  and takes a runner input;
* the macos matrix calls the SAME workflow on macos-latest, deleting
  install-e2e-macos-run.yml: installer-script -> installer-script /
  hermes-update flip from grey to live, app-update pairs stay TODO
  inside the shared gate;
* install.sh is no longer curl'd through a fake CA - each leg runs
  the copy from the ref a user of that version actually executed;
* scripts/dev-sandbox.sh becomes the minimal isolation sandbox from
  ab6b9492f (separate HERMES_HOME / Electron userData / app name,
  same CLI surface: --persistent, --from, --delete), keeping its
  .hermes-sandbox dir name so gitignore and docs hold;
* nix/sandbox.nix drops the bwrap/proxy closure and keeps only the
  Electron runtime LD_LIBRARY_PATH the desktop app needs.

Verified: nix build .#sandbox + smoke run (isolated HERMES_HOME
created, ephemeral cleanup), shellcheck/bash -n on both scripts,
actionlint on all three workflows, and the new driver ran the full
v0.20.2 -> HEAD hermes-update pass locally before this commit.
2026-08-11 21:20:42 -04:00
ethernet
6d7ba86963 ci(install-e2e): collapse the method vocabulary - 3 install ids, install+2 update ids
Per review the unions were overcomplicated. Install methods are now
just: installer-script (the platform one-liner - curl | bash on
linux/macos, irm | iex on windows), desktop-installer, and
packaged-app (declared, unused). Update methods are every install
method (re-run it over the existing install) plus hermes-update and
app-update. desktop-installer-rerun, desktop-app, curl-bash, and
irm-iex are gone as ids; the windows driver's ValidateSet, switch
arms, and both run workflows' gates renamed to match. tsc --checkJs
clean; generator output re-verified (4 linux / 16 windows / 6 macos
legs for 2 tags).
2026-08-11 17:14:17 -04:00
ethernet
17dea59026 ci(install-e2e): generator drops runtime validation for jsdoc type unions
Per review: the method/version vocabulary is now closed TYPE unions
(@ts-check + jsdoc typedefs - InstallerVersion, InstallMethod,
UpdateMethod - checked with tsc --checkJs, which rejects a SPEC entry
outside the unions; verified by corrupting a copy: 6 errors) instead
of runtime KNOWN_METHODS/ALLOWED_VERSIONS sets. validateEntry and
routeWants are deleted with all the paranoia: the generator always
emits every OS matrix and the dispatch route filter moved to plain
job-level ifs in install-e2e.yml, where the OS jobs already live.
secondUpdate is typed never[] so declaring one is a type error until
a leg implements it. Anything types cannot catch is self-evident on
the next CI run.
2026-08-11 17:11:07 -04:00
ethernet
1c60bdc30d ci(install-e2e): back to the generator - workflows stay generic, names carry everything
Revert the hardcoded 16-job experiment: the combination spec belongs
in scripts/sandbox/generate-e2e-matrix.mjs (restored), not copy-pasted
YAML blocks. What survives from the experiment:

* leg names carry everything - 'os: install -> update (tag -> HEAD)' -
  generated per entry, since slash-joined names are all the graph
  renders;
* pick-releases annotates each tag ({ref, desktop}) and the generator
  threads tag_has_desktop onto windows entries, so the windows run
  workflow still gates pre-desktop tags without a probe job;
* the per-OS run workflows are untouched: single job, static 'e2e'
  name, native skip gates own all capability knowledge.

install-e2e.yml is one generate job + three per-OS matrix fanouts.
Generator shape (4/16/12 legs for 2 tags), annotation threading, and
all six error paths verified; all four workflows pass actionlint;
driver parses clean pure-ASCII.
2026-08-11 16:54:54 -04:00
ethernet
6515f8132a ci(install-e2e): hardcode combination jobs in the primary workflow - one matrix box per combo
GitHub only draws matrix boxes for the PRIMARY workflow's matrices;
everything inside a called workflow flattens into slash-joined names.
The generator + per-tag sub-workflow therefore bought no structure in
the graph and hid the support matrix in a script.

Invert it: install-e2e.yml now declares one job per {os,
install-method -> update-method} combination (2 linux + 8 windows +
6 macos - same 16 the generator produced, verified by inventory
before/after), each a matrix over the picked release tags. The graph
now renders one titled box per combination whose legs read
'... from vX' - the tag axis inside the combo axis. The per-OS run
workflows are unchanged: they own capability knowledge and natively
skip unimplemented method pairs and pre-desktop tags.

install-e2e-tag.yml and generate-e2e-matrix.mjs are deleted; adding a
method is now adding one job block here, implementing one is flipping
the run workflow's gate.
2026-08-11 16:35:03 -04:00
ethernet
6d94678e62 ci(install-e2e): jobs own their skips - generator is pure expansion
Remove all capability knowledge from the combination generator: no
IMPLEMENTED table, no skipped matrix, no per-OS special cases. It now
only declares and expands - every {os, install-method, update-method}
combination is dispatched to its OS's run workflow, and each run
workflow natively skips (grey, job-level if on the method inputs) the
pairs its driver cannot run yet:

* install-e2e-run.yml gains install-method/update-method inputs,
  gates on the supported pairs (curl-bash -> hermes-update/curl-bash),
  and maps the method id to the sandbox script's --route internally;
* install-e2e-macos-run.yml is new - all pairs skip until a macOS
  driver exists, and implementing one flips its job-level if;
* install-e2e-windows-run.yml already worked this way;
* install-e2e-skip.yml is deleted - nothing special-cases macOS
  anymore, so the tag workflow is three identical OS fanouts.

Structure is now uniformly matrix(tag) -> matrix(combination) ->
run-or-skip, with capability knowledge living only next to each
driver. Generator shape/route filters/error paths re-verified; all
five workflows pass actionlint.
2026-08-11 16:17:48 -04:00
ethernet
83d009ae82 ci(install-e2e): nest the graph per starting tag; skips live where the knowledge lives
Two structural changes to the combination fanout:

1. Tags become the OUTER axis, as a sub-graph per starting version:
   install-e2e.yml fans a plain matrix over the picked tags into a new
   per-tag reusable workflow (install-e2e-tag.yml), which runs the
   combination generator for that one tag and fans out one job per
   {os, install-method, update-method}. The Actions graph now reads
   'from vX -> windows: install -> update' per leg. Nothing is
   hardcoded in the workflows: the tag workflow calls the generator
   itself.

2. Native skips move to the point that owns the capability knowledge:
   macOS combos (no driving workflow exists) grey out in the tag
   workflow via install-e2e-skip.yml, untouched by the tag axis; ALL
   windows combos dispatch to install-e2e-windows-run.yml, which takes
   install-method/update-method inputs and natively skips the pairs
   its driver cannot run yet - so implementing a windows method is a
   change in the run workflow + driver only. The driver's -Route ids
   now match the generator's method ids verbatim.

Generator output shape, route filters, and all error paths re-verified
locally; all four workflows pass actionlint; driver re-parses clean
pure-ASCII.
2026-08-11 16:01:36 -04:00
ethernet
1bf0a1c95b ci(install-e2e): generate every install/update combination - one job per combo
Replace the hand-enumerated update/installer/windows-desktop jobs with
a support-matrix generator (scripts/sandbox/generate-e2e-matrix.mjs).
The spec declares every {os, install-method, update-method} combination
a user could be on; generate-matrix expands it and fans out ONE JOB PER
COMBINATION:

* linux combos (curl-bash install x hermes-update/curl-bash rerun)
  drive install-e2e-run.yml, still multiplied by the sampled release
  tags from pick-releases;
* the windows combo (desktop-installer@latest -> desktop-app) drives
  install-e2e-windows-run.yml - the real GUI flow;
* every declared-but-unimplemented combo (all of macOS, the remaining
  windows methods) becomes its own visible skipped job, so the
  coverage gap is enumerable from the Checks tab and implementing one
  is a one-line move into IMPLEMENTED.

Strictness carried into the generator: method ids validate against a
closed set, installer 'versions' arrays only allow 'latest' until a
versioned archive exists, secondUpdate must stay empty until a chained
second-update leg is implemented, and unknown spec keys/routes throw.
Expansion, route filtering, empty-matrix gating, and all seven error
paths verified locally; the dispatch route choice keeps its exact
previous semantics (all/both/update/installer/windows-desktop).
2026-08-11 15:21:24 -04:00
ethernet
e3d22b5b29 fix(update): drain uv/pip stderr - undrained pipe deadlocked windows desktop updates
Two consecutive Windows E2E runs hung inside dependency install until
the job timeout: 31449642122 65 minutes in uv pip install .[all],
31453853006 43 minutes in the SQLite-repair uv sync (RUST_LOG=uv=debug
made THAT run hang earlier and with more stderr - the tell).

Root cause: scripts/desktop-update.ps1 redirects both child pipes but
only pumps stdout while the child runs; stderr is ReadToEnd()'d after
exit. uv and pip write progress to stderr. Once that pipe hits the
~64KB buffer, uv blocks on write, hermes update blocks on uv, the
hand-off blocks on hermes update: deadlock. Slower stderr producers
survive by finishing before the buffer fills, which is why the linux
sandbox never sees this.

Fix both sides of the class:
- managed_uv.py candidate sync + main.py _run_install_with_heartbeat:
  merge stderr into stdout (the pipe that IS drained). This arm heals
  EXISTING installs, whose old hand-off script drives the NEW python
  after the git reset.
- desktop-update.ps1: drain stderr concurrently via ReadToEndAsync so
  future bases never block regardless of what a child writes there.

_run_logged_subprocess and _run_npm_install_deterministic already
merge or capture both pipes; the two fixed sites were the only update-
path spawns that redirect stderr without draining it live.
2026-08-11 00:03:47 -04:00
ethernet
8789cf9f0c fix(sec): patch the npm advisories main left open
Main (7537de9e7) moved most of the vulnerable locked versions, but some
fixes live only in the lockfiles and some advisories stayed open. This
commit closes the rest:

website/package.json gets durable overrides for js-yaml 4.3.1,
dompurify 3.4.13, mermaid 11.16.1, and tar 7.5.22. The root workspace
gets the same tar override, which moves the tar 6.2.1 copies under
get-windows and @mapbox/node-pre-gyp past twelve open advisories.
Without an override, a reinstall can pull an old transitive copy back
in.

image-size <=2.0.2 has two infinite-loop DoS advisories and no fixed
release upstream. An override points it at @nous-research/image-size
2.0.3, our maintained fork of the real repo. The OSV scanner resolves
the aliased fork cleanly, so no ignore entries are needed.

The photon sidecar moves @opentelemetry/core to 2.10.0. The
whatsapp-bridge gets a body-parser 1.20.6 override, so the lockfile-only
fix from main cannot regress on reinstall.

website/.npmrc gets matching min-release-age exclusions for the fix
releases that are less than two weeks old.

electron stays at 40.10.2. The 41.x fix for GHSA-9f4c-93c8-jc8g brings
back the install failure that bb8280b75 reverted: install.js in 40.10.3+
extracts with an MSVC native binding, which fails on Windows machines
without the VC++ Redistributable. Upstream tracks this in
electron/electron#52481, with no fix released.
2026-08-10 13:49:37 -04:00
ethernet
d5ddd442d7 fix(ci): review comment poller deadlocked on its own run
The poller job set GITHUB_RUN_ID in env: to point at the CI run.
The Actions runner sets the GITHUB_* defaults itself and ignores
the override. Thus the poller read its own run id and watched
itself. Its own run stays in_progress while the poller runs, so
runs_all_completed() was never true. The comment froze at
'waiting for jobs to start' and the job burned its full 3000s
timeout on every PR.

Rename the variable to CI_RUN_ID. Also drop the GITHUB_REPOSITORY
override — it was a no-op for the same reason, and the runner
default already holds the correct value.
2026-08-10 13:48:21 -04:00
ethernet
8359e760be fix(ci): don't report all-good before jobs start
The live comment poller inferred completion from the job list. An empty
job list looks the same as a finished run: GitHub has not spawned the
jobs yet, so nothing is pending, and the poller posted a final
"all good!" comment and exited.

The run status is now the authoritative signal. collect_run_jobs()
returns whether the CI run and every watched sibling run report
status=completed, and the loop exits only when no job is pending AND
all runs are complete. While a run is still queued or in progress with
no visible jobs, the comment shows "waiting for jobs to start" instead
of a final banner.
2026-08-09 22:24:19 -04:00
ethernet
641d254db4 ci: add macos and windows test lanes for the os-marked tests
the markers from the previous commit skip off-host. without a host to
run them on, every marked test is a silent skip. this commit adds the
hosts.

- tests-os.yml runs -m macos_only on macos-latest and -m windows_only
  on windows-latest. ci.yml requires both lanes in all-checks-pass.
- a lane fails on pytest exit code 5 (zero tests selected). a renamed
  marker cannot produce a green job that ran nothing.
- each lane repeats 'not integration' because a command-line -m
  replaces the addopts filter.
- scripts/ci/list_os_marked_tests.py selects which files each lane
  imports. -m filters after collection, and collection imports every
  module. without this helper, one unrelated ImportError on the
  foreign host fails a job whose own tests passed. the helper exits
  non-zero when a marker matches no file, and writes bytes with
  explicit lf so windows crlf translation cannot corrupt the bash
  file list. it has its own tests in tests/ci/.
- the local runner now reports the skipped count and prints a note:
  macos_only/windows_only tests were skipped on this host, and this
  ci lane runs them. a green local run on linux no longer reads as
  coverage of the other hosts.
- the runner default job count is now #cpu, not #cpu*2.
2026-08-09 22:09:49 -04:00
ethernet
7aecab56db ci: move the review comment and the image build out of the CI run
The CI run stayed in progress until its last job ended. Two advisory jobs
set that time: the review-comment poller (40 minutes) and the Docker image
build (45 minutes). Neither job was required to merge.

GitHub refuses `gh run rerun` on a run that is in progress. Thus a reviewer
who added the `ci-reviewed` label had to wait for the two slow jobs, and
label-rerun.yml carried a 2100-second wait loop for this reason. The fast
required jobs were ready long before.

Each slow job now runs in its own workflow:

- docker.yml owns its `pull_request` trigger and does its own change
  detection. The new `detect` job runs the same composite action with the
  same condition that ci.yml applied, so a tests-only PR still skips the
  build. The `workflow_call` trigger is gone.
- ci-review-comment.yml starts on `workflow_run` when CI starts. It reads
  the workflow and the scripts from the default branch, which is the trust
  boundary that the old job got from its `ref: default_branch` checkout.

The poller reads job results through the API, so it can report on a run
that it does not belong to. `WATCH_WORKFLOWS` names sibling workflows for
the same commit, and `select_watched_runs` keeps the newest run for each
name. Thus the comment still shows the Docker results. The list is
newline-separated, because a workflow name can contain a comma.

The poller always exits 0 now. It reports on the CI run from a different
run, so a failed CI job is not a failure of the poller. The CI run has its
own gate for that.

Also correct a parse error in label-rerun.yml. STATUS came from the already
truncated RUN_ID, so its value was the run id and never "completed". Thus
the wait branch always ran.

ci.yml no longer needs `packages: write`, because the image build has left.
2026-08-09 15:40:04 -04:00
Teknium
952f44f841 fix(desktop): focus the update progress window, then hand focus to the relaunched Desktop
Two focus polish items from the first fully-working hand-off run
(ryanc, 2026-08-09):

1. The progress window came up backgrounded: the script is spawned via
   `cmd start /min`, and Form.Show() + TopMost keeps it above other
   windows without ACTIVATING it. Claim activation explicitly
   (Form.Activate + SetForegroundWindow) right after Show.

2. The relaunched Desktop came up behind whatever the user had focused:
   a WMI-spawned process starts unfocused and cannot take foreground by
   itself. Since the hand-off owns foreground while its progress window
   is up, delegate it: AllowSetForegroundWindow(new pid), poll up to 20s
   for Electron's MainWindowHandle, then ShowWindow(SW_RESTORE) +
   SetForegroundWindow. Best-effort at every step -- a focus failure
   never affects the update result.

Sequence on success: progress window foreground during the update ->
window closes -> freshly relaunched Hermes.exe takes foreground.

Verified live on the incident machine: Add-Type shim compiles under
PS 5.1; WMI spawn + AllowSetForegroundWindow + MainWindowHandle poll +
ShowWindow all execute against a real spawned window. (In the bg test
shell SetForegroundWindow returns False by OS design -- only the
current foreground owner may delegate; the real flow's TopMost progress
window IS that owner.) PS parse clean, check-windows-footguns clean.
2026-08-09 02:01:57 -07:00
Teknium
36eda6112b fix(desktop): detach relaunched Desktop from the hand-off console + UTF-8 child streams
First real-world run of the #82328/#82366 hand-off (2026-08-09, ryanc)
surfaced two defects:

1. The console window never closes after the update finishes -- and
   closing it manually KILLS the freshly relaunched GUI. Root cause:
   Start-DesktopRelaunch spawned Hermes.exe as a child of the console
   PowerShell. Electron/Chromium calls AttachConsole(ATTACH_PARENT_
   PROCESS) at boot, so the new Desktop latched onto the hand-off's
   console: the console can't close while an attached process lives,
   and closing it takes the attached GUI down with it. Fix: create the
   process via WMI (Win32_Process.Create) -- parent becomes WmiPrvSE,
   no console to inherit or attach, same detachment explorer.exe gives
   a normal launch. Start-Process fallback retained (tethered Desktop
   beats no Desktop).

2. Both the console and the progress box render hermes update's UTF-8
   glyphs (checkmarks, arrows) as mojibake. PS 5.1 defaults redirected
   child streams to the OEM codepage. Fix: StandardOutput/ErrorEncoding
   = UTF8 on the child, PYTHONIOENCODING/PYTHONUTF8 so Python emits
   UTF-8, and [Console]::OutputEncoding = UTF8 for our own echo.

Verified live on the incident machine: WMI-created process parents to
WmiPrvSE.exe (not the shell); UTF-8 glyph round-trip through the exact
ProcessStartInfo shape reads back byte-correct (15/15 chars). PS 5.1
parse clean, check-windows-footguns clean.
2026-08-09 01:43:04 -07:00
Teknium
6495ef82f7 fix(desktop): hand-off hardening - fail-closed gates, truthful completion, progress UI, result surfacing
Review feedback on the #82328/#82366 hand-off, all four points plus the
missing progress GUI:

1. FAIL CLOSED. Both preflight gates aborted-open: a Desktop still alive
   after 30s proceeded anyway, and a shim locked after 20s proceeded
   with --force - both mutate a potentially locked install (the exact
   Access-denied brick class). Now: desktop-alive -> exit 4, nothing
   changed; shim-locked -> exit 5, nothing changed. Both relaunch the
   Desktop so the user is never stranded.

2. TRUTHFUL COMPLETION. `hermes update` treats a Desktop GUI build
   failure as non-fatal (warns, exits 0) - correct for CLI use, a lie
   for a Desktop-driven update that then relaunches the OLD exe as
   "success". The script now detects the warning in the update output,
   retries the build once (`hermes desktop --force-build --build-only`),
   and exits 6 with an honest message when it still fails.

3. MARKER OWNERSHIP. Cleanup now removes the marker only while OUR pid
   still owns it - a handoff partner that rewrote the marker keeps its
   claim (same rule as UpdateLock.release).

4. RESULT SURFACING. The script writes .hermes-update-result.json on
   every exit path (ok, exit_code, message, branch, finished_at). New
   electron/handoff-result.ts consumes it exactly once at the boot
   update-gate: success logs, failure shows a real dialog pointing at
   desktop-update-handoff.log. Stale (>30min) and malformed results are
   consumed silently. Previously a failed detached update was
   indistinguishable from "nothing happened" - the exact live report
   that triggered this work.

5. PROGRESS UI. The old Tauri updater showed a window; the script ran
   in a hidden console with zero feedback. It now shows a WinForms
   progress window (marquee bar + streaming log) pumped via DoEvents
   during the update; -NoUi keeps tests/headless sessions clean, and a
   WinForms-unavailable session degrades to log-only.

Also: subprocess execution moved from Start-Process (ExitCode
unreliably $null under PS 5.1 even with the Handle workaround -
observed live: happy path reported "failed (exit )") to
System.Diagnostics.Process with synchronous stdout pumping, which
keeps the UI alive and the exit code real.

E2E on a real Windows box, sandbox HERMES_HOME + compiled fake
hermes.exe, all five paths:
- happy: exit 0, result {ok:true, "Update complete."}
- shim held open via O_RDWR: exit 5, nothing mutated, honest result
- desktop pid alive (60s ping child): exit 4 after the 30s gate
- update exits 0 printing "Desktop build failed" + rebuild fails:
  exit 6, result names the stale build and the retry command
- foreign-owned marker: overwritten by step-0 claim, removed as owner;
  ownership check verified in the cleanup path
vitest 18/18 (5 new handoff-result tests), typecheck 3 projects clean,
eslint clean, PS 5.1 parse + footguns + ASCII-only clean.

Remaining known gap (deliberate): the full click-to-relaunch lifecycle
through a REAL Desktop build still needs one live Windows verification
after this lands - tracked in the PR body.
2026-08-09 01:26:14 -07:00
Teknium
3b08a0f9b5 fix(desktop): give the update hand-off script its own console - a detached hidden powershell dies before -File runs
Live failure on the first real use of #82328 (2026-08-09): clicking
Update closed the Desktop with "an updater will happen", then nothing.
desktop.log showed `launched repo hand-off script`, but
desktop-update-handoff.log was never created - PowerShell exited 0
without executing a single line.

Root cause, isolated by spawning the exact production shape against a
sandbox HERMES_HOME: `spawn('powershell', [..., '-File', script],
{ detached: true, stdio: 'ignore', windowsHide: true })` kills
powershell.exe during console-subsystem init, before -File processing.
Variant matrix: plain pipes -> runs; hide only -> runs; detached only ->
runs; detached+hide -> exits 0, script never starts. Unit tests and
foreground invocations can't see this class of bug.

Fix: wrapHandoffForDetachedConsole() routes the invocation through
`cmd /d /s /c start "" /min powershell ...` - `start` allocates the
script its own minimized console and fully detaches it; the cmd wrapper
exits immediately. Verified the wrapped form survives the full
detached+hidden production spawn.

Knock-on: child.pid is now the short-lived wrapper, not the script, so
the Electron-side marker pre-write can't represent the script. The
script now claims the update marker itself as step 0 (its own $PID,
byte-exact "<pid>\n<ts>\n" via WriteAllText - Set-Content emits CRLF
and would break the three readers' framing). The Electron pre-write is
kept as a bridge for the spawn window: the script overwrites it, and if
the script never starts the wrapper's dead pid reads as stale and
self-deletes (no wedge). `hermes update` adopts the script's claim via
update_lock.py's process-ancestry rule, unchanged.

E2E in exact production shape (cmd start wrapper, detached, hidden,
parent exits 1.5s after spawn) against a sandbox HERMES_HOME with a
compiled fake hermes.exe: script ran, claimed marker with its own pid
(fake observed "<script-pid>|<ts>|" LF-framed DURING the update),
desktop-pid wait worked, update invoked with correct argv, marker
removed on completion. vitest 13/13 (new wrapper-shape test), 3-project
typecheck clean, eslint clean, PS 5.1 parse + windows-footguns clean.
2026-08-09 01:26:14 -07:00
Teknium
92be912d73 feat(desktop): repo-owned Windows update hand-off script - stop depending on the frozen hermes-setup binary
The Desktop's Update button hands off to the staged Tauri binary
(HERMES_HOME/hermes-setup.exe). That binary has no self-update path
(copy_self_to_hermes_home no-ops during --update), so every updater-side
fix only reaches users when a new installer is built, signed, and
published. In practice the published binary lags main by months and
users hit long-fixed bugs on every GUI update: the 2026-08-09 incident
chain was four distinct failures (stale install.ps1 cache resolver
pre-#67369, marker adoption pre-#74782, straggler teardown) all caused
by a June 4 binary running against an August repo.

This inverts ownership: scripts/desktop-update.ps1 lives in the repo
checkout, so every `hermes update` refreshes the code that drives the
NEXT update. Only PowerShell itself - an OS component - stays frozen.

Desktop side (apps/desktop/electron):
- resolveUpdateScriptHandoff() (updater-process.ts): returns the spawn
  recipe when scripts/desktop-update.ps1 exists in the checkout;
  Windows-only (POSIX updates in place via applyUpdatesPosixInApp);
  null on old checkouts -> caller falls back to the staged binary path
  completely unchanged.
- applyUpdates() prefers the script hand-off. The marker pre-write is
  ALWAYS safe on this path - no stagedUpdaterSupportsPrewrittenMarker()
  mtime heuristics - because hermes_cli/update_lock.py's UpdateLock
  adopts a live marker held by a process ANCESTOR, and the script is
  the `hermes update` child's parent. This closes the unguarded
  marker-gap window that pre-#74782 binaries force today (the 23:56
  failure in the incident: 'skipping marker pre-write: staged updater
  predates self-adopt' -> renderer respawned a backend into the gap ->
  update refused).
- CLI-installed users (no staged binary) now get the script hand-off
  too instead of the manual `hermes update` card, when the script
  exists.

Script (scripts/desktop-update.ps1): waits for the Desktop pid to exit
(bounded 30s), waits for the venv shim to unlock (mirrors the Rust
is_locked probe, bounded 20s), runs `hermes update --yes --gateway
--force --branch <ref>` from the CURRENT checkout with one retry for
the update-boundary class (skipped for exit 2), removes the marker on
every exit path, relaunches the Desktop. ASCII-only (the #67193
lesson), logs to logs/desktop-update-handoff.log.

Verification (real Windows box):
- apps/desktop: typecheck (3 projects) clean, eslint clean, vitest
  updater-process.test.ts 12/12 (3 new resolver tests).
- Script E2E against a sandbox HERMES_HOME with a compiled fake
  hermes.exe: correct argv (update --yes --gateway --force --branch
  main), stale marker removed, exit code propagated (0 and 1 paths),
  retry-once fires exactly once on failure, PS 5.1 parse + windows
  footguns check clean.
- Contract E2E with the real UpdateLock: ancestor-owned marker adopted
  (True), left in place on release, foreign live holder still refused.
2026-08-09 00:27:06 -07:00
Teknium
1d45e62f30 fix(install): replay npm's debug log into the bootstrap stream on failure
On Windows npm prints only a terse summary on failure; the actual cause
(postinstall stderr like Electron's install.js, network traces, EBUSY
retries) lives in npm-cache\_logs\<ts>-debug-0.log, which never reached
the Tauri bootstrap log. Field report: a fresh-VM desktop install died
with 'npm error command node install.js' and zero actionable detail.

Adds Write-NpmDebugLogTail: locates the debug log from npm's 'A complete
log of this run' line (fallback: newest _logs/*-debug-*.log under 'npm
config get cache') and replays its last 200 lines through our output
stream, which the bootstrap installer's streaming sink captures.

Wired at all four npm failure sites: desktop workspace npm ci/install,
_Run-NpmInstall (browser tools), Install-AgentBrowser (--silent global
install), and the desktop 'npm run pack' build step.
2026-08-08 22:55:00 -07:00
Gille
a09124cea5 fix(install): stop managed runtime child trees on Windows 2026-08-08 19:46:35 -07:00
Teknium
7537de9e74 fix(deps): patch 31 known CVEs across Python and npm lockfiles
OSV weekly scan reported 50 known vulnerabilities in pinned deps.
This bumps everything with a released, semver-compatible fix:

Python (uv.lock):
- aiohttp 3.14.1 -> 3.14.3 (GHSA-cq5v-8q36-5273, GHSA-mfx4-hv73-q22v,
  GHSA-mq44-7p77-q5h7)
- h2 4.3.0 -> 4.4.1 (CVE-2026-71554 request smuggling; exclude-newer
  exception documented in pyproject, remove after 2026-08-17)

npm (root workspace):
- brace-expansion 5.0.8 -> 5.0.9, undici 6.27->6.28 / 7.28->7.29,
  js-yaml 4.3.1, nanoid 3.3.17/3.3.18, ip-address 10.4.0,
  mermaid 11.16.1 + dompurify 3.4.13 (root overrides so the
  streamdown transitive copy is pinned too)
- electron 40.10.2 -> 40.10.6 (GHSA-r4w5-6pfg-jxp5; the 41.x major
  for GHSA-9f4c-93c8-jc8g is deferred to its own PR)

npm (website): mermaid, dompurify, js-yaml, nanoid, fast-uri 3.1.5,
postcss 8.5.23, undici 7.29.0
npm (photon sidecar): @opentelemetry/core 2.8.0 via override, undici
npm (whatsapp-bridge): body-parser 1.20.6

min-release-age excludes added to .npmrc/website/.npmrc for the
sub-2wk CVE-fix releases, each with a removal date.

Remaining findings are blocked upstream: cryptography <49 cap
(alibabacloud-tea-openapi), image-size (no fixed release), tar 6.x
transitive majors, electron 41.

Local rescan: 50 -> 19 known vulns, 0 introduced.
2026-08-08 14:06:48 -07:00
Jeff Watts
298ef06458 fix(tests): Windows-aware path-list split and UTF-8 progress output in parallel runner
Two Windows bugs in scripts/run_tests_parallel.py:

- --files/--paths/HERMES_TEST_PATHS were split on ':', which shreds
  absolute Windows paths at the drive letter ('C:\repo\tests' ->
  ['C', '\repo\tests']): the drive letter became a phantom discovery
  root and the rooted remainder only resolved by WindowsPath
  re-anchoring it onto repo_root's drive. New _split_pathspec() keeps
  drive-letter colons glued to their path and accepts ';' (os.pathsep)
  on Windows, while ':'-joined lists (CI generate job) keep working.

- With piped stdout (CI, subprocess capture) Windows encodes the
  runner's output as the ANSI code page, so printing the per-file
  progress glyphs raised UnicodeEncodeError inside the executor
  done-callback and every progress line was silently lost -- which is
  also why test_bare_value_flag_keeps_its_value failed on win32 (no
  '1[check]' line, and the summary says '1 tests passed', which does not
  contain '1 passed'). The runner now reconfigures its own
  stdout/stderr to UTF-8 on Windows, and the tests decode the captured
  output as UTF-8.

Adds regression tests: os.pathsep-joined absolute roots (all
platforms) and no-phantom-drive-root (win32).

Fixes #57149
2026-08-08 12:33:19 -07:00
Teknium
7b1f02377f feat(lint): close the fdopen + chained-call gaps in the encoding footgun gate
ruff PLW1514 (already enforced repo-wide via the blocking lint step)
covers open()/Path.open()/read_text()/write_text() but NOT os.fdopen —
the exact hole the AlexFucuson9 sweep PRs (#56033 #56940 #65565) kept
patching by hand. Add an fdopen rule to check-windows-footguns.py, which
also runs as a blocking CI step, so a bare text-mode fdopen fails CI.

Also fix a false-negative in the read_text/write_text rule: chained
forms like `read_text()[:4000]` or `read_text().splitlines()` never end
the line with `)` and slipped past the multi-line-call heuristic.
Replace the endswith check with a paren-balance walk (keeps multi-line
calls with encoding= on a continuation line unflagged — verified against
the full tree). This makes the rule the effective standing replacement
for the standalone checker proposed in PR #66669: R1-style coverage now
lives in PLW1514 + this script, both blocking in .github/workflows/lint.yml.

Sabotage-verified: reverting agent/shell_hooks.py's fdopen encoding or
tools/skills_tool.py's read_text encoding now fails the gate.

Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Paulo Nascimento <pnascimento9596@gmail.com>
2026-08-08 12:32:23 -07:00
Ken Weiner
ff7af1cbaf chore: map kweiner contributor email 2026-08-08 14:06:30 +05:30
kshitij
5ded99af45 chore: add prashantjain25 to AUTHOR_MAP
Needed before salvaging PR #80740 (contributor audit runs against main).
2026-08-07 20:06:53 +05:30
kshitij
ca120413fc fix(tests): forward HERMES_TEST_* knobs through the hermetic runner
scripts/run_tests.sh runs the suite under `env -i` with an explicit
allowlist. The runner's own documented environment knobs were never on
that list, so all of them were silent no-ops for anyone invoking the
canonical wrapper:

  * HERMES_TEST_WORKERS / PATHS / FILE_TIMEOUT / FILE_RETRIES / SLICE
    are read by run_tests_parallel.py at argparse-default time — inside
    the stripped environment.
  * HERMES_TEST_IMAGE is read by tests/docker/conftest.py to skip its
    session-scoped docker build.

The HERMES_TEST_IMAGE strip is the expensive one, and it's been biting
CI since docker.yml switched from bare pytest to run_tests.sh
(f0cb04921): the workflow sets HERMES_TEST_IMAGE to the image the build
step just loaded, the wrapper drops it, and every per-file pytest
subprocess falls back to building hermes-agent-harness:latest itself.
The job log timing shows it plainly — the first 8 files dispatched (the
LPT-heaviest) all report 248-297s, which is them waiting out the
concurrent initial `docker build` (~4 min on a cold local builder);
every file dispatched after that rides the layer cache and finishes in
4-38s (e.g. test_dump_build_sha.py, a single `docker run --entrypoint
cat`, reported 256.6s). ~4 min of pure waste per docker job, on both
arches — and the tests exercised a locally-rebuilt image WITHOUT the
HERMES_GIT_SHA build-arg the workflow bakes in, not the artifact being
shipped.

Fix: forward the six knobs the same way the Windows location vars are
forwarded (66c4c9c0b) — an explicit compute-before-drop allowlist, each
var only when set, so POSIX runs without them are byte-for-byte
unchanged and the 'no credential can leak' property stays auditable.

Verified empirically via a probe test through the wrapper:
  before: HERMES_TEST_IMAGE=None inside the subprocess
  after:  HERMES_TEST_IMAGE='sentinel-image', HERMES_TEST_FILE_TIMEOUT
          forwarded, HERMES_TEST_WORKERS=3 yields '(3 workers)' in the
          summary, and an unrelated SOME_SECRET stays stripped.
bash -n clean; shellcheck: no new findings (SC2046 on the pre-existing
compileall line predates this change).
2026-08-05 19:46:43 -04:00
Teknium
ced8e30217 feat(scripts): reproducible core-toolset A/B eval harness (toolperf_abeval)
Ships the hard A/B evaluation used for the August 2026 core-toolset
performance batch (#77056) as a reusable harness: 9 error-inducing trap
tasks derived from measured production waste classes, two-arm
PYTHONPATH-only comparison, ATOF-trace-based scoring, resume-safe
batteries.

Hardened from the original one-off: paths de-hardcoded (ABEVAL_ROOT /
ABEVAL_HOME), encoding= on all file IO, startup crashes retry on resume
instead of polluting cells, post-hoc grading fix for err_inline_script
baked in. Live-smoked end to end (baseline arm, qwen3-coder-30b,
err_multi_dir: exit 0, correct on-disk verification, resume record
written).
2026-08-05 13:43:30 -07:00
Jeffrey Quesnelle
6564f319a6 Merge pull request #69416 from afourniernv/feat/hermes-relay-install-activation-metrics
feat(observability): add Relay active install metrics
2026-08-05 14:09:25 -04:00
Jeffrey Quesnelle
edf0a7e14b Merge pull request #68978 from afourniernv/feat/hermes-relay-client-dimensions
feat(observability): add Relay client resource metrics
2026-08-05 14:02:28 -04:00
Jeffrey Quesnelle
0531aad55d Merge pull request #68883 from afourniernv/feat/hermes-relay-skill-metrics
feat(observability): aggregate bounded skill metrics
2026-08-05 13:20:57 -04:00
ethernet
ee7c614eef fix(ci): follow artifact download redirect without auth
The artifact download URL returns a 302 redirect to a signed blob URL.
urllib sent the Authorization header to the blob, and the blob rejected it
with a 401 error. The download now has two hops. The first hop authenticates
to the API. The second hop follows the redirect without the auth header.

The query runs?event=workflow_call returns nothing for this repository.
GitHub flattens reusable-workflow jobs and their artifacts into the caller
run. The fetch now lists the artifacts on the orchestrator run only. The
dead sub-run enumeration is gone. Two API calls per cycle are gone with it.

The 'artifact statuses updated' reason never appeared. The code updated the
count before the comparison. Now the code compares first and updates after.

The code rejects zip members that contain '..' or start with '/'.

tests/ci/test_live_comment.py is deleted. This repository does not keep
tests for CI infrastructure.
2026-08-05 11:16:18 -04:00
ethernet
1d7d0e41af ci: poll review statuses from artifacts every cycle
The live comment poller got its review statuses from two sources. The first
was the REVIEW_STATUSES environment variable, fixed at the start of the
comment-live job. The second was one ci-timings artifact, downloaded at the
end of the run. Status details (error messages, action_required items)
appeared only after all jobs finished. The job pass/fail results were visible
as each job completed.

Now every status-producing workflow_call uploads a small review-status
artifact when it completes. The poller lists all review-status-* artifacts
from the orchestrator run and its workflow_call runs every cycle. It
downloads each artifact and merges the statuses into the comment. A status
appears as soon as its job finishes.

Changes:
- live_comment.py: _fetch_artifact_statuses became fetch_all_review_statuses.
  The new function lists the artifacts via the API, downloads each one, and
  parses it. Removed the review_statuses_json parameter, the
  --review-statuses-file argument, and the subprocess import.
- ci.yml: removed the REVIEW_STATUSES environment variable, the inline Python
  merger, and the --review-statuses-file argument. Renamed the
  ci-timings-review-status artifact to review-status-ci-timings.
- Eight workflow_call files: added a step that writes review-status.json and
  uploads it as an artifact after each review_status output.
- test_live_comment.py: added tests for _parse_status_file and
  _merge_statuses.
2026-08-05 11:16:18 -04:00
ethernet
949babd083 ci: add detailed logging to live comment poller
The poller logs transitions between polls. It reports newly completed jobs
(with their results), newly appeared jobs, and jobs that left the pending
list. Each comment update shows the reason for the change. For example:
'1 new completion(s); artifact statuses updated'. When nothing changed, the
poller lists the jobs that are still pending. The status line shows the raw
job count from the API and the number of infra jobs that the filter removed.
2026-08-05 11:16:18 -04:00
brooklyn!
ae6c2e57e1 Merge pull request #78682 from NousResearch/bb/win-8dot3-profile-paths
fix(install): Windows install completes on profiles with a space, dot, or accent in the username
2026-08-05 08:21:23 -06:00
Alex Fournier
806c2b1fdc Merge updated client resource metrics into active-install metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-04 15:10:41 -07:00
Alex Fournier
e7eaae2bd3 Merge latest skill metrics into client resource metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-04 15:07:34 -07:00