Commit Graph

619 Commits

Author SHA1 Message Date
teknium1
3859e5c5b2 test(e2e/windows): a stuck install.ps1 fails with its process table and Python stacks
install.ps1 now runs under the same watchdog as `hermes update` (40 min),
and the watchdog adds a py-spy stack for every Python process of the leg
before it stops them. The workflow installs py-spy beside pywinpty.
2026-09-27 04:15:09 -07:00
teknium1
840f0fc9cc test(e2e/windows): one PM tool store per leg
`windows: desktop-installer@latest -> hermes-update (HEAD -> NEXT)` and
`installer-script+desktop -> hermes-update (HEAD -> NEXT)` failed with
"source launcher publication failed". install.ps1 and `hermes update` ran
with setup-pm's HERMES_RUNTIME_DIR, while the desktop smoke launched the app
without it (as a user's app runs) and settled the checkout onto a second
store at <HERMES_HOME>\tools. The update then re-pointed its own running,
locked hermes.exe.

windows-e2e.ps1 now drops HERMES_RUNTIME_DIR, HERMES_PYTHON and VIRTUAL_ENV
on entry, so every product step of a leg (installer, app, update, chat)
resolves the one store a user's machine has. The driver keeps setup-pm only
through $DriverPython (and PATH, which older installers need for uv and
ripgrep).

`hermes update` also runs under a 45-minute watchdog. Past it, the driver
writes every process (pid, parent, start time, command line) to
logs\update-hang-processes.txt, stops this leg's processes and fails with
that table, instead of being cancelled blind at the job cap.
2026-09-27 04:15:09 -07:00
teknium1
bb1e5680d6 ci(install-e2e): give Windows legs a 90-minute job budget
windows: desktop-installer@latest -> hermes-update (v2026.9.24 -> HEAD)
was cancelled at the 60-minute cap (job 108542725043) with nothing wrong:
its hermes update ran the historical venv->PM takeover, npm ci and a
Desktop rebuild in 24 minutes (one npm ci stalled 11 of them), exited 0 and
passed every post-update check, and the desktop smoke was cut off. Green
Windows legs already take up to ~40 minutes and the app-update driver may
wait 30. Real hangs stay bounded by pty-run.py and the per-step budgets.
2026-09-27 04:15:09 -07:00
teknium1
03ffc307ee ci(install-e2e): a rate-limited result chart no longer turns the run red
The Result chart job lists the run's jobs and artifacts through the repo's
GITHUB_TOKEN, which shares one hourly budget with every other workflow. In
a busy hour it answered "API rate limit exceeded for installation" and
failed a run whose legs were all green (36283843710, 36285499760). Retry
after 30/60/90 s. If the API still refuses, write a warning and a summary
line instead of failing. Each leg's own conclusion still stands.
2026-09-27 04:15:09 -07:00
teknium1
5afba1e8d9 ci(install-e2e): keep one install-e2e-red issue in step with the scheduled matrix
A red scheduled run blocked nothing and told nobody. install-e2e-red.yml
runs after every scheduled "Install & Update E2E" run: red opens one issue
labelled install-e2e-red (or rewrites the open one's body in place) with the
red legs grouped by failure class and linked; the first green run closes
it. No per-run comment, never a second issue.

It is its own workflow_run workflow because install-e2e.yml is also called
by stable-release.yml with read-only permissions, and a nested job asking
for issues: write would fail that call at startup. workflow_dispatch with a
dry-run default previews the change for any run id.
2026-09-27 04:15:09 -07:00
teknium1
7605349f25 ci(install-e2e): run a path-filtered four-leg subset on pull requests
install-e2e.yml only ran on the clock, so nothing in front of a merge
installed a release and updated it on a real OS. A pull_request trigger,
path-filtered to the install/update surface (derived from 60 days of
update/install/pm commits), runs the new `pr` route of
generate-e2e-matrix.mjs with only the newest release tag sampled:

  linux   installer-script -> hermes-update (newest release -> PR)
  linux   installer-script -> hermes-update (PR -> NEXT)
  windows installer-script -> hermes-update (PR -> NEXT)
  macos   installer-script -> hermes-update (newest release -> PR)

The bundle-manifest validation job is skipped on PRs (bundled legs need
dispatch-only manifests). The full matrix stays on schedule and release.
2026-09-27 04:15:09 -07:00
teknium1
7b4211eee2 test(e2e/desktop): stop the installer's nested lazy-fetch chain on CI
Trail from CI: 78+ nested 'git fetch origin --filter=blob:none --stdin'
lazy fetches, each with its own upload-pack on the seeded origin, until
the runner is OOM-killed. The origin serves full clones, so nothing is
legitimately missing; GIT_NO_LAZY_FETCH=1 for the installer run only.
2026-09-27 00:42:43 -07:00
teknium1
9bc96b8502 ci(e2e-desktop-update): trail the upload-pack requester, not the oldest git 2026-09-27 00:42:43 -07:00
teknium1
78084696a4 ci(e2e-desktop-update): trail git transfer ancestry
The second CI run showed the OOM cause: hundreds of concurrent
upload-pack/pack-objects processes on the seeded origin.git, reaching
105 GB RSS about 90 s into seeding. Log the client-side git command and
its parent chain to find what drives them.
2026-09-27 00:42:43 -07:00
teknium1
3b1da278d1 ci: Desktop update E2E is advisory via continue-on-error, not by leaving all-checks-pass
tests/ci/test_stable_release_graph.py requires every reusable-workflow
job in all-checks-pass; keep it there and let a red shard show without
failing the gate. Log a resource trail during the run: every shard's
runner died ~2 minutes into seeding on the first CI run with only
'The runner has received a shutdown signal'.
2026-09-27 00:42:43 -07:00
teknium1
e4c4d95d29 ci: keep Desktop update E2E out of all-checks-pass until it is stable on CI 2026-09-27 00:42:43 -07:00
Teknium
183e0d4adc test(e2e/desktop): PR-time Desktop install/update suite on a real install
Drive the packaged app a real scripts/install.sh + hermes desktop install
built, against a local bare origin, a local api.github.com stand-in and a
scripted model: first run, Update now hand-off/relaunch/chat, a failed
Desktop build inside that update, and a second update while one runs.
Open bugs are message-gated (gated on #122991).
2026-09-27 00:42:43 -07:00
teknium1
81f873b334 test(e2e/pm): PM environment lifecycle across real hermes updates
tests/e2e/core/upgrade/pm/: 22 cells in 7 files, one per failure class, each
driving the real entry points (install.sh, hermes update, hermes pm
repair/status/install, hermes doctor, hermes gateway run/install --force,
hermes kanban dispatch, hermes -z through the loopback provider) inside the
existing bwrap sandbox, seeded with an N-1 install and a dependency-changing
release:

- configured features survive update, repair and legacy-venv migration
- SIGKILL at the stage/publish boundary of a dependency update recovers
- stamps, receipts and doctor tell the truth (healthy, drifted, failed update)
- no stray in-tree venv / repo-local 3.11 .venv reaches any interpreter;
  execute_code and Kanban workers spawned after an update boot and import deps
- repeated updates/repairs/restarts don't accumulate generations
- gateway install --force keeps the install bootable

Every non-gated cell has a recorded sabotage proof; open bugs are
message-gated with known_failure (gated on #124228, #122627, #122425,
#124214, #124075, #124049, #124668).

Harness: make_origin clones --single-branch and allows filters, so a
blobless developer checkout can serve the installer's clone.

CI: e2e-upgrade is sharded per suite directory (an e2e-upgrade-plan job
lists "core" plus every subdirectory holding test_*.py); a shard that
collects no files fails; HERMES_TEST_WORKERS=6.
2026-09-27 00:42:13 -07:00
teknium1
15ae3f94b2 ci(windows-install-update-e2e): headroom for slow Windows runners 2026-09-27 00:41:41 -07:00
teknium1
16da7f1b38 test(e2e/windows): real C:\Users profiles, serialized gateway phases, Git-for-Windows machines 2026-09-27 00:41:41 -07:00
teknium1
9129897dd1 test(e2e/windows): PR-time install.ps1 -> hermes update journey on native Windows 2026-09-27 00:41:41 -07:00
teknium1
79872aaf3f test(e2e): messaging-adapter contract suite across Telegram/Discord/Slack 2026-09-26 11:57:54 -07:00
ethernet
b5d583c4ac feat(release): start the gates and the signed candidates together
Every gate and every candidate now needs only admit, so the signed macOS
and Windows builds and the docker image stop waiting behind the full CI
run. acceptance stays the one join: it still needs ci, docker and all six
candidates, so what can be promoted is unchanged.

The four gates that skipped under --skip-tests only because ci skipped
(nix, termux-checks, windows-live, install-e2e) and pm-bundle needed their
own condition. SKIPPED_BY requires the gate to observe 'skipped', so
dropping the edge alone would have left them running and blocked the
release instead of failing it.
2026-09-25 12:50:53 -04:00
Hermes Agent
5307e93252 ci(e2e): keep the upgrade suite out of the e2e job again
27df3b8847 dropped the exclusion 65e79880c8 added, so the 30-minute e2e
job (shallow checkout, no bwrap, 900s per file) ran tests/e2e/core/upgrade
next to its own e2e-upgrade job. Its install/update files then timed out
and the job was cancelled at the step limit on main and every Python PR.
2026-09-25 10:15:47 -05:00
ethernet
2ab7d9bfe3 chore(ci): remove publish-e2e-evidence pipeline
gh v2.99 gained a native --attach flag for issues, PRs and comments, so
the custom trusted-publisher chain (gh-image extension + GH_IMAGE_SESSION_TOKEN
workflow_run job + attachment-upload script) is superseded.

Remove:
- .github/workflows/publish-e2e-evidence.yml (workflow_run publisher)
- scripts/ci/publish_e2e_evidence.py + its tests
- e2e-evidence-* artifact upload + evidence staging in e2e-desktop.yml /
  e2e_screenshot_status.py, incl. the 'inline evidence is publishing...'
  marker placeholder in the CI review comment status

Keep: the review-comment screenshot/diff counts and artifact links produced
by e2e_screenshot_status.py.
2026-09-24 23:53:17 -04:00
ethernet
3aa215fdc9 build: remove leftover references to the dropped hindsight extra
The extra removal left CI, the Docker image and the nix package still
requesting `hindsight`. Once the extra is gone, `--extra hindsight` and
extraDependencyGroups = [ "hindsight" ] ask for something that no longer
exists. Drop them the same way 73c598e319 originally did: remove it from
the CI extras lists, the Docker sealed-venv build and the nix default
groups, and point the nix examples/check at honcho. Also remove the
stray blank line left in the exclude-newer table.
2026-09-24 22:15:15 -04:00
ethernet
ac4181fdfa Merge pull request #122030 from NousResearch/ci/bootstrap-installer-build
ci: dispatch-only signed builds of the bootstrap installer
2026-09-24 20:02:06 -04:00
ethernet
13cd98b06b ci: dispatch-only signed builds of the bootstrap installer
Hermes-Setup has never had a CI build; every published copy was built by
hand. This workflow_dispatch lane builds the Windows x64 exe (signed via
the desktop MSIX's Azure Trusted Signing path, batch-sign-binaries.mjs)
and the macOS arm64 dmg (Developer ID signed, app + dmg notarized and
stapled) from any ref, and uploads them as run artifacts. Nothing is
published. No caches: no actions/cache or setup-node cache, and fresh
npm/cargo/electron-builder cache dirs.
2026-09-24 19:39:19 -04:00
ethernet
4b7229d612 fix(icons): commit generated icons; installs and regular builds never render
User installs failed with 'resvg-py is missing' because the web/desktop
source builds rendered icons on whatever python was on PATH. The default
brand outputs are now committed; source_build, apps/desktop build.mjs and
the npm/docusaurus pre-hooks consume them directly. Flavored release
bundles (canary/commit) still render into their own product dir.

icons-freshness-check now regenerates and fails on any byte diff.
2026-09-24 19:17:34 -04:00
ethernet
b5dcdbf151 ci(e2e): run the e2e lane at 3 workers 2026-09-24 18:45:32 -04:00
ethernet
d6e78f37b6 ci(e2e): run fewer process-tree suites at once
The e2e lane at 4 workers still starved the PTY turn and serve-SIGTERM
deadlines intermittently; drop it to 2. The upgrade lane had no cap and
ran cpu_count real-updater trees at once; cap it at 4.
2026-09-24 18:40:44 -04:00
ethernet
b2866924df ci(pm-bundle): skip the linux legs until the larger runners run 24.04
The ubuntu-latest-32-* larger runners boot ubuntu-22.04 (glibc 2.35), and
the prebuilt llama-server in the native payload needs GLIBC_2.38, so the
staged-entry verification fails. Restore the two entries once the runner
images are updated.
2026-09-24 17:16:54 -04:00
ethernet
473cac85e7 ci(pm-bundle): pin the bootstrap python to 3.13
The ubuntu-latest-32-* larger runners boot ubuntu-22.04 images, whose
system python3 is 3.10. The bundle bootstrap imports tomllib (3.11+), so
linux-arm64 failed before doing any work. Pin the interpreter instead of
trusting the image.
2026-09-24 16:58:12 -04:00
ethernet
fcd4e40dfc ci(e2e): provision node through PM so /api/pty can start the TUI
The dashboard PTY test failed on this branch with "node: not installed and
lazy installs are disabled". On this branch _make_tui_argv asks PM for node
(hermes_cli/main_tui_launch.py _tui_node_bin). The E2E sandbox forbids lazy
installs (HERMES_DISABLE_LAZY_INSTALLS=1 in tests/e2e/core/dashboard
hermetic_env). It gets node only when the runner's PM store has it
(tests/e2e/core/_pm_dependencies.py). The e2e job got node from
actions/setup-node, so node was on PATH but not in the PM store. main passes
the test because main's TUI launch takes node from PATH.

Install the locked toolchain (setup-pm toolchain: all) before npm ci, in
place of setup-node. The locked Node 26 also satisfies the boot-contract
suite. When HERMES_E2E_REQUIRE_TUI=1, the fixture now fails at once if the
store has no node, not later with an opaque 1011 close.
2026-09-24 16:49:58 -04:00
ethernet
1fc6bc8cee feat(ci): add windows-latest-32-arm-core and use it
also bigger runs where needed
2026-09-24 16:49:26 -04:00
ethernet
99768fd127 feat(pm): hermes pm lock relocks uv.lock after a pyproject edit
Contributors had to run `python -m pm.build_env --source . --lock-only`, then
re-source activate, then call sync_venv for opt-in extras, while `hermes pm
lock` only pinned tool artifacts in pm/lock.json. Now `hermes pm lock` with no
arguments is the one step: it runs the same PM check_project_lock /
lock_project operations as build_env's --check-lock / --lock-only (same
exclude-newer quarantine), writes nothing when the lock is current, changes
no environment, and prints the activate command to run next. Activation
syncs [all] plus recorded extras only, and --test-extras replaces the
default, so a newly added extra outside [all] gets
`source ./activate --test-extras all,NAME` (`-TestExtras` on PowerShell).

`hermes pm lock --bump NAME VERSION` keeps writing pm/lock.json; NAME and
VERSION are now only accepted together. build_env --lock-only is unchanged
for scripts. Contributor docs, AGENTS.md, CONTRIBUTING.es.md, the pyproject
comment, and the uv.lock CI remediation text point at `hermes pm lock`.
2026-09-24 15:55:08 -04:00
ethernet
e78e3a77be refactor(release): move the author map out of release.py
release.py was 2,804 lines, 2,076 of them the frozen legacy author dict.
The map and its resolver now live in scripts/releases/: authors.py holds
the directory loader, the merged AUTHOR_MAP and resolve_author, and
authors_legacy.py holds only the frozen LEGACY_AUTHOR_MAP literal (same
1,895 entries, same order). release.py drops to 707 lines.

The contributor-check job and audit_pr_attribution.py grep the legacy
file for quoted emails, so both now read authors_legacy.py. The
importers (contributor_audit.py, add_contributor.py), contributors/README
and the tests read the defining modules. No behaviour change.
2026-09-24 14:08:22 -04:00
ethernet
dc11e3b3bc feat(release): add --skip-bundles and --skip-tests to stable releases
`release.py release` gains two flags. They can be used together.

--skip-bundles ships only the claim, the GitHub release, the final tag
and the Docker image. No desktop, Termux or PM bundle job runs. The
final tag records candidateManifestSha256: null. Publication moves only
the Docker stable/latest aliases. The R2 stable head, feeds, APT, the
downloads page, the signed-package baseline and the Store stay on the
previous bundle release.

--skip-tests builds, signs and publishes every artifact and runs no
test job: source CI, Nix, PM bundle check, Termux, Windows live,
install/update E2E, bootstrap identity, native smokes, upgrade
acceptance, tests/docker and the in-build vitest step. The candidate
manifest records each smoke as skipped, never as passed.

The flags live in the claim message (skipBundles, skipTests), next to
autopublish. They are not workflow inputs, so a rerun cannot change
them. admit emits them, and every job condition and gate reads them.
stable.validate_claim and stable.validate_final are now the one shape
check for stable.py and the sequencer.

The gates stay strict. SKIPPED_BY in stable.py maps each job to the
flags that remove it. `gate` requires those jobs to report skipped and
every other gated job to report success. A job that ran although a flag
removes it blocks the release.

A release that skipped bundles never moves the R2 stable head. Two
readers depended on that head:

- The next version was derived from it, so the next cut would reuse the
  version. It now takes the newer of the R2 head and the newest
  published non-prerelease GitHub release with a vX.Y.Z tag. Bare v*
  tags do not count, because those refs are not protected yet.
- The sequencer used it to decide which published releases still need
  their publication pass, so a bundle-less release would re-advance
  every 15 minutes. The head is now the newer of the R2 head and the
  published release whose final tag binds the Docker stable alias
  digest.

`release` also refuses a cut when its next version already has a final
tag. That closes the window between the final tag and the public
release, where the published identity still names the old version.

Tests: 42 release test files, 546 passed. Three tests fail on this
Windows host, and they fail the same way on a clean HEAD worktree:

- test_stable_release_graph::test_docker_recovery_refuses_to_replace_a_divergent_version_tag
- test_release_artifacts::test_windows_metadata_is_read_from_package_and_stale_stamp_is_rejected
- test_tag_builds_summary::test_admitted_failure_publishes_tag_info_without_promoting_channel[True]

Not verified: no real Stable Release dispatch ran with either flag, and
actionlint is not installed on this host. The workflow changes are
checked by the graph tests and by running the phase-result step script.
2026-09-24 13:31:33 -04:00
ethernet
6f14f6001d ci(lazy-deps): describe tools.lazy_deps as the stub it is
tools/lazy_deps.py was not deleted; it survives as an old-updater stub
that raises or stops for relaunch. The job display name and the
checker/test prose claimed it was deleted, which sends readers looking
for a missing file. The display name is not a required status check
(main requires only "All required checks pass", which keys on job ids),
so rename it to "No production imports of the tools.lazy_deps stub".
2026-09-24 11:50:30 -04:00
ethernet
845b61f2b8 fix(release): rerun failed stable runs on the failure event, drop the cron
Stable Release Publication ran every 15 minutes (96 runs a day, each
checking out full history, setting up node and buildx, logging into
Docker Hub, and taking the release-signing environment) only because the
sequencer held a failed run for a 15-minute backoff that the failure
event could never satisfy, so the cron was what actually retried.

Drop the backoff: the reconcile pass started by a failed Stable Release
reruns its failed jobs right away. MAX_ATTEMPTS burning, oldest-first
retry ordering, the attempt-entry check, and the needs_retarget repair
stay. The schedule trigger goes; workflow_run and workflow_dispatch
remain the recovery paths.

The shared stable-release concurrency group cannot deadlock: the rerun
waits as pending behind this job, and the sequencer only confirms the
new attempt is queued before it exits and frees the group.
2026-09-24 11:50:30 -04:00
ethernet
890e9b73d4 test(desktop-e2e): run the core Desktop E2E on the PM toolchain
Upstream's e2e-desktop-core job provisioned setup-node + uv sync and pointed the
packaged smoke at a checkout .venv. Provision through setup-pm, hand its selected
interpreter to Desktop via HERMES_DESKTOP_PYTHON, and read the packager contract
from electron-builder.config.cjs, where the build config now lives.
2026-09-24 11:12:27 -04:00
ethernet
427d0883be Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 09:33:58 -04:00
ethernet
5c062d0cc3 docs: clarify Python updater range and test sandbox marker 2026-09-24 09:20:50 -04:00
ethernet
e8a5e0978c chore(desktop): require prepared build Python and refresh release comment 2026-09-24 09:20:50 -04:00
teknium1
65b24e9600 test(desktop): packaged-app smoke — asarUnpack contract + packaged binary boots to a first chat (#121097) 2026-09-24 06:06:00 -07:00
kshitijk4poor
a84e4ea8bc ci: drop the venv-e2e workflow edit; tests-os.yml already runs the live test
The workflow edit trips the review-label gate. The new test is marked
windows_only, so tests-os.yml's '-m windows_only' job on windows-latest
already collects it.
2026-09-24 18:08:24 +05:30
JoaoMarcos44
6c71ff2bd7 fix(state): detect Windows database holders before maintenance
(cherry picked from commit b006ae2dcf6b210d3db3dfbe78c5aac040044644)
2026-09-24 18:08:24 +05:30
ethernet
5ef56ede40 fix(ci): bound hosted e2e process-tree concurrency
The 32-file Linux lane runs many more child processes than CPU cores. Hosted children remained alive without readiness output and exhausted per-child deadlines across unrelated suites; four concurrent files complete the focused tenancy, terminal model and compaction cases locally without weakening their assertions.
2026-09-24 08:34:49 -04:00
ethernet
f689044ef2 ci(windows): run E2E on PM test environment and platform markers 2026-09-24 07:49:06 -04:00
ethernet
43c105d5ba Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 07:39:33 -04:00
teknium1
9521ee098b ci(windows-e2e): run the e2e-windows job on windows-latest-32-core
Same runner class as the os-tests Windows row. Seven files run in
parallel, each spawning process trees; measured on the fixed suite the
test step drops from 153-171 s to 124 s and the job from ~5 min to 3.1.
2026-09-24 04:25:48 -07:00
teknium1
0616f72049 test(windows-e2e): find serve/gateway orphans by profile ownership, not ancestry
The serve tree-kill test snapshotted process_tree(first.pid) and then
checked that same tree after `taskkill /T /F` - the snapshot is exactly
what taskkill /T kills, so the check could not fail. Windows never
re-parents: a grandchild spawned detached (cmd /c start /b ...) keeps a
dangling ppid and is invisible to both Process.children() and
taskkill /T. Reviewer sabotage (web_server spawns two detached sleepers
before READY) passed while both orphans survived.

Survivors are now every live process created since the spawn whose
HERMES_HOME, cwd or argv points into the test's scratch profile
(owned_processes), with a positive control that the scan sees the
backend itself before the kill. Same check for gateway stop. Cleanup
kills everything the profile owns.

Also: state-db guard matches the holder PID as a whole number, and the
e2e-windows job ends with a scan that fails on any python.exe /
hermes.exe left running.
2026-09-24 04:25:48 -07:00
teknium1
8c44895e7b test: native Windows E2E suite on real Hermes processes (e2e-windows job)
Adds tests/e2e/core/windows (windows_only + integration; 18 tests, 5 strict-xfail KNOWN
entries for #120504 #121150 #121015 #121114 #120205) and an e2e-windows job in
tests-os.yml running it on windows-latest, one pytest process per file, no retries.
2026-09-24 04:25:48 -07:00
ethernet
4ffba2c882 test(ci): prepare pinned Git before Windows installer stages 2026-09-24 05:45:13 -04:00
ethernet
1aeb53a40b fix: align upgrade CI and installer fixture with PM bootstrap 2026-09-24 02:50:25 -04:00