Commit Graph

44927 Commits

Author SHA1 Message Date
teknium1
41755854ad test(e2e/windows): accept Hermes-Setup.exe's legacy receipt as a permanent shape
#125053 fixed the exe's receipt writer, but the installer .exe is not being
rebuilt, so every Desktop-installer machine keeps the old receipt
(completedAtUnix, and a null pinnedCommit without git on PATH) until
`hermes update` rewrites it. No reader depends on completedAt or
pinnedCommit: Desktop's launch gate is runtime usability, and every other
reader only checks that the file exists. So the stamp phase stops
presenting this as a gate that lifts when a fixed exe ships and asserts the
contract that holds with the released exe: only those two FAIL lines, an
integer completedAtUnix, and the checkout at the expected commit. Any
other FAIL line stays red.

Refs #124949, #125053.
2026-09-27 04:15:09 -07:00
teknium1
61b97ffc12 test(e2e/windows): gate the Hermes-Setup.exe receipt on #124949
With one PM store per leg the Desktop app no longer re-bootstraps on first
launch, so the receipt Hermes-Setup.exe writes survives to verify-stamp:
pinnedCommit null on a git-less machine and completedAtUnix instead of
completedAt, replacing install.ps1's correct receipt. Gate only that
writer's exact two FAIL lines, and then require the checkout itself to be
at the expected commit. Any other FAIL line stays red.
2026-09-27 04:15:09 -07:00
teknium1
85a1dd1cd2 test(e2e/windows): derive HERMES_HOME from a normalized workroot
With one store per leg, HEAD's install.ps1 never finished on Windows: the
source completion relaunched its store Python into itself every ~19 s, 35
levels deep after 12 minutes (job 108573522351). The workflow's workroot is
<workspace>\..\hermes-desktop-gui-e2e, so the store path carried the "..";
Windows reports sys.executable normalized, and prepare_launch()'s
Path.absolute() comparison never matched (#122513, fix PR #122518). The
job's own store (setup-pm) had no "..", which hid this until now. Normalize
the workroot, as a user's HERMES_HOME is.
2026-09-27 04:15:09 -07:00
teknium1
3859e5c5b2 test(e2e/windows): a stuck install.ps1 fails with its process table and Python stacks
install.ps1 now runs under the same watchdog as `hermes update` (40 min),
and the watchdog adds a py-spy stack for every Python process of the leg
before it stops them. The workflow installs py-spy beside pywinpty.
2026-09-27 04:15:09 -07:00
teknium1
93105d7d87 test(e2e/windows): undo installer CRLF churn through a pathspec file
`windows: {installer-script, installer-script+desktop, desktop-installer@latest}
-> hermes-desktop-app-update (v2026.7.1 -> HEAD)` failed with "Program
'git.exe' failed to run: The filename or extension is too long". A v2026.7.1
clone leaves hundreds of CRLF-churned paths, and Clear-HistoricalInstallerChurn
passed every one to `git checkout --` on the command line. Pass them through
a NUL-separated --pathspec-from-file with --literal-pathspecs instead.
2026-09-27 04:15:09 -07:00
teknium1
840f0fc9cc test(e2e/windows): one PM tool store per leg
`windows: desktop-installer@latest -> hermes-update (HEAD -> NEXT)` and
`installer-script+desktop -> hermes-update (HEAD -> NEXT)` failed with
"source launcher publication failed". install.ps1 and `hermes update` ran
with setup-pm's HERMES_RUNTIME_DIR, while the desktop smoke launched the app
without it (as a user's app runs) and settled the checkout onto a second
store at <HERMES_HOME>\tools. The update then re-pointed its own running,
locked hermes.exe.

windows-e2e.ps1 now drops HERMES_RUNTIME_DIR, HERMES_PYTHON and VIRTUAL_ENV
on entry, so every product step of a leg (installer, app, update, chat)
resolves the one store a user's machine has. The driver keeps setup-pm only
through $DriverPython (and PATH, which older installers need for uv and
ripgrep).

`hermes update` also runs under a 45-minute watchdog. Past it, the driver
writes every process (pid, parent, start time, command line) to
logs\update-hang-processes.txt, stops this leg's processes and fails with
that table, instead of being cancelled blind at the job cap.
2026-09-27 04:15:09 -07:00
teknium1
08f0691fa3 Revert "test(e2e/windows): hermes update follows the store the desktop smoke settled on"
This reverts commit c7883282c91. On real Windows (runs 36293787001 and
36293789061, c7883282 + #124702), dropping HERMES_RUNTIME_DIR for the update
turned a 40-second red ("source launcher publication failed") into a silent
40-minute hang: the update finished its code, Desktop rebuild and skills
steps, printed the curator notice, and never exited until the job was
cancelled. The store it then used (<HERMES_HOME>\tools) was created by the
desktop smoke and never received the install's PM tools, so the update's
post-update default-tool install ran against a half-populated second store.
Both one-sided variants (keep the variable in the smoke; drop it for the
update) fail. The real fix is one store per leg, and it touches every
Windows leg, so it goes in on its own with a full Windows run.
2026-09-27 04:15:09 -07:00
teknium1
bb1e5680d6 ci(install-e2e): give Windows legs a 90-minute job budget
windows: desktop-installer@latest -> hermes-update (v2026.9.24 -> HEAD)
was cancelled at the 60-minute cap (job 108542725043) with nothing wrong:
its hermes update ran the historical venv->PM takeover, npm ci and a
Desktop rebuild in 24 minutes (one npm ci stalled 11 of them), exited 0 and
passed every post-update check, and the desktop smoke was cut off. Green
Windows legs already take up to ~40 minutes and the app-update driver may
wait 30. Real hangs stay bounded by pty-run.py and the per-step budgets.
2026-09-27 04:15:09 -07:00
teknium1
2d6620f668 test(e2e/windows): hermes update follows the store the desktop smoke settled on
`windows: desktop-installer@latest -> hermes-update (HEAD -> NEXT)` and
`installer-script+desktop -> hermes-update (HEAD -> NEXT)` failed with
"source launcher publication failed". The desktop smoke runs the app
without the job's HERMES_RUNTIME_DIR, the way a user's app runs, so it
settles the checkout onto <HERMES_HOME>\tools and republishes .hermes\bin
for that store's Python. The driver's `hermes update` then ran from that
hermes.exe with the job's store (setup-pm\tools) back in scope. The update
re-pointed both launchers, and Windows refused to replace the running
hermes.exe.

Invoke-HermesUpdate now drops HERMES_RUNTIME_DIR while <HERMES_HOME>\tools
exists, so the update stays on the one store the install actually uses.
Keeping HERMES_RUNTIME_DIR in the smoke instead fixed these two legs but
broke eight others whose app never reached its backend (run 36288957697),
so that variant was reverted.
2026-09-27 04:15:09 -07:00
teknium1
a965c4052e test(e2e/windows): the app-update driver waits out the real update and checks staged main
Once launch capture works, two Windows hermes-desktop-app-update failures
show up (run 36286580917):

- v2026.9.24 -> HEAD (installer-script, installer-script+desktop,
  desktop-installer@latest): the app's update is correct but slow. From
  v2026.9.24 it runs the historical venv->PM takeover and a full Desktop
  rebuild (the `hermes update` alone took 9m43s). launch-from-spec's 10 min
  default gave up 9m55s after the Update now click while the updater window
  showed "Updating code and dependencies 9m 22s elapsed". windows-e2e.ps1
  now passes --timeout-ms 1800000, in line with open-app-update's 35 min
  wait. The driver's self-deadline is now measured after the update wait
  instead of being a flat 20 min inside it.
- HEAD -> NEXT: Desktop offered real GitHub main (0f4a98f8) instead of the
  staged NEXT, so assertStagedBranch refused to click. HEAD's Electron main
  is one ESM bundle, and checkout-source.ts binds promisify(execFile) at
  load, before installSourceBranchProbe patches execFile in the packaged
  app. POSIX avoids this by wrapping the launcher script, which Windows
  skips. The probe now also hooks ChildProcess.prototype.spawn, which every
  child passes through, and reuses branchProbeArgs, which already knows both
  the --run-module and cmd.exe .cmd shapes. Verified locally: a promisify
  captured before the hook now gets `--git <real> --branch main`, and other
  commands are untouched.
2026-09-27 04:15:09 -07:00
teknium1
03ffc307ee ci(install-e2e): a rate-limited result chart no longer turns the run red
The Result chart job lists the run's jobs and artifacts through the repo's
GITHUB_TOKEN, which shares one hourly budget with every other workflow. In
a busy hour it answered "API rate limit exceeded for installation" and
failed a run whose legs were all green (36283843710, 36285499760). Retry
after 30/60/90 s. If the API still refuses, write a warning and a summary
line instead of failing. Each leg's own conclusion still stands.
2026-09-27 04:15:09 -07:00
teknium1
9f78424dbb test(install-e2e/windows): capture the packaged launch of a PM-launched hermes desktop
`windows: * -> hermes-desktop-app-update (HEAD -> NEXT)` failed with
"a launch was actually captured (exit 0 without a launch must not pass)":
`hermes desktop` built the app and launched the real Hermes.exe, and the
driver's hook never saw it. There were two gaps, and either one alone is
enough to miss the launch:

- HEAD installs publish the PM launcher (.hermes\bin\hermes.exe), which
  runs its interpreter with -I and drops PYTHONPATH, so sitecustomize never
  loads. windows-e2e.ps1 now does what installer-script-e2e.sh already does:
  run launch-capture/pm-launch.py for a PM launcher, and keep PYTHONPATH
  for pre-PM venv console scripts.
- Since 3ddf82fe24 the Windows packaged launch is a detached
  subprocess.Popen, not subprocess.run. sitecustomize now also intercepts a
  launch-shaped Popen: it records the same spec and stands in for a child
  that already exited with 0. Every other Popen, including the one inside
  subprocess.run, spawns for real.
2026-09-27 04:15:09 -07:00
teknium1
5f9c5a2c38 test(install-e2e/windows): tolerate v2026.9.24's CRLF clone churn before the GUI update
Three `-> hermes-desktop-app-update (v2026.9.24 -> HEAD)` legs failed with
"installed source has only installer-generated changes (other: 1 .M N...
<same blob> <same blob> website/i18n/zh-Hans/.../user-stories.mdx)".

v2026.9.24's install.ps1 clones under Git for Windows' system
core.autocrlf=true and pins false only afterwards, so the ~1700 text files
.gitattributes does not pin to LF (*.mdx, LICENSE, ...) sit on disk as CRLF
over LF blobs. The stat cache hides them, except the last few the clone
wrote inside the index's racy timestamp window, which read as modified.
Which files show up varies (1 or 3 per leg); main's installer no longer
pins autocrlf.

Clear-HistoricalInstallerChurn now also restores a `1 .M` entry whose
index and worktree modes match and that has no `diff --numstat
--ignore-cr-at-eol` record (a line-ending-only difference). A real content
edit, an edit inside a CRLF file, a mode change, or an untracked file still
fails the assertion, listed.
2026-09-27 04:15:09 -07:00
teknium1
5afba1e8d9 ci(install-e2e): keep one install-e2e-red issue in step with the scheduled matrix
A red scheduled run blocked nothing and told nobody. install-e2e-red.yml
runs after every scheduled "Install & Update E2E" run: red opens one issue
labelled install-e2e-red (or rewrites the open one's body in place) with the
red legs grouped by failure class and linked; the first green run closes
it. No per-run comment, never a second issue.

It is its own workflow_run workflow because install-e2e.yml is also called
by stable-release.yml with read-only permissions, and a nested job asking
for issues: write would fail that call at startup. workflow_dispatch with a
dry-run default previews the change for any run id.
2026-09-27 04:15:09 -07:00
teknium1
7605349f25 ci(install-e2e): run a path-filtered four-leg subset on pull requests
install-e2e.yml only ran on the clock, so nothing in front of a merge
installed a release and updated it on a real OS. A pull_request trigger,
path-filtered to the install/update surface (derived from 60 days of
update/install/pm commits), runs the new `pr` route of
generate-e2e-matrix.mjs with only the newest release tag sampled:

  linux   installer-script -> hermes-update (newest release -> PR)
  linux   installer-script -> hermes-update (PR -> NEXT)
  windows installer-script -> hermes-update (PR -> NEXT)
  macos   installer-script -> hermes-update (newest release -> PR)

The bundle-manifest validation job is skipped on PRs (bundled legs need
dispatch-only manifests). The full matrix stays on schedule and release.
2026-09-27 04:15:09 -07:00
teknium1
116b3e2f9e test(install-e2e/windows): the first chat turn exits under the pseudoconsole
Every Windows leg of every scheduled run since 2026-09-25 was cancelled at
the 60-minute job timeout in its Install step. The user-state phase runs
`hermes chat -q "..."` under pty-run.py (a ConPTY, so the CLI sees a real
TTY). Since a5c7eed (shipped in v2026.9.21) `-q` on a TTY seeds an
interactive session and only answers-and-exits with --oneshot; the turn
answered ("Hello from the mock inference server!") and then sat at the
prompt. pty-run.py's --timeout 900 never fired either: it checked the
deadline only between blocking PtyProcess.read() calls, and an idle prompt
never returns from read().

- windows-e2e.ps1 passes --oneshot when `chat --help` advertises it (older
  tags have no flag and exit after -q on their own), lowers the turn budget
  to 300s, and names a 124 as "never exited" instead of a generic failure.
- pty-run.py drains the pty on a daemon thread and waits on the queue with
  the deadline, so --timeout holds whatever the child does, and writes a
  TIMEOUT line into the captured log before killing it.
2026-09-27 04:15:09 -07:00
teknium1
d25bbd01b7 test(e2e/network): drop the #124653 gate; the channel 503 cell passes
The known_failure wrapper on test_channel_503_with_retry_after_is_retried_then_updates keyed on the one-attempt channel read. With update-path retries it passes, so the cell asserts plainly again.
2026-09-27 03:41:42 -07:00
teknium1
d1484ab44b fix(update): only an explicit hermes update retries channel reads
Passive checks (`hermes --version`, the banner, the Desktop/dashboard
update check) now make one channel-read attempt again. Offline usually
surfaces as DNS EAI_AGAIN or ENETUNREACH, which pm.network.is_transient
classes as transient, so wrapping every read in retry_network added 7 s
of backoff to the synchronous version line, stretched a hung CDN from
30 s to ~127 s, and logged a WARNING to stderr on every retry.

`release_channels.retrying_reads()` opts a block in; update_cmd wraps
the channel resolution of `hermes update` in it. Tests: keep the
Retry-After retry test (now scoped to the update path) and replace the
budget-exhaustion test with one pinning that passive reads and 404s make
exactly one attempt.
2026-09-27 03:41:42 -07:00
JoaoMarcos44
0da04c192c fix(update): retry transient release channel reads 2026-09-27 03:41:42 -07:00
teknium1
54388a21ca fix(cli): compare interpreter paths lexically, not through symlinks
resolve() on both sides also collapses a venv interpreter onto the store
binary it links to, so a caller still running in its old venv stopped
re-executing into the store interpreter (two tests in
tests/pm/test_source_update_launch.py go red). The loop in #122513 is a
spelling difference, not a symlink: PM spells the store Python through a
HERMES_HOME that may contain '..', and the OS reports sys.executable
normalized. normcase(abspath()) matches those spellings and keeps a venv
interpreter distinct.

Tests: the '..' spelling is current (red on main); a venv python symlinked
to the store binary still relaunches (red with resolve()).
2026-09-27 03:38:52 -07:00
Konstantin Khlopkov
1d6e53eb05 fix(cli): compare resolved interpreter paths so a current venv launch does not relaunch 2026-09-27 03:38:52 -07:00
teknium1
5a0225dfff fix(gateway): restart watcher names its checkout instead of inheriting the cwd
The detached watcher runs as `<python> -c <program>`, so `hermes_cli` resolved only
because update_completion happened to spawn it with cwd=<checkout>. The program now puts
the checkout on sys.path itself, and the bare-Python test runs it from an unrelated cwd
(red without the sys.path line).

Also drops the two unused re-export aliases in gateway.status (`_posix_is_zombie`,
`_pid_exists_win32_ctypes`): nothing imports them and neither is in the old-updater
compat surface.
2026-09-27 03:37:09 -07:00
teknium1
029445545c fix(gateway): keep the update restart watcher stdlib-only so it survives the bare store Python
After the package-manager handoff, hermes update finishes on the bare store
interpreter and spawns the detached restart watcher as sys.executable -c.
The watcher imported gateway.status (utils -> hermes_yaml -> ruamel) and
hermes_cli.config, died with ModuleNotFoundError before relaunching, and
left every manually started gateway down after the update (gated on #124649).

Move the stdlib liveness probe (zombie-aware POSIX kill(0), Windows
OpenProcess) into hermes_cli._subprocess_compat, have gateway.status's
fallback delegate to it, and import only stdlib-backed modules in the
watcher.

Fixes #124649.
2026-09-27 03:37:09 -07:00
teknium1
07be7d47ed fix(bootstrap-installer): write the shared receipt schema and resolve the commit without git
Hermes-Setup.exe replaced install.ps1's .hermes-bootstrap-complete with its
own copy: completedAtUnix instead of the completedAt every other writer and
reader uses (install.ps1, Electron main, hermes_cli/source_stamp.py,
scripts/verify-bootstrap-version-stamp.py), and a pinnedCommit from a bare
`git rev-parse HEAD`, which is null on a machine where git exists only as
the PM-staged binary install.ps1 uses.

The receipt now carries completedAt (UTC ISO-8601, milliseconds), and the
commit resolves like install.ps1's Stage-Complete: the pinned commit, else
the checkout HEAD read from its ref files (loose ref or packed-refs), else
the full sha install.ps1 already recorded (read BOM-tolerant).

Fixes #124949
2026-09-27 03:35:26 -07:00
teknium1
80512c90dc fix(website): batch the plugin stars probe at 100 repos per GraphQL request
A single request for all 317 catalog repos now exceeds GitHub's per-query
resource limit: the reply carries partial data plus an error, the script
prints 'Probed 317', and every repo past the limit keeps its stale count
(hindsight showed 26.5k stars while GitHub had 34.5k). Batch at 100; probed
live: 317/317 repos in 4 requests, no errors.
2026-09-27 03:34:56 -07:00
teknium1
e2abb460b6 fix(website): put the repo root on sys.path before importing hermes_yaml in the stars probe
The scheduled Build Skills Index run executes
`python website/scripts/fetch-plugin-stars.py --probe`, so sys.path[0] is
website/scripts and the repo-root `hermes_yaml` shim it now imports is not
importable. Every scheduled probe since 2026-09-24 died with
ModuleNotFoundError, the live plugin-stars.json froze at that date, and the
41 catalog entries added since have no star count: they sort to the bottom of
'Most starred' and show no star pill regardless of real stars.

extract-plugins.py already inserts REPO_ROOT; do the same here.
2026-09-27 03:34:56 -07:00
teknium1
afc05274a2 website: keep the plugin name whole; tier/version pills move to the chip row, tools as one ellipsized list
With Community + version + stars pinned beside the title at 3 columns the name
collapsed to 'hermes-r…' — the one element a card must never lose. The title
now owns its row (stars stay right-aligned); tier and version pills lead the
chip row. The Tools row renders the first three names as one comma-separated
run that ellipsizes as a whole plus '+N', instead of three shrinking chips that
read 'mnemosyn… mnemosyn… mnemosyn…' for same-prefix tools.
2026-09-27 03:34:51 -07:00
teknium1
97a3607838 feat(website): uniform plugin catalog cards (clamped text, capped tool list)
Every card in a Plugin Catalog grid now renders at one height, so rows line
up instead of staggering by description length, banner presence and tool
count (the Desktop section had 28 distinct card heights, 412px to 1485px).

- grid-auto-rows: 1fr + column-flex card; the detail block (facts, install
  box, links) sinks to the bottom edge of every card
- banner slot is always 120px: entries without an image get a placeholder
  (tier tint gradient + large muted category glyph); a broken image URL
  swaps to the placeholder instead of collapsing the slot
- title row is one line (ellipsis), description clamps to 3 lines with an
  ellipsis and reserves 3 lines when shorter; full text stays in `title`
- chip row and dates row are one fixed line each (right-edge fade mask on
  chips); dates row renders empty when undated
- facts block is exactly three fixed-height rows: Maintainer, Pinned (with
  `Requires` folded in as `· hermes >=x`), Tools — first 3 tool names as
  chips plus a dashed `+N` chip (full list in `title`), or a muted dash
- install command truncates with an ellipsis instead of scrolling

`capToolChips` in PluginCatalog/catalog.ts is the cap helper (first 3 + N).
Desktop picker embed (`?embed=picker`) uses the same card and stays uniform.
2026-09-27 03:34:51 -07:00
teknium1
a3f454a287 fix(build): keep scripts/build/inputs.py free of pm imports; accept musl targets in its grammar
The Nix agent derivation builds scripts/build/*.py from a fileset that
does not include pm/, so importing pm.store.ALL_TARGETS there failed
nix flake check with ModuleNotFoundError. Extend the local target
regex with linux-(x64|arm64)-musl instead.
2026-09-27 03:24:39 -07:00
teknium1
aba7deac48 test(e2e): musl host cell passes; drop its #123682 known_failure gate
On Alpine the installer now stages musl uv/Python/Node (a full install
completes once libstdc++ is present) and, on the stock bash/git/curl
image, refuses up front naming musl and libstdc++.
2026-09-27 03:24:39 -07:00
teknium1
842f162f7d fix(installer): refuse musl hosts without libstdc++ up front, naming it
PM's musl Node is the unofficial-builds musl archive, which links the
system libstdc++. On a stock Alpine (bash, git, curl) node and npm then
fail staged verification with raw relocation errors after the clone and
downloads. Check in the prerequisites stage and name the package.
2026-09-27 03:24:39 -07:00
teknium1
5cbca02510 test(pm): keep FFmpeg's musl gaps when the re-pin test narrows its gap table
The test replaced FFmpeg's gaps with only the bionic entry, which now
turns the musl targets back on and asks BtbN for a glibc build that
FFmpeg.fetch_url refuses on musl. Extend the real table instead.
2026-09-27 03:24:39 -07:00
teknium1
f5c3c7d70a test: trim musl coverage to two invariants
Keep one PM invariant (a musl ELF userland resolves to linux-<arch>-musl
and the default closure is satisfiable with musl-native artifacts) and
one installer invariant (uv_bootstrap_target picks musl uv from a musl
ELF interpreter even when ldd says glibc). Drop the lock-content and
URL-shape change-detectors.
2026-09-27 03:24:39 -07:00
teknium1
34d52883cb fix(installer): ELF interpreter decides musl in uv_bootstrap_target, as in pm/store
install.sh consulted ldd before the ELF interpreter while pm/store.py's
_is_musl_libc reads the native userland's ELF interpreter first, so a
glibc ldd (secondary toolchain, gcompat) could make the bootstrap stage
a different libc than PM later resolves. Read /bin/sh (then /bin/ls)
PT_INTERP first in both; ldd and the loader glob are fallbacks only.
2026-09-27 03:24:39 -07:00
teknium1
ee0ad5adb2 refactor(pm): expose MUSL_TARGETS publicly; drop the per-URL pin hash cache
packages.py and security_packages.py imported the private _MUSL_TARGETS
across modules; make it a public store constant next to ALL_TARGETS.
The _pin_tool hash memo is an unrelated optimisation, out of scope here.
2026-09-27 03:24:39 -07:00
teknium1
79e6f1960d fix(installer): restore executable bits on install.sh and gen-bootstrap-pins.py
The salvaged commits dropped both scripts from 100755 to 100644.
2026-09-27 03:24:39 -07:00
JoaoMarcos44
666c65c552 fix(build): accept PM musl targets in bundle inputs 2026-09-27 03:24:39 -07:00
JoaoMarcos44
87d1243623 fix(installer): detect musl uv when ldd is unavailable
Fall back to the musl loader if ldd cannot identify libc, while honoring
an explicit GNU libc report on hosts with a secondary musl toolchain.
Cover both paths against the pinned uv URL and digest.

Refs: #123682
2026-09-27 03:24:39 -07:00
JoaoMarcos44
24487f3db6 fix(pm): keep musl installs on compatible runtime artifacts
Skip glibc-linked FFmpeg on musl in both default install and update roots.
Select musl uv during the standalone shell bootstrap, and let the native
userland resolve libc before bootstrap Python build metadata.

Add regression checks for the closure, target precedence, and installer pins.
2026-09-27 03:24:39 -07:00
teknium1
5cdfa0bde6 fix(pm): import _MUSL_TARGETS where IronProxy resolves musl targets
security_packages.py referenced _MUSL_TARGETS without importing it, so
IronProxy.fetch_url raised NameError on linux-*-musl (the PR's own
test_portable_go_tools_resolve_the_generic_linux_archives caught it).
Review fix pushed to the contributor branch.
2026-09-27 03:24:39 -07:00
JoaoMarcos44
8371dd17f6 fix(pm): select native musl artifacts on Linux 2026-09-27 03:24:39 -07:00
teknium1
60e531cb52 fix(gateway): read every Hermes inline bootstrap's argv as its identity, on all OSes
The #107002 guard keeps an inline ``-c`` program's trailing argv as data. Every
Hermes launcher runs the entry point IN the ``-c`` process, so the guard hid
real gateways on every OS (#124318, #124588):

- the store launcher / Windows updater relaunch (_launchers.runtime_command)
- the published launcher script (POSIX shell launcher: every PM-install
  systemd/launchd gateway) and its Windows .cmd base64 wrapper
- the venv_sync re-entry, whose argv is assigned inside the source

gateway.status.inline_bootstrap_argv recognises exactly those emitted source
shapes, anchored at both ends so a program merely CARRYING one (the restart
watcher's respawn argv) still never matches, and rewrites the process to the
equivalent ``python -m <entry> <argv>``. /proc, psutil and ``ps`` space-join
argv, which splits the source across tokens; the shortest token run ending
in a recognised tail is the source whichever reader joined it. Both
canonical matchers (looks_like_gateway_command_line and
update_cmd_windows._hermes_holder_subcommand) use it.

Live on a real PM install (Linux, bwrap): a gateway started through the
installed launcher script or runtime_command read "Gateway is not running"
on main; with this change both read running, find_gateway_pids and
get_running_pid see them. Drops the four #124318 known_failure gates in
tests/e2e/core/windows_update.

Co-authored-by: Hermes Agent <dmyou@users.noreply.github.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: DianaBudin <dianabudin0307@gmail.com>
2026-09-27 02:57:28 -07:00
dgpcboy
02cf40d99e fix(gateway): recognise the Windows launcher's inline-bootstrap gateway form 2026-09-27 02:57:28 -07:00
teknium1
878d902147 fix(update): judge Windows gateway ownership by home, warn on unreadable ones
Rework of the #124676 salvage onto the canonical fleet scope instead of a
second path heuristic:

- the pause, the cold-start guard and the post-relaunch readiness poll all
  filter the host-wide find_gateway_pids(all_profiles=True) scan through
  update_cmd_fleet._scoped_manual_gateway_pids, the same home scope the
  POSIX fleet restart uses (#93349). A PM install's venv lives under
  installs/, so exe-under-PROJECT_ROOT rarely proved an own gateway; its
  live HERMES_HOME (or the LOCALAPPDATA default) always does.
- a foreign gateway with no HERMES_HOME resolves to its own default home and
  no longer blocks this install's cold start (#124659, second bullet).
- a gateway whose home cannot be read is named in the pause output instead
  of skipped silently; the spawn ledger's verified (pid, create_time) entry
  proves an own gateway whose environment is unreadable.
- tests trimmed to two invariants on real processes; a real-Windows journey
  (two installs, one updates while the other serves) replaces the harness
  comment that documented the bug.
2026-09-27 02:57:28 -07:00
JoaoMarcos44
46fba4c9a4 fix(update): scope Windows gateway lifecycle to current install 2026-09-27 02:57:28 -07:00
teknium1
5b27d6c2c2 test(e2e/desktop): let the build-fail updater finish its launch acceptance before asserting it exited
On the error path posix.sh writes the result, relaunches the app, and then
sleeps 1.5 s to confirm the launch took. A relaunched app that boots fast logs
`detached update FAILED` inside that window, and the spec checked for a live
posix.sh right away. On #124878 the app logged at 08:34:06.565, about 0.5 s
after it launched, the check ran at 08:34:07.16, and posix.sh was still in
`sleep 1.5`. The spec now polls for up to 30 s.
2026-09-27 02:56:08 -07:00
teknium1
388c63e24e test(e2e/pm): drop the #124668 gate; both generation_gc cells pass
hermes update and hermes pm repair now run the generation collector, so test_repeated_updates_collect_superseded_generations and test_repeated_repairs_collect_superseded_generations assert plainly again.
2026-09-27 02:55:39 -07:00
teknium1
ccff8eb3ce test(pm): keep the takeover generation assert; drop call-order checks
Trim the salvaged tests to one invariant: after a takeover rebuild, the
superseded generation built two days ago is gone. The update-completion
event-order asserts and the repair_dependencies `events == [sync,
collect]` test pin the call sequence, not the outcome; the update and
repair outcomes are covered live by
tests/e2e/core/upgrade/pm/test_generation_gc.py. The fake venv_sync in
the completion-process fixture keeps its collect_superseded_generations
stub so the module still imports.
2026-09-27 02:55:39 -07:00
JoaoMarcos44
c11a6bc0b2 fix(pm): collect generations after maintenance syncs 2026-09-27 02:55:39 -07:00
teknium1
3c9c224e6c test(e2e): the symlinked per-task home cell is no longer a known failure
The PM runtime identity now canonicalizes the store directories, so a launch
from a per-task HERMES_HOME whose tools/ links back to the main home is current.
Drop the #123798 known_failure wrap and its "gated" wording.
2026-09-27 02:55:05 -07:00