execute_code imports third-party deps on a PM install, the PM runtime restages
through a pip mirror, and a launch on a read-only install tree fails once naming
the permission problem; the cells now assert it.
<bootstrap>` launcher, but the persisted record still carried raw sys.argv,
which under that launcher is ["-c", "gateway", "run"]. _record_looks_like_gateway
could never accept it, so readers that fall back to the record (unreadable live
cmdline on Windows/EACCES, pid-file validation) still treated a healthy gateway
as foreign, and _gateway_code_root found no checkout path in it.
_build_pid_record now records the argv `python -m hermes_cli.main` would have
(argv[0] = the checkout's hermes_cli/main.py) when argv[0] is "-c"; the rest of
argv (including --external-supervisor) is unchanged.
The handoff E2E `head` relaunch cell already passes on main after #124965, so
its RELAUNCH_GATES entry is removed (the n1 entry is #124649, a different bug).
A PM environment build keeps the recorded extras selection and never reads
config, so a home with platforms.discord (or telegram) enabled came out of
`hermes update` without the SDK: the update child detected it
(_configured_features_missing_deps) and only printed a warning, and the
gateway's lazy install at connect time is the only recovery — which fails
silently when lazy installs are off or the process is not on the selected
generation.
The update-build child (fresh, selected interpreter) now maps each missing
configured feature to the pyproject extra named after it (platform name ->
extra, MCP -> mcp), keeps only declared, platform-supported extras, and adds
them with an explicit pm.sync_venv. Detection reuses the gateway's own
connected-platform semantics (config + env credentials), so it installs exactly
what the gateway will try to load on restart. Features it could not install
still get the warning; an install failure never fails the update.
Flips tests/e2e/core/upgrade/pm/test_configured_features.py::
test_configured_gateway_platform_has_its_sdk_after_update.
The cells now assert the fixed behaviour: no false 'history diverged' or rescue
ref for an index.lock or a cut lazy fetch, a stopped rebase is refused, a parked
autostash is never reported complete, and 'hermes update --check' from the
managed environment works.
The full-clone shape (#122353), the second update after the unshallowing
update (#124272) and the --depth pre-fetch (#123346) cells run HEAD code and
now assert plainly. #123254 and #124645 live in N-1's pre-swap fetch and
check/pull code, so those gates stay until a release carrying the fix is N-1;
their reason text now says so.
pm/lock.json and the Npm/AgentBrowser templates record registry.npmjs.org, and
pinned_source handed that URL straight to the downloader, so ~/.npmrc /
npm_config_registry were ignored and closed networks could not provision npm or
agent-browser. pinned_source now routes public-npm URLs through the configured
registry (npm's own precedence: npm_config_registry, then the user npmrc); the
lock's SHA256 still verifies the bytes and progress keeps the lockfile URL.
npm_dist_tags reads from the same registry.
The E2E cell test_npm_registry_mirror_serves_pm_npm_download drops its gate.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
The E2E cells (test_update_with_corporate_root_only_in_ssl_cert_file, test_p2_env_lock_refusal_is_not_reported_as_success)
now assert the fixed behaviour directly.
Removes FALSE_SUCCESS_GATE (#109290) and the N-1 #124778 dashboard gate: the respawn runs in the
updated checkout's code in both columns, and the N-1 column is 7/7 on this branch (5/7 on
origin/main with the gates off). Removes PID_FILE_GATE (#123109): the bootstrap launcher is already
classified by #124965 and the cell passes on origin/main in both columns.
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".
- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
installed copy or None. Every caller already pm.ensure()s on None, so a
missing runtime is now provisioned instead of silently borrowing the
user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
another npm.
- source_build.source_product_current: run the freshness reader only with
PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
artifact; delete the "uv on PATH if new enough" developer shortcut.
Hermes runs stdio MCP servers on its own packaged Node only. A server whose
native addon was compiled by the user's Node (typically through the ~/.npm/_npx
cache the user's npx shares with ours) dies at startup with NODE_MODULE_VERSION /
ERR_DLOPEN_FAILED; the SDK only saw "Connection closed", Hermes retried three
times and parked the server, and the cause lived only in logs/mcp-stderr.log.
The stdio child's stderr now goes through a pipe that still copies every byte
into the shared log but keeps the last 16 KB readable. When the session fails
and that tail shows an addon load failure, the error becomes a
NodeAbiMismatchError naming the addon, both ABI versions and the remedy with
Hermes's own paths: delete the npx cache entry (Hermes's npx reinstalls it) or
`PATH=<managed node dir>:"$PATH" <managed npm> rebuild <pkg> --prefix <root>`.
It is classified permanent, so the server parks at once and self-probes back
after the rebuild. The message reaches every MCP status surface through
_format_connect_error / str(exc): startup banner, `hermes mcp test`, the TUI
and Desktop probes and the dashboard.
The hosts E2E cell that asserted the user's Node wins now asserts the ruling:
the managed Node runs the server and the ABI failure surfaces the remedy.
Refs #124264
#124940 makes the update restart a dashboard's systemd unit only when the
unit's MainPID is the dashboard, so at HEAD a dashboard started inside the
runner's hosted-compute-agent.service is respawned, never unit-restarted
(the CI logs of the failing runs show the respawn and the dashboard cell
passing). The N-1 column keeps the gate until a release carrying the fix
is N-1.
Desktop file attachments stage into <profile home>/attachments, which sits
outside the session workspace, so an agent with a restricted workspace can be
unable to read the very files the user attached to it (#110662).
Add a top-level `attachments.storage` config key ("hermes-home" default |
"workspace" opt-in). The gateway reads it per profile from the session's
profile config.yaml — file.attach runs before prompt.submit installs the
profile scope, so the process config still belongs to the launch profile
(same reason as _profile_configured_cwd). Opting in stages attachments under
<workspace>/.hermes/attachments, inside the allowed ref root, so the @file:
ref stays workspace-relative and the same profile's agent can always read its
own attachments back.
A remote (ssh) profile keeps the profile home dir either way: its workspace
lives on the execution host, and the bind-mounted <profile home>/attachments
is what container and remote backends receive (#76577). Traversal hardening
and staging name sanitization are unchanged.
Fixes https://github.com/NousResearch/hermes-agent/issues/110662
/queue has /q; /steer forced users to type the full word mid-run just to
inject a line after the next tool call. Add the same one-letter alias in
the central registry (every surface derives from it) and mirror it in the
TUI's local slash registry, the way /q is mirrored — checked that no
TUI-local /s binding shadows it (the /q vs /quit collision, #31983).
/i for interrupt stays out: /interrupt is not a registered command yet
(#69954 holds that decision), and the issue gates the alias on its
existence.
Fixes https://github.com/NousResearch/hermes-agent/issues/119176
The known_failure wrapper on test_channel_503_with_retry_after_is_retried_then_updates keyed on the one-attempt channel read. With update-path retries it passes, so the cell asserts plainly again.
On Alpine the installer now stages musl uv/Python/Node (a full install
completes once libstdc++ is present) and, on the stock bash/git/curl
image, refuses up front naming musl and libstdc++.
The #107002 guard keeps an inline ``-c`` program's trailing argv as data. Every
Hermes launcher runs the entry point IN the ``-c`` process, so the guard hid
real gateways on every OS (#124318, #124588):
- the store launcher / Windows updater relaunch (_launchers.runtime_command)
- the published launcher script (POSIX shell launcher: every PM-install
systemd/launchd gateway) and its Windows .cmd base64 wrapper
- the venv_sync re-entry, whose argv is assigned inside the source
gateway.status.inline_bootstrap_argv recognises exactly those emitted source
shapes, anchored at both ends so a program merely CARRYING one (the restart
watcher's respawn argv) still never matches, and rewrites the process to the
equivalent ``python -m <entry> <argv>``. /proc, psutil and ``ps`` space-join
argv, which splits the source across tokens; the shortest token run ending
in a recognised tail is the source whichever reader joined it. Both
canonical matchers (looks_like_gateway_command_line and
update_cmd_windows._hermes_holder_subcommand) use it.
Live on a real PM install (Linux, bwrap): a gateway started through the
installed launcher script or runtime_command read "Gateway is not running"
on main; with this change both read running, find_gateway_pids and
get_running_pid see them. Drops the four #124318 known_failure gates in
tests/e2e/core/windows_update.
Co-authored-by: Hermes Agent <dmyou@users.noreply.github.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: DianaBudin <dianabudin0307@gmail.com>
Rework of the #124676 salvage onto the canonical fleet scope instead of a
second path heuristic:
- the pause, the cold-start guard and the post-relaunch readiness poll all
filter the host-wide find_gateway_pids(all_profiles=True) scan through
update_cmd_fleet._scoped_manual_gateway_pids, the same home scope the
POSIX fleet restart uses (#93349). A PM install's venv lives under
installs/, so exe-under-PROJECT_ROOT rarely proved an own gateway; its
live HERMES_HOME (or the LOCALAPPDATA default) always does.
- a foreign gateway with no HERMES_HOME resolves to its own default home and
no longer blocks this install's cold start (#124659, second bullet).
- a gateway whose home cannot be read is named in the pause output instead
of skipped silently; the spawn ledger's verified (pid, create_time) entry
proves an own gateway whose environment is unreadable.
- tests trimmed to two invariants on real processes; a real-Windows journey
(two installs, one updates while the other serves) replaces the harness
comment that documented the bug.
hermes update and hermes pm repair now run the generation collector, so test_repeated_updates_collect_superseded_generations and test_repeated_repairs_collect_superseded_generations assert plainly again.
The PM runtime identity now canonicalizes the store directories, so a launch
from a per-task HERMES_HOME whose tools/ links back to the main home is current.
Drop the #123798 known_failure wrap and its "gated" wording.
The generation workspace now carries pm/uv.lock, so
test_pm_commands_run_from_the_managed_environment passes; drop its
known_failure wrap as tests/e2e/core/_pending_fixes.py asks.
Refs #124075
The fresh-install update and the partial-clone update no longer fail with
[WinError 2] on a machine without system Git (#124634), so their
known_failure gates come off.
On CI every job runs inside hosted-compute-agent.service, so the hand-started
dashboard sits in that unit's cgroup and hermes update restarts the unit
instead of respawning the dashboard (both columns). The verdict names the unit
the update restarted, the N-1 column accepts #124778 or #124938, and the
failure report shows the sandbox's cgroup. Reproduced locally by running the
suite inside a transient systemd user service.
Real hermes update in a bwrap/PID-namespace sandbox with a manual gateway, a
manual dashboard, a cron job due mid-update and an in-flight kanban worker, from
N-1 (v2026.9.24) and from HEAD->NEXT; plus refused-fetch cells. Open bugs are
message-gated: gated on #124649, #124029, #123109, #124778, #113293, #109290.
tests/e2e/core/upgrade/pm/: 22 cells in 7 files, one per failure class, each
driving the real entry points (install.sh, hermes update, hermes pm
repair/status/install, hermes doctor, hermes gateway run/install --force,
hermes kanban dispatch, hermes -z through the loopback provider) inside the
existing bwrap sandbox, seeded with an N-1 install and a dependency-changing
release:
- configured features survive update, repair and legacy-venv migration
- SIGKILL at the stage/publish boundary of a dependency update recovers
- stamps, receipts and doctor tell the truth (healthy, drifted, failed update)
- no stray in-tree venv / repo-local 3.11 .venv reaches any interpreter;
execute_code and Kanban workers spawned after an update boot and import deps
- repeated updates/repairs/restarts don't accumulate generations
- gateway install --force keeps the install bootable
Every non-gated cell has a recorded sabotage proof; open bugs are
message-gated with known_failure (gated on #124228, #122627, #122425,
#124214, #124075, #124049, #124668).
Harness: make_origin clones --single-branch and allows filters, so a
blobless developer checkout can serve the installer's clone.
CI: e2e-upgrade is sharded per suite directory (an e2e-upgrade-plan job
lists "core" plus every subdirectory holding test_*.py); a shard that
collects no files fails; HERMES_TEST_WORKERS=6.
The interrupted-fetch cell went red on CI only ("the interrupted fetches left no
temp packs"). upload-pack relays whatever pack-objects has buffered, so the same
pack arrives as 8 KiB sideband packets when it keeps up and as 65520-byte ones
when it lags (a loaded runner). The client acts only on whole pkt-lines: a
20000-byte cut inside a coalesced first packet means fetch-pack never sees the
PACK header ("bad pack header", "45532 bytes of body are still expected") and
never starts index-pack, so no temp pack exists. hermes update failed cleanly
all three times; only the cell's precondition depended on the timing.
Faulted (cut/stall) responses now have their sideband-1 pack data re-framed to
8 KiB pkt-lines before the byte limit applies (a valid framing, pack bytes
untouched), and the request log records the pack bytes the client received in
whole pkt-lines before the drop, so the precondition failure explains itself.
Proof: coalescing upstream packets to 65515 bytes (a lagging upload-pack)
without re-framing reproduces the CI message exactly; with re-framing the cell
passes.
CI diagnostics (tmux 3.4 on ubuntu-24.04) show the /exit hang is tmux's:
after /exit the pane process is a single-threaded zombie of the tmux
server (Threads:1, PPid = tmux server), the server is running with
SIGCHLD caught, not blocked and not pending, and pane_dead_status never
fills for 60s. The TUI had already exited (its epilogue is in the PTY
transcript). A zombie tmux has not reaped after 2s now gets a SIGCHLD
nudge so tmux runs its waitpid loop and reports the real exit status; a
process that is still running is never touched and still times out.
Main now appends the install's first-contact onboarding note to the first
user message on the wire only (per-turn sidecar, never persisted). The
clarify cell's invariant is that the answer is not a second user turn, so
compare what the user typed. An un-reaped pane zombie on CI now reports
the pane process threads (state/wchan) and the tmux server's signal masks.
tmux marks a pane dead on pty EOF, which the kernel delivers when the exiting
process closes its last tty fd -- before SIGCHLD lets tmux reap it and fill
pane_dead_status. On a loaded CI runner the harness read the status in that
gap and failed first_exits_clean with an empty '/exit status '. Poll until
tmux reports the exit status or signal, and make every exit failure report
status/signal, raw tmux answer, pane pid state, elapsed time, the frame
before /exit and the tail of a pipe-pane PTY transcript.
- wait for the startup session (status bar 'ready') before the first submit;
PHASE names in harness errors
- known_failure pins (merge-order safe) instead of strict xfail; new cell for
/resume typed during startup being undone (#121456)
- verbose tool progress so tool output is rendered and checked exactly once
- width cells check the paragraph layout (whole words, rows read on) so a
stale-width frame goes red on shrink
- tmux socket under the test root (-S), removed on close
- persisted/screen, summary, exit and raw interrupt-partial checks