Commit Graph

274 Commits

Author SHA1 Message Date
teknium1
be84a141f6 test(e2e/pm): drop the #124214 gate in test_stamps_and_doctor
hermes doctor reports a dashboard web surface that cannot import.
2026-09-28 06:42:00 -07:00
teknium1
de7ed4cc3b test(e2e): drop the #124049, #124418 and #124635 gates
execute_code imports third-party deps on a PM install, the PM runtime restages
through a pip mirror, and a launch on a read-only install tree fails once naming
the permission problem; the cells now assert it.
2026-09-28 06:42:00 -07:00
teknium1
e13f4212c2 fix(gateway): launcher-started gateways persist an identifiable argv (#124029)
<bootstrap>` launcher, but the persisted record still carried raw sys.argv,
which under that launcher is ["-c", "gateway", "run"]. _record_looks_like_gateway
could never accept it, so readers that fall back to the record (unreadable live
cmdline on Windows/EACCES, pid-file validation) still treated a healthy gateway
as foreign, and _gateway_code_root found no checkout path in it.

_build_pid_record now records the argv `python -m hermes_cli.main` would have
(argv[0] = the checkout's hermes_cli/main.py) when argv[0] is "-c"; the rest of
argv (including --external-supervisor) is unchanged.

The handoff E2E `head` relaunch cell already passes on main after #124965, so
its RELAUNCH_GATES entry is removed (the n1 entry is #124649, a different bug).
2026-09-28 06:42:00 -07:00
teknium1
a153adcc2e fix(update): install the SDKs of configured gateway platforms instead of only warning (#124228)
A PM environment build keeps the recorded extras selection and never reads
config, so a home with platforms.discord (or telegram) enabled came out of
`hermes update` without the SDK: the update child detected it
(_configured_features_missing_deps) and only printed a warning, and the
gateway's lazy install at connect time is the only recovery — which fails
silently when lazy installs are off or the process is not on the selected
generation.

The update-build child (fresh, selected interpreter) now maps each missing
configured feature to the pyproject extra named after it (platform name ->
extra, MCP -> mcp), keeps only declared, platform-supported extras, and adds
them with an explicit pm.sync_venv. Detection reuses the gateway's own
connected-platform semantics (config + env credentials), so it installs exactly
what the gateway will try to load on restart. Features it could not install
still get the warning; an install failure never fails the update.

Flips tests/e2e/core/upgrade/pm/test_configured_features.py::
test_configured_gateway_platform_has_its_sdk_after_update.
2026-09-28 06:42:00 -07:00
teknium1
0b5bb59a42 test(e2e): flip the update gates for #124642, #124644, #122557 and #122627
The cells now assert the fixed behaviour: no false 'history diverged' or rescue
ref for an index.lock or a cut lazy fetch, a stopped rebase is refused, a parked
autostash is never reported complete, and 'hermes update --check' from the
managed environment works.
2026-09-28 05:33:25 -07:00
teknium1
01a5549faf test(e2e): shallow/partial-clone update cells assert the fixed behaviour
The full-clone shape (#122353), the second update after the unshallowing
update (#124272) and the --depth pre-fetch (#123346) cells run HEAD code and
now assert plainly. #123254 and #124645 live in N-1's pre-swap fetch and
check/pull code, so those gates stay until a release carrying the fix is N-1;
their reason text now says so.
2026-09-28 03:08:56 -07:00
teknium1
b22631ec4f test(e2e): drop stale gates for #124643 (detached HEAD) and #123424 (duplicate ~/.local/bin)
Both issues are already resolved on main and both cells pass there ungated; the gates only hid a
future regression.
2026-09-28 02:00:41 -07:00
teknium1
a31e618e70 fix(pm): download npm-hosted tools through the user's npm registry
pm/lock.json and the Npm/AgentBrowser templates record registry.npmjs.org, and
pinned_source handed that URL straight to the downloader, so ~/.npmrc /
npm_config_registry were ignored and closed networks could not provision npm or
agent-browser. pinned_source now routes public-npm URLs through the configured
registry (npm's own precedence: npm_config_registry, then the user npmrc); the
lock's SHA256 still verifies the bytes and progress keeps the lockfile URL.
npm_dist_tags reads from the same registry.

The E2E cell test_npm_registry_mirror_serves_pm_npm_download drops its gate.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-28 02:00:41 -07:00
teknium1
86f20a367a test: trim salvaged SSL/config unit tests to two invariants each; flip the #124654 and #119928 E2E gates
The E2E cells (test_update_with_corporate_root_only_in_ssl_cert_file, test_p2_env_lock_refusal_is_not_reported_as_success)
now assert the fixed behaviour directly.
2026-09-28 02:00:41 -07:00
teknium1
9b02a977bd test(e2e): assert the dashboard handoff cells the respawn fixes turned green
Removes FALSE_SUCCESS_GATE (#109290) and the N-1 #124778 dashboard gate: the respawn runs in the
updated checkout's code in both columns, and the N-1 column is 7/7 on this branch (5/7 on
origin/main with the gates off). Removes PID_FILE_GATE (#123109): the bootstrap launcher is already
classified by #124965 and the cell passes on origin/main in both columns.
2026-09-28 01:52:54 -07:00
teknium1
b63c138d78 fix: Hermes never falls back to the user's node/npm/npx/uv
Maintainer ruling: only Hermes and only its packaged package managers
(uv/pip/node/npm) are ever used; no PATH fallback when the managed tool is
missing, no "prefer the user's if new enough".

- hermes_constants.find_node_executable: node/npm/npx resolve to PM's
  installed copy or None. Every caller already pm.ensure()s on None, so a
  missing runtime is now provisioned instead of silently borrowing the
  user's Node (native-addon ABI / npm cache mismatches).
- agent/lsp/install._install_npm: pm.ensure('npm') when PM npm is absent,
  instead of failing over to whatever npm is on PATH.
- gateway._append_node_dir_for_service: stop baking the invoker's PATH node
  dir into generated systemd/launchd units.
- main_install_repair._resolve_node_runtime_npm: drop the PATH re-scan for
  another npm.
- source_build.source_product_current: run the freshness reader only with
  PM's node.
- doctor: Node/npm rows and npm audit use PM's copies (Termux APT distro
  keeps its system Node).
- install.sh ensure_uv / install.ps1 Get-Uv: always stage the pinned uv
  artifact; delete the "uv on PATH if new enough" developer shortcut.
2026-09-27 22:04:26 -07:00
teknium1
5b770c8bf5 fix(mcp): name the rebuild when a stdio server's native addon was built for another Node
Hermes runs stdio MCP servers on its own packaged Node only. A server whose
native addon was compiled by the user's Node (typically through the ~/.npm/_npx
cache the user's npx shares with ours) dies at startup with NODE_MODULE_VERSION /
ERR_DLOPEN_FAILED; the SDK only saw "Connection closed", Hermes retried three
times and parked the server, and the cause lived only in logs/mcp-stderr.log.

The stdio child's stderr now goes through a pipe that still copies every byte
into the shared log but keeps the last 16 KB readable. When the session fails
and that tail shows an addon load failure, the error becomes a
NodeAbiMismatchError naming the addon, both ABI versions and the remedy with
Hermes's own paths: delete the npx cache entry (Hermes's npx reinstalls it) or
`PATH=<managed node dir>:"$PATH" <managed npm> rebuild <pkg> --prefix <root>`.
It is classified permanent, so the server parks at once and self-probes back
after the rebuild. The message reaches every MCP status surface through
_format_connect_error / str(exc): startup banner, `hermes mcp test`, the TUI
and Desktop probes and the dashboard.

The hosts E2E cell that asserted the user's Node wins now asserts the ruling:
the managed Node runs the server and the ABI failure surfaces the remedy.

Refs #124264
2026-09-27 21:30:10 -07:00
teknium1
3249512632 test(e2e/handoff): drop the HEAD column's foreign-unit dashboard gate
#124940 makes the update restart a dashboard's systemd unit only when the
unit's MainPID is the dashboard, so at HEAD a dashboard started inside the
runner's hosted-compute-agent.service is respawned, never unit-restarted
(the CI logs of the failing runs show the respawn and the dashboard cell
passing). The N-1 column keeps the gate until a release carrying the fix
is N-1.
2026-09-27 21:29:39 -07:00
Brooklyn Nicholson
7a21d15c35 feat(tui): add attachments.storage config for workspace attachment staging
Desktop file attachments stage into <profile home>/attachments, which sits
outside the session workspace, so an agent with a restricted workspace can be
unable to read the very files the user attached to it (#110662).

Add a top-level `attachments.storage` config key ("hermes-home" default |
"workspace" opt-in). The gateway reads it per profile from the session's
profile config.yaml — file.attach runs before prompt.submit installs the
profile scope, so the process config still belongs to the launch profile
(same reason as _profile_configured_cwd). Opting in stages attachments under
<workspace>/.hermes/attachments, inside the allowed ref root, so the @file:
ref stays workspace-relative and the same profile's agent can always read its
own attachments back.

A remote (ssh) profile keeps the profile home dir either way: its workspace
lives on the execution host, and the bind-mounted <profile home>/attachments
is what container and remote backends receive (#76577). Traversal hardening
and staging name sanitization are unchanged.

Fixes https://github.com/NousResearch/hermes-agent/issues/110662
2026-09-27 19:07:07 -05:00
Brooklyn Nicholson
a0ec0d286f feat(cli): /s as a one-letter alias for /steer
/queue has /q; /steer forced users to type the full word mid-run just to
inject a line after the next tool call. Add the same one-letter alias in
the central registry (every surface derives from it) and mirror it in the
TUI's local slash registry, the way /q is mirrored — checked that no
TUI-local /s binding shadows it (the /q vs /quit collision, #31983).

/i for interrupt stays out: /interrupt is not a registered command yet
(#69954 holds that decision), and the issue gates the alias on its
existence.

Fixes https://github.com/NousResearch/hermes-agent/issues/119176
2026-09-27 19:04:20 -05:00
teknium1
d25bbd01b7 test(e2e/network): drop the #124653 gate; the channel 503 cell passes
The known_failure wrapper on test_channel_503_with_retry_after_is_retried_then_updates keyed on the one-attempt channel read. With update-path retries it passes, so the cell asserts plainly again.
2026-09-27 03:41:42 -07:00
teknium1
aba7deac48 test(e2e): musl host cell passes; drop its #123682 known_failure gate
On Alpine the installer now stages musl uv/Python/Node (a full install
completes once libstdc++ is present) and, on the stock bash/git/curl
image, refuses up front naming musl and libstdc++.
2026-09-27 03:24:39 -07:00
teknium1
60e531cb52 fix(gateway): read every Hermes inline bootstrap's argv as its identity, on all OSes
The #107002 guard keeps an inline ``-c`` program's trailing argv as data. Every
Hermes launcher runs the entry point IN the ``-c`` process, so the guard hid
real gateways on every OS (#124318, #124588):

- the store launcher / Windows updater relaunch (_launchers.runtime_command)
- the published launcher script (POSIX shell launcher: every PM-install
  systemd/launchd gateway) and its Windows .cmd base64 wrapper
- the venv_sync re-entry, whose argv is assigned inside the source

gateway.status.inline_bootstrap_argv recognises exactly those emitted source
shapes, anchored at both ends so a program merely CARRYING one (the restart
watcher's respawn argv) still never matches, and rewrites the process to the
equivalent ``python -m <entry> <argv>``. /proc, psutil and ``ps`` space-join
argv, which splits the source across tokens; the shortest token run ending
in a recognised tail is the source whichever reader joined it. Both
canonical matchers (looks_like_gateway_command_line and
update_cmd_windows._hermes_holder_subcommand) use it.

Live on a real PM install (Linux, bwrap): a gateway started through the
installed launcher script or runtime_command read "Gateway is not running"
on main; with this change both read running, find_gateway_pids and
get_running_pid see them. Drops the four #124318 known_failure gates in
tests/e2e/core/windows_update.

Co-authored-by: Hermes Agent <dmyou@users.noreply.github.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: DianaBudin <dianabudin0307@gmail.com>
2026-09-27 02:57:28 -07:00
teknium1
878d902147 fix(update): judge Windows gateway ownership by home, warn on unreadable ones
Rework of the #124676 salvage onto the canonical fleet scope instead of a
second path heuristic:

- the pause, the cold-start guard and the post-relaunch readiness poll all
  filter the host-wide find_gateway_pids(all_profiles=True) scan through
  update_cmd_fleet._scoped_manual_gateway_pids, the same home scope the
  POSIX fleet restart uses (#93349). A PM install's venv lives under
  installs/, so exe-under-PROJECT_ROOT rarely proved an own gateway; its
  live HERMES_HOME (or the LOCALAPPDATA default) always does.
- a foreign gateway with no HERMES_HOME resolves to its own default home and
  no longer blocks this install's cold start (#124659, second bullet).
- a gateway whose home cannot be read is named in the pause output instead
  of skipped silently; the spawn ledger's verified (pid, create_time) entry
  proves an own gateway whose environment is unreadable.
- tests trimmed to two invariants on real processes; a real-Windows journey
  (two installs, one updates while the other serves) replaces the harness
  comment that documented the bug.
2026-09-27 02:57:28 -07:00
teknium1
388c63e24e test(e2e/pm): drop the #124668 gate; both generation_gc cells pass
hermes update and hermes pm repair now run the generation collector, so test_repeated_updates_collect_superseded_generations and test_repeated_repairs_collect_superseded_generations assert plainly again.
2026-09-27 02:55:39 -07:00
teknium1
3c9c224e6c test(e2e): the symlinked per-task home cell is no longer a known failure
The PM runtime identity now canonicalizes the store directories, so a launch
from a per-task HERMES_HOME whose tools/ links back to the main home is current.
Drop the #123798 known_failure wrap and its "gated" wording.
2026-09-27 02:55:05 -07:00
teknium1
c1e4d4d274 test(e2e): the managed-environment pm cell is a plain assertion again
The generation workspace now carries pm/uv.lock, so
test_pm_commands_run_from_the_managed_environment passes; drop its
known_failure wrap as tests/e2e/core/_pending_fixes.py asks.

Refs #124075
2026-09-27 02:54:35 -07:00
teknium1
95fdcf92b0 test(e2e): enforce the Windows update cells the PM git fix flips
The fresh-install update and the partial-clone update no longer fail with
[WinError 2] on a machine without system Git (#124634), so their
known_failure gates come off.
2026-09-27 02:54:03 -07:00
teknium1
679a7a7314 test(e2e): enforce the untracked-collision cell the autostash fix flips
An untracked file at a path upstream adds is now kept in the parked stash
(#124641), so its known_failure gate comes off.
2026-09-27 02:53:32 -07:00
teknium1
fb2dded3d1 test(e2e/handoff): gate the dashboard respawn on #124938 when the runner's systemd unit owns the cgroup
On CI every job runs inside hosted-compute-agent.service, so the hand-started
dashboard sits in that unit's cgroup and hermes update restarts the unit
instead of respawning the dashboard (both columns). The verdict names the unit
the update restarted, the N-1 column accepts #124778 or #124938, and the
failure report shows the sandbox's cgroup. Reproduced locally by running the
suite inside a transient systemd user service.
2026-09-27 01:27:09 -07:00
teknium1
c02f6b0f5c test(e2e/handoff): processes running across hermes update (gateway, dashboard, cron, kanban)
Real hermes update in a bwrap/PID-namespace sandbox with a manual gateway, a
manual dashboard, a cron job due mid-update and an in-flight kanban worker, from
N-1 (v2026.9.24) and from HEAD->NEXT; plus refused-fetch cells. Open bugs are
message-gated: gated on #124649, #124029, #123109, #124778, #113293, #109290.
2026-09-27 01:27:09 -07:00
teknium1
81f873b334 test(e2e/pm): PM environment lifecycle across real hermes updates
tests/e2e/core/upgrade/pm/: 22 cells in 7 files, one per failure class, each
driving the real entry points (install.sh, hermes update, hermes pm
repair/status/install, hermes doctor, hermes gateway run/install --force,
hermes kanban dispatch, hermes -z through the loopback provider) inside the
existing bwrap sandbox, seeded with an N-1 install and a dependency-changing
release:

- configured features survive update, repair and legacy-venv migration
- SIGKILL at the stage/publish boundary of a dependency update recovers
- stamps, receipts and doctor tell the truth (healthy, drifted, failed update)
- no stray in-tree venv / repo-local 3.11 .venv reaches any interpreter;
  execute_code and Kanban workers spawned after an update boot and import deps
- repeated updates/repairs/restarts don't accumulate generations
- gateway install --force keeps the install bootable

Every non-gated cell has a recorded sabotage proof; open bugs are
message-gated with known_failure (gated on #124228, #122627, #122425,
#124214, #124075, #124049, #124668).

Harness: make_origin clones --single-branch and allows filters, so a
blobless developer checkout can serve the installer's clone.

CI: e2e-upgrade is sharded per suite directory (an e2e-upgrade-plan job
lists "core" plus every subdirectory holding test_*.py); a shard that
collects no files fails; HERMES_TEST_WORKERS=6.
2026-09-27 00:42:13 -07:00
teknium1
8d70712124 test(e2e/windows): snapshot the installed checkout before the update moves it 2026-09-27 00:41:41 -07:00
teknium1
3b05737cad test(e2e/windows): drop the managed-Python cell (no sabotage drives it red for the right reason) 2026-09-27 00:41:41 -07:00
teknium1
1e6a3bbeab test(e2e/windows): inline message gates, AccessCheck-based standard-user probe, strict status check 2026-09-27 00:41:41 -07:00
teknium1
2f2e465f1d test(e2e/windows): stamped transcripts, claims log, observed-only gates, standard-user control dir 2026-09-27 00:41:41 -07:00
teknium1
16da7f1b38 test(e2e/windows): real C:\Users profiles, serialized gateway phases, Git-for-Windows machines 2026-09-27 00:41:41 -07:00
teknium1
9129897dd1 test(e2e/windows): PR-time install.ps1 -> hermes update journey on native Windows 2026-09-27 00:41:41 -07:00
teknium1
b7990b6f7e test(e2e/hosts): split symlinked-home cell into its own file; gate #123424 on the PATH count only 2026-09-27 00:40:15 -07:00
teknium1
04b13e3c00 test(e2e/hosts): read-only checkout for update refusal; fold symlinked-home check into per-task cell; order shim-log check first 2026-09-27 00:40:15 -07:00
teknium1
b15bec7e5e test(e2e/hosts): install/update on odd host shapes 2026-09-27 00:40:15 -07:00
teknium1
024e0cfefa test(e2e/network): drop the backend re-run; name the partial-clone cause in the edge log 2026-09-27 00:39:45 -07:00
teknium1
7f97475689 test(e2e/network): fake forge survives a concurrent repack of the borrowed object store; log backend stderr 2026-09-27 00:39:45 -07:00
teknium1
3063a7a4d9 test(e2e/network): assert the unchanged-install contract before message wording 2026-09-27 00:39:45 -07:00
teknium1
c528509f65 test(e2e/network): gate on filed issues 2026-09-27 00:39:45 -07:00
teknium1
11a84d7d71 test(e2e/network): hostile-network install/update suite under a private netns 2026-09-27 00:39:45 -07:00
teknium1
fb00c0ba3b test(e2e/git): cut faulted responses at a fixed pkt-line framing
The interrupted-fetch cell went red on CI only ("the interrupted fetches left no
temp packs"). upload-pack relays whatever pack-objects has buffered, so the same
pack arrives as 8 KiB sideband packets when it keeps up and as 65520-byte ones
when it lags (a loaded runner). The client acts only on whole pkt-lines: a
20000-byte cut inside a coalesced first packet means fetch-pack never sees the
PACK header ("bad pack header", "45532 bytes of body are still expected") and
never starts index-pack, so no temp pack exists. hermes update failed cleanly
all three times; only the cell's precondition depended on the timing.

Faulted (cut/stall) responses now have their sideband-1 pack data re-framed to
8 KiB pkt-lines before the byte limit applies (a valid framing, pack bytes
untouched), and the request log records the pack bytes the client received in
whole pkt-lines before the drop, so the precondition failure explains itself.

Proof: coalescing upstream packets to 65515 bytes (a lagging upload-pack)
without re-framing reproduces the CI message exactly; with re-framing the cell
passes.
2026-09-27 00:39:16 -07:00
teknium1
a39de43f85 test(e2e/git): link the #124641 gate to its fix PR 2026-09-27 00:39:16 -07:00
teknium1
59b14f69a0 test(e2e/git): git transport and working-tree edge cases through the real hermes update over smart HTTP 2026-09-27 00:39:16 -07:00
teknium1
4127d78da8 test(e2e): TUI /exit re-signals tmux when it leaves the exited pane unreaped
CI diagnostics (tmux 3.4 on ubuntu-24.04) show the /exit hang is tmux's:
after /exit the pane process is a single-threaded zombie of the tmux
server (Threads:1, PPid = tmux server), the server is running with
SIGCHLD caught, not blocked and not pending, and pane_dead_status never
fills for 60s. The TUI had already exited (its epilogue is in the PTY
transcript). A zombie tmux has not reaped after 2s now gets a SIGCHLD
nudge so tmux runs its waitpid loop and reports the real exit status; a
process that is still running is never touched and still times out.
2026-09-26 13:18:16 -07:00
teknium1
13bf2ebaaa test(e2e): TUI clarify cell tolerates the API-only first-contact note; exit failures dump threads + tmux server signals
Main now appends the install's first-contact onboarding note to the first
user message on the wire only (per-turn sidecar, never persisted). The
clarify cell's invariant is that the answer is not a second user turn, so
compare what the user typed. An un-reaped pane zombie on CI now reports
the pane process threads (state/wchan) and the tmux server's signal masks.
2026-09-26 13:18:16 -07:00
teknium1
ceecaa98f1 test(e2e): TUI /exit waits for tmux to reap the pane before reading its exit status
tmux marks a pane dead on pty EOF, which the kernel delivers when the exiting
process closes its last tty fd -- before SIGCHLD lets tmux reap it and fill
pane_dead_status. On a loaded CI runner the harness read the status in that
gap and failed first_exits_clean with an empty '/exit status '. Poll until
tmux reports the exit status or signal, and make every exit failure report
status/signal, raw tmux answer, pane pid state, elapsed time, the frame
before /exit and the tail of a pipe-pane PTY transcript.
2026-09-26 13:18:16 -07:00
teknium1
9acdab8f87 test(e2e): docstring wording 2026-09-26 13:18:16 -07:00
teknium1
f0640122e4 test(e2e): harden Ink TUI tmux suite per review
- wait for the startup session (status bar 'ready') before the first submit;
  PHASE names in harness errors
- known_failure pins (merge-order safe) instead of strict xfail; new cell for
  /resume typed during startup being undone (#121456)
- verbose tool progress so tool output is rendered and checked exactly once
- width cells check the paragraph layout (whole words, rows read on) so a
  stale-width frame goes red on shrink
- tmux socket under the test root (-S), removed on close
- persisted/screen, summary, exit and raw interrupt-partial checks
2026-09-26 13:18:16 -07:00
teknium1
cccadaf170 test(e2e): TUI tmux harness uses a private prebuilt bundle; retry empty tmux queries 2026-09-26 13:18:16 -07:00