Commit Graph

45579 Commits

Author SHA1 Message Date
teknium1
e790ef4e31 fix(cron): trim owner-adapter resolution comment to the why (#124248, salvage #116316) 2026-09-28 05:28:05 -07:00
Fabio Ito
7d7166be5e fix(cron): do not fall back to the default adapters when owner-profile resolution fails
Review feedback: the previous try/except restored runner.adapters on any
resolution error, which is the cross-profile misdelivery this fix
prevents. Let the error propagate to _run_claimed_job's handler, which
marks the run failed and surfaces the message. Runners without
_adapters_for_profile (shims, tests) still keep runner.adapters.

Adds a regression test where resolution raises: run not fired, marked
failed, error surfaced.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 183547ea6b3798874dc8a98f943ef98ee710911b)
2026-09-28 05:28:05 -07:00
Fabio Ito
8be50587d4 fix(cron): deliver manual runs through the owner profile's adapters under multiplex
A manual `cronjob(action="run")` fired from a secondary profile's agent
resolved delivery through `runner.adapters`, which is the default
profile's map. Under `gateway.multiplex_profiles` HERMES_HOME is
overridden to the owner profile mid-turn, so the result left through the
default profile's bot (a Telegram DM arrived from the wrong bot while the
run was correctly recorded on the secondary profile's cron).

Resolve the owner profile from HERMES_HOME and use
`runner._adapters_for_profile()`, the same fail-closed path notifications
and goal loops already use. Runners without that method (tests, older
shims) keep the previous behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 350cfc1436504b92d5f9e924f43f81b232beefd7)
2026-09-28 05:28:05 -07:00
teknium1
dd00040182 fix(cron): liveness answers "is a scheduler ticking THIS profile" and names it (#121881, #99631, #99579, salvage #121921 #99613)
Trim of the #121921 salvage plus the rest of the class:

- `_builtin_gateway_liveness` (shared by `cron list/create`, `cron status` and the
  `cronjob` tool via tools/cronjob_job_args.py): own-home lock/pid, else a fresh
  ticker heartbeat in THIS home's store. The multiplexer-record rung was redundant
  once the heartbeat is the evidence (served + no heartbeat was already False), and
  it wrongly said "dead" when the record excluded the profile or predated
  `served_profiles` while `hermes serve` / the Desktop backend was ticking the home
  in-process (#121881) or a shared gateway was firing the profile's jobs (#99631).
- `cron status`: the in-process ticker is a rung beside the multiplexer rung instead
  of a re-indented copy of the "no gateway" block; its restart hint names the
  Desktop app / serve backend.
- The red verdict names the inspected profile and claims only what the probe can
  know ("No scheduler is serving profile 'X'"): from a named profile it cannot see a
  gateway alive for other profiles, so "No gateway is running on this host" was a
  host-wide negative it could not make (#99579). `cron list` names the profile in
  its banner and its empty-state line, and the list/create warning names it too
  (idea from #99613 by @webtecnica; its `_active_profile_label` wrapper, banner
  rewrite and unrelated kanban/script-validation scope were not taken).

WHEN L3 (ticker liveness read by a CLI command), WHERE T7 (Desktop/serve in-process
ticker) and T2/T3 (`hermes -p X` against the host multiplexer).

Tests: the salvaged serve-ticker test keeps the fresh->alive / stale->dead
invariant; one named-profile test proves the profile is named on every verdict and
that probe's own heartbeat counts when the default record excludes probe. Both red
on origin/main (`_builtin_gateway_liveness()` False with a fresh heartbeat;
"No scheduled jobs." without the profile). Existing assertions on the old host-wide
string were moved to the new one.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-28 05:28:05 -07:00
Bartok9
fc15ec8fe3 fix(cron): treat a fresh serve ticker as scheduler-alive (#121881)
Desktop topology ticks cron in-process inside `hermes serve`, so
`hermes cron status`/`list` no longer report a missing CLI gateway
when the ticker heartbeat is fresh.

(cherry picked from commit 583de75430adf6272eca932bf72665148df52101)
2026-09-28 05:28:05 -07:00
teknium1
01d85137bd fix(review): MCP owner rebuild keeps the launch profile's env-only credentials (#119092 review)
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.

Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
2026-09-28 05:24:26 -07:00
kokhlo
0543bfd6fd fix(tui_gateway): tools.list/tools.show read back the session's effective toolsets (#117977, salvage #118020)
`profiles.configure` pins `platform_toolsets.cli` in the profile's own config.yaml, the key
`_load_enabled_toolsets` reads under the session's profile scope when the agent is built. The
read-back RPCs consulted only a BUILT agent: a session whose agent had not been built yet
(`session.create` + `profiles.configure` + `tools.list {session_id}`, the editor's flow) read
`getattr(None, "enabled_toolsets", [])` and reported every toolset enabled, and `tools.show`
listed the gateway-global catalog. A third-party WS client could not read back what the chat
may use.

`_session_toolsets(session)` is the one seam both RPCs answer from: the live agent's
`(enabled, disabled)` when built, else `_load_enabled_toolsets(platform)` /
`_load_disabled_toolsets()` under `_session_profile_runtime_scope(session)` — the same read and
scope `_build_agent` and `_refresh_live_sessions` use, so a secondary-profile session served by
one multiplexing backend resolves its own home, not the launch home. No new RPC, result shapes
unchanged (row L6, topology T7).

Salvaged from #118020 by @kokhlo (mechanism kept; scope widened from HERMES_HOME-only to the
full profile runtime scope, `disabled_toolsets` resolved on the same path, tests trimmed to two
invariants driving the real handlers against two homes A→B→A).

Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
2026-09-28 05:24:26 -07:00
teknium1
b74ff36401 chore(contributors): map me@spencermcguire.com -> angel12 (salvage #120734) 2026-09-28 05:24:26 -07:00
teknium1
e5947817be fix(tui_gateway): profiles.describe binds the profile scope via _session_profile_runtime_scope (#120726, salvage #120734)
Trim of the salvaged fix: `_hermes_home_scope` returns the existing
`_session_profile_runtime_scope` context manager (home + secret + terminal,
the composition `@_profile_scoped` binds) instead of re-implementing the
token/release pair around `_profile_runtime_scope_tokens`. Behaviour is
unchanged: the launch home maps to None (frozen launch env), no external
source hydration for these config read/write bodies.

Class: every `_hermes_home_scope` caller in methods_profiles.py (describe,
configure cfg sections, create-time voice/model mirroring) — row L6,
topology T7. Red on base: describe answered `set()` enabled toolsets under
multi-profile hosting; green on head.

Co-authored-by: angel12 <angel12@users.noreply.github.com>
2026-09-28 05:24:26 -07:00
ang3l12
49064ce4dd fix(profiles): describe/configure bind the profile's secret scope, not just its home
`_hermes_home_scope` set only the HERMES_HOME override. Under multi-profile
hosting the toolset snapshot's XAI_API_KEY probe raised UnscopedSecretError,
which `_describe_toolsets`' best-effort `_try` turned into an empty enabled
set: every unpinned profile described as "all toolsets off", and a Desktop
editor save from that snapshot pinned whatever subset the user re-checked.

Bind the full runtime scope via `_profile_runtime_scope_tokens` (the same
home + secret + terminal binding `@_profile_scoped` uses), launch profile
included, without external-source hydration since these bodies only read
and write config.

Fixes #120726

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 05:24:26 -07:00
teknium1
c8756dbd61 fix(mcp): remote headers render under the owner's fresh secret scope at connect and every reconnect (#119092, salvage #119097)
Under multiplex a served secondary profile's remote MCP server whose config has a
`${VAR}` Authorization header was brought up at gateway boot with the placeholder
unresolved (HTTP 401) and retried that same rendering forever.

Mechanism: `_owner_scope_home()` trusted any caller-bound secret scope, and the
gateway's boot-time `_profile_runtime_scope` binding is a SNAPSHOT taken before the
profile's external secret source (`secrets.command`) may have answered. The run
task copies that context, and `_refresh_remote_config` re-read config.yaml on every
probe but rendered it under the same frozen mapping, so a parked server sent the
literal `Bearer ${VAR}` every ~5 min for the life of the process.

- `_owner_scope_home()` now returns the owner home (the scope's stamped home, else
  the registry scope) even when a scope is bound: MCP rebuilds the owner's scope
  fresh, which retries hydration (cached once it succeeds). Single-profile
  processes (scope key None) are unchanged.
- `MCPServerTask.run()` binds that fresh owner scope around the rebuild-time config
  refresh, so a reconnect heals once the source hydrates.
- `_require_rendered_remote` fails closed: a remote `url`/`headers` still carrying
  a `${VAR}` after rendering raises naming the variable instead of sending it.
- Salvaged from #119097 (@JoaoMarcos44): the connect-time re-render, which also
  covers a lazy server whose config was rendered at boot; its mechanism alone was a
  no-op for the reported topology because the owner-scope install was skipped
  whenever a (frozen) scope was already bound.

Row L1/L3, topology T2. Tests: A→B→A under set_multiplex_active(True) with two
temp homes, red on origin/main (literal placeholder sent), green on head.

Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
2026-09-28 05:24:26 -07:00
JoaoMarcos44
c418294eca fix(mcp): reinterpolate config after owner secret hydration 2026-09-28 05:24:26 -07:00
liuhao1024
079ff12f96 fix(buzz): scoped requirements gate resolves the profile's buzz section like the runtime loader (#125985, salvage #126060)
Inside a multiplexed secondary profile scope, `_profile_buzz_extra` read the buzz block only from
`gateway.platforms.buzz`, while the runtime loader (`gateway/config_loader.py::platform_section`,
the same seam that feeds this plugin's `apply_yaml_config_fn`) also resolves a top-level `buzz:`
block and the documented `platforms.buzz` shape. A fully configured secondary profile therefore
failed `check_requirements` closed on every boot ("Platform 'Buzz' requirements not met").

Resolve the section through `platform_section` so the gate sees exactly what the loader sees.
Row L1 (profile-scoped gate reads config), topology T2 (multiplexed secondary profile).

Salvaged from #126060 by @liuhao1024 (authorship kept); the test is trimmed to one A->B->A
invariant across two profile homes. Supersedes #126067 (same mechanism, later).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-28 05:24:26 -07:00
teknium1
e785ced297 test(gateway): trim plugin-injector salvage to two invariants (#125102, salvage #125114)
Keep the pre-existing newer-owner invariant in the gateway test (only the patch seam moves to the
process-wide slot the runner now publishes through) instead of a mock-based change-detector, drop
the hermes_cli duplicate of that invariant, and bind set_multiplex_active(True) around the
A->B->late->A profile-manager test so it exercises the multiplex topology the bug lives in (T2, L1/L3).
2026-09-28 05:24:26 -07:00
JoaoMarcos44
8a462159c1 fix(gateway): publish plugin injector across profiles 2026-09-28 05:24:26 -07:00
teknium1
4bd1f96c90 chore(contributors): map JE4NVRG email for the #122128 salvage 2026-09-28 05:24:26 -07:00
teknium1
75b6d946cd fix(a2a): scope the client-tools A2A_PORT gate and prove A→B→A under multiplex (#122126, salvage #122128)
Follow-up trim to @JE4NVRG's salvaged commit (is_connected → env_is_connected("A2A_PORT")):

- plugins/platforms/a2a/tools.py::_a2a_tools_available read os.getenv("A2A_PORT") too — the
  same unscoped gate in the same plugin, so every secondary profile also paid for the a2a
  toolset whenever the launch profile set A2A_PORT. Now get_scoped_secret, same policy as
  the adapter's port read (scoped miss fails closed; default/T1 reads its own environ).
- Tests trimmed to two invariants: A→B→A over two temp profile homes under
  set_multiplex_active(True) with the launch value in os.environ (profile B with no
  A2A_PORT of its own is NOT connected and gets no tools; red on origin/main at index 1),
  and the standalone T1 case where os.environ IS the profile.

Row L1 (boot-time platform enablement probe), topologies T2/T3 (multiplexed host serving
secondary profiles). No other plugins/platforms/*/__init__.py is_connected/check_requirements
does an unscoped env read (buzz #125985 is handled separately).

Co-authored-by: JE4NVRG <jean.v1803@gmail.com>
2026-09-28 05:24:26 -07:00
JE4NVRG
68c205b5b9 fix(a2a): resolve the inbound enablement gate through the profile-scoped reader
is_connected() gated on a raw os.getenv("A2A_PORT"), which under multiplexing
holds the launch (default) profile's value. Every secondary profile therefore
instantiated the inbound A2A server, all of them fell back to the module
default port 9900, and all but one died with "could not bind ... Address
already in use" (fatal / bind_failed in gateway_state.json) on every boot.

Resolve the gate through gateway.platforms._shared.env_is_connected, the same
scope-aware idiom the sibling adapters use -- the adapter's own port read
already goes through _get_scoped_secret for exactly this reason.

Tests: tests/plugins/test_a2a_plugin.py pins the behavior. A profile with no
A2A_PORT in its own scope is not connected and never reads os.environ (a raw
os.getenv fails the test), a scoped port connects, and extra.enabled still
short-circuits.

Fixes #122126
2026-09-28 05:24:26 -07:00
hermes-seaeye[bot]
365259378d fmt(js): npm run fix on merge (#126363)
Some checks failed
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
Skills Index Freshness Check / check-freshness (push) Has been cancelled
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-28 12:14:07 +00:00
teknium1
ad4e349df7 fix(review): supersede displaced PTY viewer, route project tile by active source
Review findings on #126144 (independent review, all minor):

- pty_session.close_other_sessions closed a sibling tab's PTY under it without
  closing its websocket: two tabs sharing one attach token on different
  profiles left tab A silent until its next keystroke returned 1013. Close the
  displaced viewer with the documented 4409 before closing the session. Still
  closes attached siblings (skipping them would reopen the #125287 lease hole
  via the old handler's finally: detach).
- openNewSessionTile paired projectProfile() with $newChatConnectionId — the
  source captured with a stale new-chat pin — so a pin made on another
  connection produced the right profile on the wrong host. Build the route
  from the active source (resolveActiveSourceOwnerRoute) for the
  project-derived profile.
- windows.test.ts: restore the no-parentSessionId case asserting the
  viewed-profile fallback alongside the parent-routed case.
2026-09-28 04:59:38 -07:00
teknium1
7bb4e32c56 chore: map contributor email for salvaged #121771 commits 2026-09-28 04:59:38 -07:00
teknium1
df05d0a8e2 fix(desktop): explain the stale-open gateway reconnect guard (#121680, salvage #121771)
Trim-follow-up to the salvaged PR: document why ensureGatewayOpen asks the
socket itself instead of trusting the effect-mirrored gatewayStateRef, and
why the registry backs a not-yet-populated ref. No behaviour change.
2026-09-28 04:59:38 -07:00
Joel
c39ecff00f test(desktop): cover stale-open closed socket recovery 2026-09-28 04:59:38 -07:00
Joel
150f28a8bb fix(desktop): verify gateway socket before skipping reconnect 2026-09-28 04:59:38 -07:00
Joel
341766ebb8 fix(desktop): verify gateway socket before skipping reconnect 2026-09-28 04:59:38 -07:00
Joel
4185428c2d test(desktop): cover stale gateway ref recovery 2026-09-28 04:59:38 -07:00
Joel
e56ec05fb8 fix(desktop): recover from stale gateway ref 2026-09-28 04:59:38 -07:00
teknium1
3361ac4a82 fix(desktop): trim #120899 salvage — owner route only, one window per chat (#120213, salvage #120899)
Keep the substantive fix from #120899 (renderer carries the session owner's
connectionId, buildSessionWindowUrl serializes it, boot honors it for
win=secondary windows, child watch windows route via their parent) and trim
to the salvage bar:

- windows.ts: drop the fail-closed assertSessionOwnerResolved rung; an owner
  that resolves nowhere keeps the documented active-profile fallback.
- main.ts: keep the session-window registry keyed by sessionId (one window per
  chat) instead of (connectionId, profile, sessionId).
- windows.test.ts: fold the connectionId invariant into the existing owner
  test instead of adding a second one.
2026-09-28 04:59:38 -07:00
itsflownium
13919e52a8 style: separate secondary window URL assertions 2026-09-28 04:59:38 -07:00
itsflownium
25bc086d3f fix: preserve connectionId for secondary chat windows (Fixes #120213) 2026-09-28 04:59:38 -07:00
teknium1
f8deae40fd fix(desktop): own a project-cwd tile by the project's profile (#124265, salvage #124296)
The #79005 pin (projectProfile + pinNewChatProfile) lives in
startWorkspaceSession, which the sidebar project "+" only reaches while no
chat is loaded. With a chat occupied, startSessionInWorkspace stacks a tile
via openNewSessionTile({ cwd, listed: false }) and the owner fell through to
$newChatProfile || $activeGatewayProfile — a stale new-chat pin routed
session.create to the wrong (default) profile with the project's cwd.

Fix it once at the tile create: a tile anchored at a project path defaults
its profile to projectProfile() unless the caller named one. That also
covers the other project-cwd tile paths (project-row "+" drags via
onNewSessionSplit, the "new project" drop via $newProjectSessionRequest).
All-profiles view has no owner and keeps the old fallback.

Direction credit: @kokhlo (#124296) identified the skipped pin; this lands
the profile on the tile call instead of adding a DI'd door module.
2026-09-28 04:59:38 -07:00
teknium1
252e316c85 test(dashboard): profile switch closes the previous profile's keep-alive PTY (#125287, salvage #125486)
End-to-end receipt through /api/pty: one attach token, profile alpha then beta;
the alpha PTY (which held the chat's active-session lease) must be closed when the
tab moves to beta instead of idling detached for the 30-minute registry TTL.
2026-09-28 04:59:38 -07:00
Nagisa-3000
ad06769eb1 fix(dashboard): release stale PTY sessions on profile switch 2026-09-28 04:59:38 -07:00
teknium1
e55c1cb879 fix(desktop): pin /messages/around reads to the session owner and trim the pin seam (#125372, salvage #125414)
Follow-up trim of the salvaged #125414 commits:
- fetchHistoryWindow (GET /api/sessions/{id}/messages/around) is the same
  session-scoped read class and now carries sessionReadOwnerPin too.
- client.ts: drop the unused mode field and the sessionOwnerForRead wrapper,
  shorten the seam comment.
- sessions.test.ts: keep two invariant tests (owner pin applied; explicit
  caller pin authoritative).
- use-timeline-history.test.tsx: stub sessionReadOwnerPin in its client mock
  so the hook test compiles against the new import.
2026-09-28 04:59:38 -07:00
Yuan Li
36d59a2ca3 test(desktop): stub the session-owner pin in the timeline-index client mock 2026-09-28 04:59:38 -07:00
Yuan Li
af9d7a6318 fix(desktop): pin session REST reads to the session's owner connection
A session-scoped read (detail / messages / timeline) dispatched on the
window's ambient connection scope: with a registered remote that exposes
a same-named profile, both directions answered 404 "Session not found"
for rows that verifiably exist on the other machine (#125372).

The renderer already persists the owner (tile route -> owner hint ->
connection-tagged row, resolved by knownOwnerForSession); the read path
never consulted it. Wire that resolver into api/ through a setter seam
(same pattern as setApiRequestProfile, no store->api cycle) and pin the
resolved connection on the three read helpers. An explicit (connection,
profile) caller scope stays authoritative; a bare-profile scope and the
ambient dial gain the owner's connection pin; reads with no resolvable
owner keep today's ambient behavior.
2026-09-28 04:59:38 -07:00
teknium1
dd3ba4b4b8 fix(pm): close the lease create→flock race on the writer side, no grace window
A time-based grace kept a just-abandoned lease 'held', so back-to-back `hermes pm repair` runs
never collected the generation the previous run execv'd out of (e2e test_generation_gc red).
Rename+replace is out (msvcrt byte locks refuse renames on Windows), so the writer now re-checks
its lease path after taking the flock and retakes when a peer's prune unlinked it in between; the
pruner only unlinks while holding the lock, so a visible path after the lock is always ours.
2026-09-28 04:44:29 -07:00
teknium1
7627eb466f fix(review): all-external fleet proves nothing for a SHA-less marker; holder argv strips both quote kinds
Review minors on #126179:

(a) update_cmd_fleet._marker_only_restart_obsolete: a fleet whose every
    row is EXTERNAL (served from another checkout root) fell through the
    row loop and cleared the marker with zero evidence about this checkout.
    For the SHA-less, inventory-less record it is now treated like an empty
    fleet and stays armed.

(b) update_cmd_windows._hermes_holder_subcommand stripped only double
    quotes before the inline-bootstrap check, so a shlex.join'ed
    runtime_command (`python -I -c '<bootstrap>' dashboard`, the shape
    launchd ProgramArguments take through main_dashboard) was never
    inventoried. Strip both quote kinds, as _is_entry already does.
2026-09-28 04:44:29 -07:00
teknium1
6da0d5a887 fix(review): fail closed on unopenable or just-created runtime leases
Review majors on #126179 (runtime_state._prune_unlocked_leases):

1. The boot-time prune opened every sibling lease with O_RDWR and caught
   only FileNotFoundError, so one lease this user cannot open (a root-owned
   0o600 file from `sudo hermes` on the same checkout, the #125525 class)
   raised PermissionError inside activate_dependencies and hermes_bootstrap
   exited 1 with "run hermes pm repair" on every launch. Any other OSError
   now counts the lease as HELD: GC stays fail-closed, boot never aborts.

2. lease_directory creates (O_CREAT|O_EXCL) and flocks in two syscalls, and
   pm/worker.py, pm/launch.py and pm/environments.py call it outside
   runtime_lock, so a peer's prune in that window could take the
   non-blocking lock and unlink the file; the first process then held a
   lock on an unlinked inode and its pin was invisible for its lifetime
   (collect_generations could rmtree a running generation). A lease younger
   than LEASE_GRACE_SECONDS (60 s) is never a prune candidate and counts as
   held. Chosen over a temp-name + os.replace scheme because it is three
   lines, keeps a crash-between-create-and-flock leak collectable, and a
   60 s margin dwarfs a microsecond window.

test_next_reader_removes_lease_left_by_hard_exit ages the hard-exited
lease past the grace window before asserting the next reader removes it;
its invariant is unchanged.
2026-09-28 04:44:29 -07:00
teknium1
05c5cbf190 test(update): follow the SHA-less obligation rule in the reconciliation matrix and completion harness
Two existing tests pinned the pre-#125952 shape and went red on the salvage:

- tests/hermes_cli/test_update_scoped_reconciliation.py `marker-no-sha`: expected an
  inventory-less marker armed with expected_sha="" to stay pending with every live gateway
  current on the checkout. That was the "no expected_sha — nothing to verify against"
  fail-closed default from 8130274b37, which #125952 deliberately changes: with no owed set
  and no SHA the checkout is the only code the record can be held to, so the case now settles
  (False) exactly like `marker-external-restart`. Added `marker-no-sha-stale` so the boundary
  stays pinned: the same record keeps the warning while a live row is stale.

- tests/hermes_cli/test_update_completion_process.py
  `test_interrupt_after_child_success_demotes_gateway_marker_at_boundary`: the harness hooked
  the FIRST subprocess.Popen as the completion child. The armer now falls back to
  _current_checkout_sha() when the request carries no SHA, which spawns `git rev-parse`
  before the completion child, so the hook wrapped git's Popen and the recorded waits were
  subprocess.run's own (`[2.99, 2.99, None]`). The hook now matches the completion child by
  its argv; the invariant under test (interrupt after published success kills, waits with a
  bounded timeout and closes the child's stdout) is unchanged.

No product change.
2026-09-28 04:44:29 -07:00
teknium1
a767190e84 fix(desktop): trust a remote HERMES_HOME only when it exists on the remote (#118988, salvage #119022)
probeWindowsRemote took `$env:HERMES_HOME` at face value. A stale
User-scope value on the remote (install.ps1 persisted one before
47f4ab3a17 stopped doing so; nothing removes it) or a client path
leaked over the SSH channel then flowed straight into
assertSafeRemoteHome and every Windows remote failed with
"Unsafe remote Hermes home" even though the remote install was fine.

Only honour HERMES_HOME when it names an existing directory on the
remote; otherwise fall back to the remote's default
%LOCALAPPDATA%\hermes, exactly as an unset variable does. The
installer/updater half of the report is already fixed on main:
scripts/install.ps1 no longer calls
[Environment]::SetEnvironmentVariable("HERMES_HOME", ..., "User")
(47f4ab3a17), and the updater's "Refreshed Windows gateway launcher
scripts" step only rewrites gateway.cmd/.vbs with a process-scoped
`set HERMES_HOME=`; it never persisted anything.

Slim redo of #119022 by @MohamadKanso, which reached for the User/Machine
registry scopes instead of a remote existence check.
2026-09-28 04:44:29 -07:00
chelsealong
42941c0a37 fix(pm): put the checkout launcher ahead of the venv's own hermes script (#124627, salvage #124828)
The runtime venv installs the core project editable against the PM
build snapshot (workspace/), so venv/bin/hermes always imports that
frozen copy. PM only rebuilds the venv on lock/extras/python-pin/plugin
changes, so a source-only update leaves the copy stale while the
gateway keeps running the checkout via activate_dependencies's
sys.path override. Any child that resolves `hermes` off PATH (terminal
tool, cron, kanban workers) still finds venv/bin first and runs the
stale snapshot instead — including `gateway install`/`restart`, which
then pins launchd/systemd to the snapshot's launcher path.

Prepend the checkout's own .hermes/bin ahead of venv/bin in
activate_dependencies so `hermes`/`hermes-acp` always resolve to the
launcher that puts the checkout first, regardless of what else lives
in the venv. This also fixes tools/environments/local.py's
_resolve_hermes_bin_dir(), whose shutil.which("hermes") now finds the
launcher dir instead of the venv's console script, so no separate guard
on PROJECT_ROOT in generate_launchd_plist / systemd unit generation is
needed: the snapshot's hermes_cli is never the running one.

Fixes #124627
2026-09-28 04:44:29 -07:00
Ahmett101
b23d4f4b32 fix(pm): collect abandoned runtime leases (#125609, salvage #125632)
lease_directory releases its per-invocation lease file only through atexit,
which os.execv (hermes_bootstrap's re-exec into the managed environment) and
os._exit (gateway shutdown_watchdog._hard_exit, hermes_cli/main.py past
finalization) never run: a 4s supervisor restart loop left ~860 lease files an
hour in the SELECTED generation, which generation GC never visits.

The kernel lock dies with the process even when atexit does not, so the file
itself is the only leak: every reader (the next lease_directory and
leases_held) now unlinks lease files nobody holds a lock on. Both run under
runtime_lock at boot / in the collector, so a lease between O_CREAT|O_EXCL and
its lock is never pruned from under a live process.
2026-09-28 04:44:29 -07:00
teknium1
96475946c1 fix(update): trim default-config migration salvage to two tests (#105013, salvage #105014)
Follow-up trim of the #105014 salvage: keep the named-active-migrates-default and
default-active-still-migrates-named-siblings tests; the missing-config.yaml skip is
unchanged pre-existing code.
2026-09-28 04:44:29 -07:00
Adolanium
9c02d6868e fix(update): migrate default config.yaml on named-profile update
Sibling config migration only walked profiles/. Default lives at the
install root, so hermes -p work update left it on the old
_config_version. Snapshots already treat default as a sibling.
Reuse _sibling_profile_homes so the two lists match.
2026-09-28 04:44:29 -07:00
teknium1
f46c7159e4 fix(update): trim SHA-less obligation salvage to the checkout fallback (#125952, salvage #125964)
Follow-up trim of the #125964 salvage:

- Armer falls back to _current_checkout_sha() (the identity the reader compares the fleet
  against) instead of a new receipt helper; at arm time the receipt has no post_update yet.
- Reader: the SHA-less, inventory-less record follows the existing inventory-less rule
  (live fleet all current on the checkout) with no extra pid-liveness gate, matching the
  invariant the SHA-bearing inventory-less record already has. The gatewayless branch stays
  fail-closed for an SHA-less record (checkout_contains("") is not evidence).
- Drop the obligation_fields pid plumbing and update_receipt helper; keep 2 tests:
  discharge on live-fleet evidence, and the armer recording the checkout SHA.
2026-09-28 04:44:29 -07:00
kokhlo
e57f6d0f2d fix(update): discharge an SHA-less host restart obligation on live-fleet evidence
A no-op `hermes update` (pre == post SHA) whose head capture failed armed the
host update-restart record with `expected_sha: ""` and no inventory. The reader
bailed out on the empty SHA before the live-fleet reconciliation, so the
"gateways may still be serving pre-update modules" warning could never clear
even when every live gateway was current on the checkout.

Reader (hermes_cli/update_cmd_fleet.py):
- An inventoried obligation without its SHA stays fail-closed.
- An inventory-less, SHA-less record now falls through to the existing
  live-fleet check once the process that armed it is gone (pid liveness via
  pid_exists_stdlib), so a status call cannot discharge a restart phase that
  is still running.

Armer (hermes_cli/update_cmd.py):
- When the completion request carries no SHA, fall back to the receipt's
  post_update.sha (then pre_update.sha) instead of arming with "".

obligation_fields() now reports the arming pid alongside expected_sha.

Tests: discharge on all-current fleet, kept when the armer pid is alive, kept
on a stale fleet, kept with an inventory but no SHA; the discharge test fails
on the pre-fix reader with the issue's exact warning.
2026-09-28 04:44:29 -07:00
teknium1
7bf1373cb8 fix(cli): select dashboard/serve backends by entrypoint + subcommand tokens, not argv substrings (#121156, salvage #121243)
`hermes update`'s post-update dashboard cleanup and `hermes dashboard --stop`
picked their kill list with `any(p in cmd for p in _DASHBOARD_PATTERNS)`,
i.e. "hermes serve" anywhere in the command line. "hermes serve" is a prefix
of "hermes server", so a terminal multiplexer started as
`herdr --session hermes server` was SIGTERMed by the updater, and (before
#124940) its unrelated systemd unit was restarted with every shell in it.
`_parse_dashboard_runtime` (launchd backend inventory, `--status`,
uninstall) used the same substrings.

Both now defer to the canonical token matcher
`update_cmd_windows._hermes_holder_subcommand`: a Hermes entry token
(`hermes`, `-m hermes_cli.main`, `.../hermes_cli/main.py`) followed, after
value flags, by exactly the `dashboard`/`serve` token. The matcher learns
the `hermes_cli/main.py` script entry (the one shape the substring table
covered that it did not) and strips single quotes, since launchd
`ProgramArguments` reach it `shlex.join`ed. `_DASHBOARD_PATTERNS` is
deleted; the spawn-ledger augmentation is unchanged.

The home-ownership fallback (`_hermes_home_for_pid` -> $HOME/.hermes) is
left as is: it is only unsafe when the selection has false positives, and
a process that passes the token matcher IS a Hermes backend whose default
home resolves that way. Point 3 of the issue (foreign unit restart) landed
in #124940.

Slim redo of #121243 by @MohamadKanso (same core hunks; that PR also
rewrites the entry detection positionally and reshapes the Windows holder
classifiers, beyond what this issue needs); #121781 by @salch-cred is a
later duplicate of the scan hunk mixed with an unrelated auth change.

Red on origin/main (2 failed) / green on head (2 passed):
tests/hermes_cli/test_process_identity_canonical_matchers.py -k "decoy or never_substrings"
2026-09-28 04:44:29 -07:00
teknium1
c935aa328f fix(tui_gateway): bind the session's home for the cwd record only when the caller is not already in it
_register_session_cwd runs from prompt_turn inside the turn's profile scope as well as from the plain
session.create @method; binding unconditionally added a redundant override/reset pair to every turn
(test_prompt_submit_releases_old_history_before_heap_trim red). Bind only when the routed home differs
from the session's profile_home, via hermes_constants directly rather than the server-bound names.
2026-09-28 04:34:30 -07:00
teknium1
124b5c7d45 fix(review): key cwd/override records by the session's own profile home; keep /p/<launch>/ ids stable
Independent-review minors on #126157 (#123989 class):

- session.create is a plain @method, so tui_gateway/session_workdir.py::_register_session_cwd
  wrote the cwd record under the RAW session key while the scoped turn read
  profile:<p>:<key> and missed it until the first `cd`. The writer now binds the
  session's own profile_home around register_task_env_overrides.
- tools/terminal_tool.py::_task_env_overrides stayed raw-keyed, so two profiles
  registering the same task id (docker_image / cwd) still collided. Writer, clear
  and both readers (_has_isolation_overrides, resolve_task_overrides) now use
  _qualify_task_key; the isolation-override branch of _resolve_container_task_id
  returns the qualified key so the env cache and the override record agree.
- gateway/platforms/api_server.py::_derive_chat_session_id prefixed the seed for
  the NAMED LAUNCH profile addressed via /p/<launch>/, so prefixed and
  un-prefixed requests on one profile derived different ids. The prefix is now
  skipped when the routed name is the process's launch profile
  (_names_launch_profile vs get_routing_process_hermes_home()).

Tests (red on 3720519199c, green here):
  tests/tui_gateway/test_session_cwd_profile_key.py::test_secondary_profile_session_cwd_is_found_inside_its_scope
  tests/gateway/test_api_server.py::TestDeriveChatSessionId::test_launch_profile_prefix_keeps_the_unprefixed_id
2026-09-28 04:34:30 -07:00