Commit Graph

40693 Commits

Author SHA1 Message Date
Austin Pickett
bd945ec384 fix(state): surface corrupt state.db as one degraded session-storage state (#120274)
* fix(state): publish structural state.db corruption as one profile-level state

A structurally corrupt state.db showed up differently on every surface: the
sidebar endpoint returned 200 with empty slices plus an errors row, /api/sessions
returned 500, /api/status said components.storage ok and readiness was green.
None of them said the store was damaged, so Desktop rendered it as deleted
history (#72046).

hermes_state_health is now the single latch, keyed by resolved state.db path:

- SessionDB._halt_db_corrupt, the SessionDB read helpers, the web profile
  reader and the readiness probe publish into it, only for structural
  corruption (not FTS-scoped damage, not the malformed-schema case the web
  open path heals).
- gateway.readiness reports it (state_db degraded/corrupt, and session_store
  unavailable/corrupt even when the handle cache says ok), which also feeds
  /api/status components.storage (now with reason: corrupt).
- /api/sessions, /api/profiles/sessions and /api/profiles/sessions/sidebar
  carry storage: {profile: "corrupt"}; /api/sessions returns 503
  state_db_corrupt instead of 500.
- A peer SessionDB handle in the same process refuses writes on a latched
  path with the existing StateDbCorruptError, so gateway/agent transcript
  diversion and classify_persistence_error keep working unchanged.

The latch never clears on its own and resets on restart, the recovery boundary
StateDbCorruptError already documents.

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>

* fix(desktop): say the session store is damaged instead of an empty sidebar

The sidebar reads the list endpoints' new storage map into
$corruptSessionStores and renders a persistent destructive Alert above the
session list naming the affected profile(s). The copy says missing chats were
not deleted and points at the non-destructive path (quit Hermes, then
`hermes sessions recover --source <state.db> --inspect-only` or restore a
snapshot) plus the recovery guide; it does not recommend `sessions repair`
for structural damage.

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>

---------

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
2026-09-23 10:17:46 -04:00
Xipong
5c6b56a866 fix(desktop): keep offset-drifted backfill rows in order (#119606)
Co-authored-by: Xipong <217837358+Xipong@users.noreply.github.com>
2026-09-23 14:12:13 +00:00
teknium1
746f7ea21b fix(dashboard): Desktop children publish their own host role; SSH spawn shape is Desktop-owned
Review follow-up on the host-rendezvous isolation:

* Excluding Desktop-owned children from ROLE_SERVE broke #119644's terminal
  path: `hermes plugins install` found "the running Desktop backend" only via
  read_record(ROLE_SERVE), so on a Desktop-only box open chats no longer lit
  up. A Desktop child now publishes ROLE_DESKTOP_SERVE (own lock, record and
  0600 token). The attach/refuse ladder (_host_backend_attachment) still reads
  ROLE_SERVE only, so a supervised public dashboard is never blocked by it;
  notify_serve_backend prefers the host owner and falls back to the Desktop
  record. /api/host/identity reports the role actually published.

* is_desktop_owned_backend() missed Desktop's SSH spawn — `env HERMES_DESKTOP=1
  hermes serve --isolated --ssh-session-token-file F` carries NO token env var
  (remote-lifecycle tests assert the var name never appears) — so the SSH child
  still claimed ROLE_SERVE on the remote host and its MCP discovery flipped to
  deferred. The predicate now accepts the token-FILE argv shape; the argv half
  lives in _startup_fast.is_desktop_ssh_backend_argv (stdlib-only, importable
  before main.py's import wall) and replaces the two duplicate substring checks
  in main.py and dashboard_procs.py.

* web_server.py's remaining three bare HERMES_DESKTOP reads (cron ticker,
  managed-gateway teardown, orphan serve reap) route through the predicate;
  the tests that modelled a Desktop child with the bare flag set the spawn
  token, as a real pool child does.
2026-09-23 07:09:45 -07:00
teknium1
9ca3b54e91 fix(dashboard): one desktop-owned-child predicate, exit 78 on host-owner refusal
Widen the two salvaged fixes to the whole class and make the refusal
supervisor-safe:

- `process_identity.is_desktop_owned_backend()` is the single discriminator
  (HERMES_DESKTOP=1 AND the per-spawn HERMES_DASHBOARD_SESSION_TOKEN). The
  attach bypass, the named-profile reroute, the env sanitizer, the MCP
  discovery timing and the host-rendezvous publish all key on it now; a shell
  that merely inherited the flag from Desktop is treated as a normal launch
  (#119210). The publish skip from #119832 keyed on the bare flag, which
  would have hidden a supervised service launched from a Desktop shell.
- `_attach_to_host_backend` refuses with exit 78 (EX_CONFIG) instead of 1.
  Exit 1 under `Restart=always` was an infinite restart loop with nothing
  listening on the ingress port (2035 restarts, #119824); 78 is the
  deliberate-refusal code `RestartPreventExitStatus=78` parks on, the same
  contract the gateway unit already uses.
- Docs: the systemd example carries RestartPreventExitStatus=78 and the
  user-side remedy (stop the owner, or `--isolated`).
- Tests trimmed to ≤2 invariants per fix; the control case in the desktop
  tests sets the token so it models a real pool child.
2026-09-23 07:09:45 -07:00
KoNit-K
269a982b64 fix(dashboard): require Desktop spawn credential for host bypass 2026-09-23 07:09:45 -07:00
KoNit-K
95d1faf291 fix(web): isolate desktop host rendezvous 2026-09-23 07:09:45 -07:00
Austin Pickett
8fb0fc6ae6 fix(desktop): a submit that resumes the selected session moves the chat to the resumed runtime (#120273)
When submit cannot prove the pane's runtime owns the selected stored session
(reverse binding lost to eviction, reconnect, or a compression rotation the
selection never followed), it resumes the stored session and continues on
the runtime id that session.resume returns. That path pinned only
activeSessionIdRef. ChatView renders the $sessionStates slice named by
$activeSessionId, so the optimistic prompt, the reply, and every later turn
landed in a slice the pane never painted, while the legacy $messages mirror
(which the ref drives) looked correct. The chat stayed frozen until a relaunch.

Rebind the atom together with the ref, as the session-not-found recovery in
the same file already does, and carry the transcript the pane was showing
into the resumed runtime's empty slice (session.resume omits messages) when
both name the same conversation by lineage. A pane runtime that belongs to a
different stored session is never carried over.

Refs #71733
Refs #117867
2026-09-23 10:06:51 -04:00
John Paul Soliva
3591bf3bb8 fix(compression): turn-start in-place compaction no longer hides the newest summarized turn (#120187)
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.

tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.
2026-09-23 10:05:30 -04:00
Austin Pickett
a1838ea87a fix(telegram): keep update receipts across adapter rebuilds and restarts (#120257)
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.

Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.

Refs #68502

Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>
2026-09-23 10:05:17 -04:00
liuhao1024
5c4db8d8b0 test(desktop): pin the failed-turn reseat after a truncating retry
Regression tests from #118719: a failed turn the re-submitted prompt truncated
out of the stored history returns to its timeline position, and a failed tail
whose retry is still local-only stays at the end.
2026-09-23 10:05:09 -04:00
Austin Pickett
865e94672e fix(desktop): keep preserved error turns in place and drop stored copies
preserveLocalAssistantErrors appended every kept local error run after the
refreshed transcript, so an older failed turn repainted below newer turns
(#118002). Reseat each run after the refreshed row that preceded it locally,
and drop a run the refresh already stored under new ids (same role/text
sequence in the gap).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-23 10:05:09 -04:00
Austin Pickett
88e3c976ae fix(desktop): skip durably completed replies compaction re-inserted 2026-09-23 10:05:09 -04:00
teknium1
b9616fe1d1 test(sessions): trim the repair-profiles title-clash tests to two invariants
Six parametrized cases collapse to two: the clash on the batch's first row with a title at
MAX_TITLE_LENGTH (the suffix must still fit) and an import-stage failure whose lineage waits
for the next run. Assertions are unchanged.
2026-09-23 07:01:11 -07:00
teknium1
6c2a94f4b6 fix(tui-gateway): stale session_id is refused once at the config.set seam for every session-scoped key
Sweep of the #119983 shape: `reasoning` had the same stale-session -> global-write
fall-through as `fast`/`yolo` (a reaped session_id wrote agent.reasoning_effort into
config.yaml and answered success), and `model` already carried its own 4001. One check
at the dispatcher replaces the three per-handler guards: a non-empty session_id this
backend no longer holds, on a key whose sessionless branch writes a wider scope
(config.yaml agent.*, the process env every later child inherits), answers 4001 so
the client resumes. `scope="global"` is still honoured; a request with no session_id
at all is still the sessionless path.

Live (temp home, stale id): base wrote reasoning_effort: high and HERMES_YOLO_MODE=1;
head leaves config.yaml and os.environ byte-identical and answers 4001 for all three.
2026-09-23 07:01:11 -07:00
John Paul Soliva
3dd848f263 fix(sessions): repair-profiles moves a session whose title the target store already holds
`hermes sessions repair-profiles --apply` copied a stranded session into
its owning profile's store verbatim, title included. Titles are unique
per store only (idx_sessions_title_unique), so when the target profile
already held an unrelated session with the same title, which is common
for generic auto-titles, the insert raised IntegrityError.

_MoveBatch.run imported every row before deleting any and had no
per-row handling. The exception escaped before the delete phase and
before the result was memoised. Rows copied earlier in the batch were
left in both stores. Every later finding re-ran the whole batch and hit
the same error. Every row in the batch was reported failed on every
run, the command exited 1 forever, and a non-colliding row ended up
duplicated across two profiles.

- import_moved_session gives a colliding title the moved row's id tail,
  the same convention import_foreign_history uses, capped at
  MAX_TITLE_LENGTH. The resident row keeps its name, so resolving it by
  title is unchanged.
- _MoveBatch.run handles each row separately. A failed import or delete
  is that row's failure alone, reported through apply(). The rest of
  the batch still moves, and the batch runs once. The failed row's
  lineage waits with it: its descendants are not imported without it,
  and the parent it still points at in the source is not deleted. The
  next run moves the lineage whole.

Measured on main: with two stranded rows and one title collision, both
rows failed on every run and one was left in both stores. With the fix,
one run moves both, and a second run finds nothing.
2026-09-23 07:01:11 -07:00
John Paul Soliva
ba5bec3f2c fix(tui-gateway): a stale session_id no longer turns config.set fast/yolo global
config.set resolves its session with `_sessions.get(session_id)`, so a
runtime id the backend no longer holds (reaped, evicted, re-minted
after a reconnect) reaches the setters as session=None, the same value
a call with no session_id gets. Two setters then take their sessionless
branch and still answer success:

- fast writes agent.service_tier into config.yaml. A per-chat Fast
  toggle or a model-preset pick moves every session, the CLI, cron and
  gateway builds on the profile onto the priority tier.
- yolo (session scope) flips HERMES_YOLO_MODE in the backend's process
  env and reports value "1" with scope "session". The in-process
  approval gate froze that variable at import, so the chat keeps
  prompting while the Desktop lights the zap, and every compute-host
  child and terminal-spawned hermes started afterwards inherits the
  flip.

Both now return _sess_nowait's 4001 "session not found" when a
non-empty session_id misses, as _set_model and tools.configure already
do. The 4001 is also what the Desktop's stale-session recovery keys on,
so it now resumes the stored session instead of trusting a false
success. Calls with no or null session_id, live sessions, and yolo
scope=global behave as before. reasoning has the same fallthrough and
is left to #98901, which gives it a stored-row override instead.
2026-09-23 07:01:11 -07:00
John Paul Soliva
fdd49a372c fix(dashboard): a served profile is not a gateway of its own
Gateway liveness reports a profile the multiplexer serves as running on
the multiplexer's PID (446f6f79a8). The dashboard read that answer as
"this profile runs its own gateway" in two places. multiplexed_profile_refusal
returned None for every served profile, so while a multiplexer was live
(the only time it matters) start/stop were never refused and a restart
was never rewritten to the multiplexer. The /api/status topology listed
one phantom gateway per served profile beside the host.

_has_own_gateway() answers the actual question: a live gateway whose PID
is not the multiplexer serving the profile. The existing lifecycle test
stubbed _check_gateway_running to False, which is what hid it; it now
runs on the real liveness.
2026-09-23 07:01:11 -07:00
John Paul Soliva
35c0ea3dff fix(dashboard): restarting a served profile names the multiplexer even from the default home
`_gateway_subcommand` rewrites a restart of a profile the multiplexer
serves into a restart of the multiplexer. From a dashboard whose own home
is the default one it emitted a bare `gateway restart`, assuming bare
means the root. It does not: the child's profile pre-parse does not trust
a root HERMES_HOME as-is and re-reads the sticky active_profile (#22502),
so the restart went to that profile, which the multiplexer serves, and
the multiplexer was never restarted. The child now always gets
`-p default`, as the pooled-backend path already did.
2026-09-23 07:01:11 -07:00
teknium1
30056d612b fix(update): a carried local commit no longer strands the fleet-restart obligation
`_marker_only_restart_obsolete` bailed out with "a newer pull moved HEAD" whenever
`checkout_sha != expected_sha`, and otherwise held the fleet to `expected_sha` exactly.
A cherry-picked hotfix on top of the pulled SHA moves HEAD without any pull arming a
fresh obligation, so the recorded one could never discharge: every CLI start warned
and every no-op `hermes update` exited 1 through the armed veto in
`_pending_fleet_restart_needed`, while the gateway verifiably served HEAD.

The bail-out now fires only when HEAD no longer CONTAINS `expected_sha`
(`checkout_contains`, `git merge-base --is-ancestor`, fail-closed on any probe
failure), and the live fleet is held to the checkout SHA — the code it actually runs.
A gateway on genuinely stale code still keeps the obligation armed.

Closes #119367
2026-09-23 07:01:11 -07:00
teknium1
8d57bfc2d7 test(update): accept the self_restart_pending keyword in the drain-triage stubs
`_drain_or_signal_gateway_for_update` gained a keyword-only `self_restart_pending`
(#119597); the unit-budget test's `lambda *a: True` stand-in rejected it.
2026-09-23 06:51:47 -07:00
teknium1
5c96be1711 test(gateway): derive the live schtasks marker from the host ANSI code page
Follow-up to the #119847 salvage: a fixed CJK literal moves the failure from
cp936 hosts to cp1252 ones (schtasks stores /TR in the system ACP, so any
character the ACP cannot represent comes back as `?`). Pick the first of
ë / 方舟 / テスト / 한글 that round-trips through GetACP()'s code page and skip
only when none does — the #116193 guard keeps its non-ASCII coverage on every
Western and CJK lane.

Closes #119845
2026-09-23 06:51:47 -07:00
teknium1
39643567f9 test(gateway): linger tests use a real pathlib.Path redirected at tmp_path
Follow-up to the #119790 salvage: the stub still returned a SimpleNamespace for
the linger-file probe, so the next Path method the code grows breaks it again
(#119786 was exactly that, `.expanduser()` from `_bare_unit_pinned_home`). Now
`gateway.Path` stays pathlib.Path and only the `/var/lib/systemd/linger/<user>`
lookup is redirected at a real file under tmp_path. The root/system-unit branch
is exercised with a tmp_path home instead of a literal /tmp path.

Also: `test_venv_repair_path_refreshes_memory_provider` recreated a venv with a
real `uv venv` — FileNotFoundError on hosts without uv on PATH (#119786,
secondary). The spawn is captured instead, keeping the assertion that the repair
path recreates the missing venv.

Closes #119786
2026-09-23 06:51:47 -07:00
teknium1
7f9b95eb8d test(update): trim the salvaged finder/hand-off tests to two invariants per fix
#119525's suites carried five finder tests and four hand-off tests; the
invariants are "a stale or unparseable finder refuses the skip" and "a covered
map still skips" (finder), "both launches pass the checkout cwd" and "a stale
finder child imports the new package only from the checkout" (hand-off). The
remaining rows were the same properties in other clothing.
2026-09-23 06:51:47 -07:00
teknium1
8680a8d1ed fix(update): the import-health probe sees the venv's editable finder, not the checkout cwd
`_critical_module_import_failures` spawned `<venv python> -c ...` with
cwd=PROJECT_ROOT; `-c` puts the cwd at sys.path[0], so every first-party import
resolved from the tree and the probe was green while the installed editable
finder's MAPPING lacked `hermes_platform` — the exact state that crash-loops a
gateway started from `/` (#119466, maintainer triage gap 1).

When the venv carries an `__editable__*hermes_agent*finder.py` the probe now
runs under `-P` (Python >= 3.11, matches requires-python), so it imports through
the same finder the gateway will. A dev checkout with no editable install keeps
the cwd path and its advisory verdict.

Live: scratch venv, `uv pip install -e .` at 9863e315 (pre-hermes_platform),
checkout a27b1305: base probe NONE (green) while `cd / && python -c 'import
hermes_constants'` dies; head probe reports ModuleNotFoundError
hermes_platform; after a plain `uv pip install -e .` the probe is green again.

Part of #119466
2026-09-23 06:51:47 -07:00
teknium1
8895b0f119 fix(update): an in-gateway cron update no longer exits STALE for the gateway it runs inside
`hermes update --yes` from a cron job inside the gateway takes the #100179
ancestor branch (`_drain_or_signal_gateway_for_update` -> self-restart request,
fire-and-forget) — the gateway can only restart after the updater exits. The
post-restart fleet matrix then saw that same gateway on the pre-update code_sha,
printed STALE + "Update not complete" and exited 1 on every nightly run (#119597).

The restart phase now records the pids that accepted a self-restart
(`_GatewayRestartOutcome.self_restart_pending_pids`, threaded from the systemd,
manual and launchd paths) and `collect_fleet_versions(self_restart_pending=)`
turns exactly those rows into `restart_pending` — rendered as "restart pending
(deferred until this process exits)" and excluded from the stale_or_down
verdict. Every other stale/down gateway keeps its verdict; the same pid without
the recorded acceptance is still STALE. The receipt keeps the row's old sha, so
the next `hermes update` / CLI startup hint still verifies the restart landed
via the live fleet (`_receipt_reports_stale_runtime` -> `_live_fleet_covers_receipt`).

Closes #119597
2026-09-23 06:51:47 -07:00
KoNit-K
3a2d86603a test(gateway): use cp936-safe schtasks marker 2026-09-23 06:51:47 -07:00
KoNit-K
5a468f1b6e fix(tests): narrow gateway linger Path stub 2026-09-23 06:51:47 -07:00
joaomarcos
cfa455e598 test(update): prove the hand-off child imports past a stale finder
Fold the finder checks into the two skip invariants, and spawn a real
child from outside the checkout so both launch paths reach the new package.
2026-09-23 06:51:47 -07:00
joaomarcos
949a8bb0ab fix(update): start the post-swap child in the checkout
A stale editable finder cannot import a new top-level name unless the
checkout is the process cwd. python -m only adds that cwd, so a hand-off
spawned from elsewhere died before the map check could refresh the finder.
2026-09-23 06:51:47 -07:00
teknium1
7cf0afdf98 fix(update): drop the --reinstall plumbing; a plain editable install already rewrites the finder
Review fix for #119525. The stale-map detection stays; the forced
`--reinstall` goes, for three reasons verified live:

- uv (0.12.13, the version the updater ships) reinstalls an editable
  source tree on every `uv pip install -e .` (`~ hermes-agent==0.21.4`
  even with no change). Adding a root module and running the plain
  install rewrote the finder MAPPING without touching packaging files,
  so once the skip is refused the existing install already refreshes it.
- `--reinstall` reinstalls EVERY package in the venv, not just
  hermes-agent; the narrow spelling would be `--reinstall-package
  hermes-agent`, and neither is needed.
- `--reinstall` is not a pip option. On the pip fallback path (no uv:
  Termux without uv, some site-packages installs) a stale map turned the
  update into a hard `no such option` failure, worse than the skip.

Also removes the second `_editable_finder_mapping_current` call in
`_sync_python_dependencies_after_pull` (the predicate already ran it)
and the flag's test.
2026-09-23 06:51:47 -07:00
joaomarcos
60a088795d fix(update): refresh a stale editable finder before skipping reinstall
The packaging-file skip treated an unchanged pyproject as proof the
setuptools finder could still import the checkout. That finder freezes
top-level names at install time, so a later package such as
hermes_platform stays missing and the gateway crash-loops after a
successful update.
2026-09-23 06:51:47 -07:00
teknium1
b88b3de6e3 fix(profile-scope): stamp profile_home only for FOREIGN homes; cache the gate-prefix scan
The profile_home stamp exists to let serves_routed_profile() see a foreign-home scope without the
HERMES_HOME override. Stamping the LAUNCH home from the own-home binders (launch_profile_runtime_scope,
_worker_profile_scope's launch arm, model_switch's launch arm) added nothing for production but, under
a test conftest where server._hermes_home differs from HERMES_HOME, flipped the launch profile to
'routed' and charged every turn a routed tool/plugin discovery. Own-home binders now leave the stamp
None; the foreign-home arms keep it.

is_profile_gate_env resolved the platform prefix set (Platform enum + bundled plugin scan + registry)
once per env key; strip_profile_gate_env walks the whole child env, so a spawn paid the scan
hundreds of times. The static half is lru_cached, the dynamic registry union is taken once per strip.
2026-09-23 06:46:02 -07:00
teknium1
ebb9bc5699 fix(env-policy): a profile gate is a platform prefix plus a gate suffix, not a bare name shape
is_profile_gate_env matched any name containing _ALLOWED_ / _ALLOW_ALL_ / ... so an operator's own
DEMO_ALLOWED_SENDER (script data, not a Hermes authorization gate) was deleted from every routed
child env built with strip_launch_profile=True, breaking no_agent cron scripts under the host
gateway. The owner prefix now has to be a platform: built-in Platform values, bundled platform
plugins (directory names and manifest aliases), runtime-registered plugin adapters, plus GATEWAY_
(pairing) and QQ_ (qqbot). Every gate the adapters read still strips; operator variables survive.

Fixes #119539
2026-09-23 06:46:02 -07:00
teknium1
a39efc9466 fix(compute-host): hydrate a routed profile's external secret sources before building its scope
The Desktop/dashboard isolated turn process (tui_gateway/compute_host.py::_build_server_session)
installed set_secret_scope(build_profile_secret_scope(home)) with no preceding
hydrate_profile_secret_sources(home). That process never ran the launch dotenv path for the routed
profile, so a provider key that lives only in an enabled external source (1Password, Bitwarden,
secrets.command) was absent from the scope and the first routed agent build failed closed.

Same order as gateway/run.py::_load_profile_secret_scope and
tui_gateway/model_switch.py::_profile_runtime_scope_tokens. Two-home A->B->A test drives the real
builder with a secrets.command helper as the only holder of each profile's key.

Part of #119521
2026-09-23 06:46:02 -07:00
KoNit-K
28c238f945 fix(auth): scope Nous keepalive on multiplexed gateways 2026-09-23 06:46:02 -07:00
beardthelion
9cae47fd90 fix(mcp): scope late OAuth attempts by profile
_LATE_ATTEMPTS was keyed by session key alone while live.py keys the
open-operation table by (profile_key, session_key) for exactly this
collision: two multiplexed profiles can carry the same session key (the
api_server binds X-Hermes-Session-Key verbatim, and it is client-chosen).

A parked OAuth attempt could therefore be adopted by a different
profile's next turn: adopt_late_connections would poll it, register the
same-named server from that profile's config, and enable it - a grant
authorized under profile A materializing under profile B, or being
discarded so the owning profile never adopts it.

Key the table by (profile_key, session_key), matching live.py's
documented identity pairing. The detached no-card path never opens an
operation, so its profile_key is empty; fall back to the calling
thread's home, which close() runs under the turn's profile scope.
2026-09-23 06:46:02 -07:00
beardthelion
0cbe552888 fix(scope): failed profile-scoped secret reads must not borrow os.environ
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.

Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
2026-09-23 06:46:02 -07:00
beardthelion
6aaa1c480b fix(profile-scope): bind home and secret tokens inside the try that resets them
Several profile-scope binders called set_hermes_home_override and
set_secret_scope before the try/finally that releases them. A raise in
scope setup (a corrupt or removed profile home) propagated with the
foreign override still bound to the caller's context, silently re-homing
every later read — and in _reregister_orphaned_adopters it also skipped
every remaining adopter. The set calls now run inside the try with
None-guarded resets; the same shape is fixed in the routed-turn scope,
the cron external worker (which also leaked the multiplex flag), the
kanban worker scope, the MCP OAuth paths, the launch-profile policy,
and model_switch, which releases partially-bound scopes on raise.
2026-09-23 06:46:02 -07:00
beardthelion
2c3a75beaa fix(secret-scope): scoped misses fail closed under a foreign-home scope
get_secret returned os.environ on a scoped miss whenever multiplex was
off, but non-multiplex hosts serve foreign homes too (dashboard/desktop
backend, per-profile cron, MCP owner scopes, kanban spawn-env builds),
where os.environ is the launch profile's. Bound scopes now carry the
home they were built for; serves_routed_profile detects a foreign scope
even when the binder deliberately skips the HERMES_HOME override, and
the miss returns the caller's default. Every production binder stamps
its home; own-home scopes keep the deliberate env overlay.
2026-09-23 06:46:02 -07:00
brooklyn!
09fbc27a4a test(desktop): preserve currency whitespace in billing queries 2026-09-23 08:40:08 -05:00
brooklyn!
cc0388e236 test(desktop): assert locale-aware counts and timeline precision 2026-09-23 08:40:08 -05:00
teknium1
2332a6433a chore: map contributor email for attribution audit 2026-09-23 06:31:43 -07:00
KoNit-K
5b635064b8 fix(gateway): slash dispatch binds the routed profile's scope once for every handler
Under multiplex the inbound handler runs under the RECEIVING bot's profile
scope (authorization needs its .env), which is not the routed RUNTIME when a
bot serves another profile's chat (profile_routes / bot_profile), and the
handler's own scope is already gone by the time a `_track_background_task`
or busy-path dispatch fires. `/memory pending|approve|reject` and
`/skills ...` then read tools/write_approval's `get_hermes_home()/pending`,
`load_on_disk_store()` and the approval-gate config under the wrong home:
a served profile's bot answers "No pending memory writes" while its own
pending/memory/ holds records, and an approve applies to another profile's
memories/ (#119915).

Fix the class at the dispatcher seam instead of per handler: both slash
dispatch tables (`_hm_dispatch_canonical_command` idle path,
`_dispatch_busy_slash_command` busy path) enter
`_async_profile_scope_for_source(source)` — the async twin of the turn
scope `_profile_scope_for_source`, launch-profile scope when multiplexing
is off — around the handler call. Every table handler (memory, skills,
model, reasoning, compress, ...) now sees the runtime home; the per-handler
wraps in /compress, /profile and /model become redundant and can be removed
in a follow-up once their own tests are re-keyed.

Salvaged from #119922 (KoNit-K: bind the routed home around all /memory
and /skills review ops incl. the approval-gate check); the handler-local
wraps moved to the dispatcher. Supersedes #119932.

Co-authored-by: chugiacan <328579510+chugiacan@users.noreply.github.com>
2026-09-23 06:31:43 -07:00
c0d1ngHUB
857e651b5f fix(relay): the logical LLM close must not pop through a concurrent turn's scope
The logical LLM scope close in ``relay_llm._complete_logical`` popped its handle with
the unguarded ``pop_relay_scope``. Two consecutive calls in one session share a physical
scope stack, so when the sibling turn's live scope sat above the handle the native
binding raised::

    RuntimeError: invalid argument: scope handle is not at the top of the stack

``_complete_logical`` catches that, logs "logical LLM finalization failed" with a
traceback and returns early, so the early-return path also skips the
``turn.logical_llm_calls`` cleanup until the handle is retried. Observed in production as
one traceback per overlapping turn (~14/day on a busy local profile).

``pop_relay_scope_if_top`` already exists for exactly this case and is used by the
shared-metrics task close (PR #116685, #115471); the logical-LLM seam was missed. Guarding
``pop_relay_scope`` itself is NOT an option — ``_pop_with_drain`` relies on that raise to
detect a stacked sibling and drain it.

The skipped scope is reclaimed by the existing session-close drain
(``RelayRuntime._close_scope_handle``), so the handle still leaves
``logical_llm_calls`` and the session unwinds without an orphan.

Test: ``test_logical_close_skips_pop_under_concurrent_turn_scope`` pushes a sibling scope
in the session context, completes the logical call, and asserts the sibling (not ours) is
still on top and that the sibling's scope survives. It fails on upstream ``main`` with the
exact ``RuntimeError`` above and passes with the guard.

Fixes #115471 (the remaining call site).
2026-09-23 06:31:43 -07:00
teknium1
15a0687fd5 test(gateway): trim handoff-owned recovery to two invariants
Keep the incident (completed handoff + cli_close stays recoverable) and the
control (a failed handoff does not widen recovery); drop the four
change-detector variants and the try/except-pass fixture from #119875.
2026-09-23 06:31:43 -07:00
teknium1
408687dd49 fix(gateway): /resume clears the persisted /model pin it used to drop by accident
With switch_session now carrying model_override across the re-point, /resume
- a real conversation boundary (#10702) - relied on the drop as a side effect.
_clear_conversation_scope clears only in-memory state, so the next turn would
rehydrate the persisted pin from the routing store. Clear it explicitly beside
the funnel call.

Invariant tests: switch_session keeps the pin; reset_session (/new) still drops it.
2026-09-23 06:31:43 -07:00
KoNit-K
8e9b3e00d2 fix(gateway): switch_session keeps the route's persisted /model pin
SessionStore.switch_session re-points a session key at another session id
for non-boundary reasons too (async-delegation re-pin, compression-tip
binding heal, CLI handoff, /branch), but _replace_route_locked built the
new SessionEntry without model_override, so the user's /model pin silently
reset to None on every such re-point (#119864).

Scoped to switch_session's kwargs, not the shared helper: reset_session
(/new) uses the same helper and is a deliberate boundary (30e947e0a0) that
must keep dropping the persisted pin, or the next turn's
_rehydrate_session_model_override resurrects it.

Salvaged from #119868 (moved from _replace_route_locked into switch_session).
2026-09-23 06:31:43 -07:00
chuckie
61274a100b fix(gateway): recover a session a completed handoff owns
A completed /handoff transfers the source session to the gateway, but the
source interface's teardown can still stamp it with a terminal end_reason
(cli_close) before the destination's first reply. The stale-route self-heal
then finds a row that is neither recoverable nor a reset boundary, silently
drops the routing entry and mints a brand-new empty session — the handed-off
leg and its whole transcript are orphaned.

Field incident (QQ DM, one session_key, 2026-09-22):

  21:33  qqbot session B (253 messages) starts
  22:29  /handoff qqbot from the CLI; the gateway claims the row,
         handoff_state='completed', synthetic turn dispatched
  22:29  the CLI teardown stamps B end_reason='cli_close'
  22:54  the user's next DM takes the stale-route self-heal
         ("routing key ... is ended in state.db but still live in
         sessions.json; dropping stale entry ... (#54878)") and creates a
         new session C; the /resume-able leg B is lost

A row a completed handoff owns is now recoverable whatever end_reason it
carries: handoff_state='completed' means ownership was transferred, and the
stamp left behind by the departing interface is not a user-facing close. Both
peer-recovery queries (exact key and peer tuple) share the widened predicate.
The existing reset-boundary fence still applies, so a later deliberate
boundary drops the row, and pending/failed/NULL handoffs are untouched.

The CLI-side guard for the teardown path is the other half of this seam and
is tracked separately (#118196 / PR #118202). This change makes losing the
handed-off leg impossible even when such a stamp lands.

Tests: tests/gateway/test_handoff_owned_recovery.py (incident case plus
controls: plain cli_close stays closed, a reset after the handoff still
fences, a failed/pending handoff does not widen recovery, and a live row still
outranks the stale handed-off row).
2026-09-23 06:31:43 -07:00
John Paul Soliva
cb819d91c3 fix(gateway): run a chat /restart outside the requester's profile scope
Under multiplex, a message from a served profile's bot (or a default-bot
chat routed to a named profile) is handled inside that profile's
_async_profile_runtime_scope. /restart calls request_restart() from
there, and the bare asyncio.create_task() copied the handler's context,
so the whole restart orchestration kept the profile's HERMES_HOME
override, secret scope and terminal scope after the handler returned.

On a host without a service manager (shell/tmux runs, WSL, Termux, the
Windows Scheduled Task launcher) that task spawns the detached watcher,
whose env comes from build_subprocess_env(inherit_profile_home=True), so
it relaunched `hermes gateway restart` with HERMES_HOME set to the named
profile. The named-profile guard refuses that with exit 78, and the host
gateway never came back for any profile. On service-managed hosts the
same leaked scope sent stop()'s pending-message flushes to
<profile>/pending_messages, which boot recovery never reads.

Spawn the restart task in an empty Context, the same isolation
_spawn_supervised already uses for host-level tasks, so the watcher env,
stop() and its flushes all resolve the launch home.
2026-09-23 06:31:43 -07:00
teknium1
1a4da74add docs(gateway): --replace beside a standalone owner; the retired 20-replace.conf drop-in 2026-09-23 06:27:24 -07:00