Commit Graph

39601 Commits

Author SHA1 Message Date
teknium1
dfaf016dbc fix(sessions): resolve the held-store scan path from the store resolver, not the db object
The admission gate read `db.db_path`, which made the refusal depend on whatever object
`SessionDB` resolved to; the CLI tests substitute a lightweight double and CI went red with
AttributeError: 'FakeDB' object has no attribute 'db_path'. The path now comes from
`_default_db_path()` — the exact resolver `SessionDB()` itself uses two lines above — so the
scan targets the same file in production and stays reachable regardless of the db object.
2026-09-20 20:13:33 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
12fb5eb476 fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb:

- The header probe used os.pread, which does not exist on Windows, while the subparser is
  registered unconditionally — the command died there with an uncaught AttributeError. It now
  reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
  admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
  running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
  (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
  silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
  header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
  bail in the command's own error style; every open/read is guarded.

The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
2026-09-20 20:13:25 -07:00
teknium1
96da5d97fc feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.

The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
2026-09-20 20:13:25 -07:00
teknium1
95b1f4c855 fix(cron): scope the yield predicate's claims to what the record proves
Review follow-ups on the stale-code tick yield gate:

- `_gate_mocks` patches `_current_gateway_code_sha` with `raising=False` so the
  base tree (where the symbol does not exist yet) fails only on the two new
  invariants instead of erroring the whole module out with AttributeError.
- Note at the pid comparison that `get_running_pid(pid_path=None)` can fall back
  to `get_runtime_status_running_pid()`, which derives the pid from the same
  record — the equality is tautological in that branch and the real proof there
  is `runtime_status_is_stale` + `code_sha`.
- Docstring says what the record actually is: `gateway_state.json` is per-HOME
  and last-writer-wins, the pid equality binds it to the lock holder, and the
  `--replace` takeover window fails open by design.
2026-09-20 20:13:19 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
fangliquanflq
153b5f0145 fix(cron): compare full gateway code identities 2026-09-20 20:13:19 -07:00
fangliquanflq
bc04c35f6a fix(cron): verify fresh gateway before yielding ticks 2026-09-20 20:13:19 -07:00
Austin Pickett
1a1f4a59e2 fix(auth): explicit-provider gate uses the credential resolver's reader
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-20 22:46:49 -04:00
xxxigm
4ea57fa61b fix(auth): count profile .env keys in the explicit provider gate
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
2026-09-20 22:46:49 -04:00
xxxigm
050ea53aba fix(desktop): honor Applies-to on Custom Endpoints
The page only showed a read-only active-profile note, so endpoint
saves followed the left-rail Bot instead of the Settings chips
Accounts and API keys already share.
2026-09-20 22:46:49 -04:00
teknium1
8ba5fe9c16 test: assert composer-paste guard on result.message/warnings, not the expanded flag
The new test read result.expanded (a bool) as if it were the expanded
message, so it raised TypeError in CI instead of exercising the guard.
Check the body on result.message and the refusal on result.warnings;
still red without the production allow-root and green with it.
2026-09-20 19:45:30 -07:00
teknium1
60dd9778c6 fix(desktop): large text pastes attach from HERMES_HOME/composer-pastes when the chat cwd is elsewhere
Desktop persisted a large paste under Electron's userData dir and attached
it as `@file:<abs path>`; the tui_gateway prompt path expands that reference
with `allowed_root=cwd`, so `_resolve_path` refused it with "path is outside
the allowed workspace" whenever the chat's cwd was not an ancestor of the
paste dir (always on Linux/macOS, and on Windows for any project cwd).

Electron now writes pastes under `<HERMES_HOME>/composer-pastes`, and
`agent/context_references.py::_resolve_path` admits exactly that anchored
directory (active profile home and global root, via file_safety._hermes_dirs)
as the one root besides `allowed_root`. A sibling directory that merely
contains the substring stays refused; the credential deny-list still runs
on the admitted path.

Supersedes the substring whitelist proposed in #117150.

Fixes #117149
2026-09-20 19:45:30 -07:00
teknium1
3d0bdce9f9 fix(desktop): prove the sign-in wall from curl's arrival URL and title too
Port the reporter's (@tianyu-liu) url_effective half: curl reports where it
landed via --write-out on a tail buffer that survives the body budget, and
isAuthWall decides on arrival URL -> sign-in title -> markup, so the Apps
Script ServiceLogin shape (no markup id) is a wall too and a wall's own title
is never used as the link label. Tests trimmed to two invariants.

Co-authored-by: tianyu-liu <tianyu-liu@users.noreply.github.com>
2026-09-20 19:37:57 -07:00
Finn763
1c81ad3485 fix(desktop): never load a sign-in wall in the hidden link-title renderer
The cookieless title partition answers a Google Workspace link with Googles

(cherry picked from commit 31e7dd19fafb13b8350bf3271d91036f243079b3)
2026-09-20 19:37:57 -07:00
liuhao1024
04845f5f3e fix(desktop): skip the local gateway restart on update when the Desktop is remote-served
A Desktop whose active connection is remote (SSH/remote/cloud, including the
registry primary) owns no local messaging gateway, yet the update hand-off
always ran `hermes update --gateway`. On hosts where launchd/service recovery
fails, the updater falls back to a detached local `gateway run --replace`;
with the same Telegram bot token as the remote VPS gateway, the two processes
compete for getUpdates and Telegram rejects one consumer, taking the
production bot offline (#117529).

Pass the ownership down the hand-off: globalRemoteActive() now adds
--no-gateway (posix) / -NoGateway (windows) when the Desktop is remote-served,
and both orchestrators drop --gateway from every update invocation (initial +
retry). The local-ownership default keeps --gateway exactly as before.
2026-09-20 19:30:16 -07:00
Chukuwebuka-2003
338b31226d fix(desktop): surface an externally-terminated renderer with a recovery page (#116472)
The OS/Chromium can SIGKILL a renderer while the window is live (memory reclaim, an external kill). Hermes never does this itself and never reloads it (a killed-after-close window must not pop back up), so the window sat silent with only a desktop.log line.

- window-renderer-lifecycle.ts: optional onRendererTerminated(details) callback, invoked for an unrecoverable render-process-gone on a live (not destroyed) window; recoverable crashes still reload; expected teardown still does nothing.
- renderer-load-error-page.ts: optional escaped title for the recovery page.
- main.ts: wire the primary window's callback to a visible recovery page (reason + Reload), guarded against intentional quit/handoff (isQuittingForHandoff || backendShutdown.hasStarted()).
- tests: three-arm lifecycle invariant + custom-title test.

Electron-only slice harvested from #116592 (commit 9bb16bd6ae) per the split advised on #117084; the agent half of #116472 lives in #117084.

(cherry picked from commit 5bdfe7ed00c4b3f764d4044fdfa527f978140cf3)
2026-09-20 19:21:46 -07:00
teknium1
c4f5dd0dad test(desktop): widen the roster-enumeration scan window past the ssh install-id branch
enumerateRegistryAgentSources grew when the ssh branch started carrying the
install id learned by the inventory probe, so the 3_700-char slice in
backend-dial-claim.test.ts stopped before getJsonForBackend('/api/profiles')
and CI went red. Same fix the sibling update-all case already applied.
2026-09-20 19:13:13 -07:00
John Paul Soliva
82e147ffb8 fix(desktop): an ssh connection learns its backend install id, so two addresses for one machine collapse
#88828 gave every connection a stable backend identity so one physical install is one roster row
per profile — but wired the probe into the non-ssh branch only. `probeConnectionInstallId` has a
single caller, in the enumerator's `else`; the other writer, `rememberConnectionInstallId`, sits
in `hermes:connections:test` after its ssh branch has already returned. So an ssh connection's
`installId` was always undefined, `buildAgentRoster` always fell back to the per-connection key,
and a Mac registered twice over ssh (alias + address, LAN + WAN) showed every bot twice with
`@name-label` handles, pushed two rows per bot into every gateway's relay roster
(`write_remote_roster` dedups by connection_id + profile), and made a bare-handle DM to it
`ambiguous` — asking the sender to disambiguate between two ids naming one machine.

The id is a plain file at the install root, and the ssh inventory probe already reads the remote
home over its own session without spawning a dashboard, so `readRemoteInstallId` reads it there:
from the INSTALL root (a connection pinned at `<root>/profiles/<name>` must report the same
backend as one pointed at the root), read-only — a missing file stays missing, since minting
identity is the install's job and not a visiting client's — and only a well-formed id counts, so
an older backend keeps today's no-id behaviour.

Fixes #117226

(cherry picked from commit 3d596f9b8fc898c74f96a378588a0b4601bfe39d)
2026-09-20 19:13:13 -07:00
John Paul Soliva
07888c880f fix(desktop): removing or re-pointing a connection forgets what was cached about it
Three main-process caches are keyed by connection id — the ssh profile list, its retry stamp and
the backend install id — and neither lifecycle event that invalidates them evicted anything.
`hermes:connections:remove` stopped pooled backends and told renderers to dispose their sockets;
`saveRegistryConnection`'s dial-material branch recycled sockets for exactly this reason, its own
comment noting they "point at the OLD target". The caches were not in scope for either.

Ids are recycled label slugs (`connectionIdForLabel` suffixes only against CURRENTLY taken ids),
and `shouldRetrySshInventory` never retries a cached success — "Cached successes never retry" —
so a connection re-added under the same label, or simply pointed at another host, kept serving the
previous machine's profiles for the rest of the app session. Clicking one of those bots dialled
the new box and failed, the relay pushed the phantom names to every peer gateway, and a peer bot
DMing one got 4092 `no profile '<name>' on this gateway` while the new machine's real bots stayed
invisible. Only a restart or a manual Test recovered.

The three maps move into `connection-caches.ts`, where the invariant they share can be stated and
tested — each is valid only while its id names the same machine — and `evictConnectionCaches()`
is called from both handlers. Every existing reference in main.ts is unchanged: the module
exports the same Map objects.

Fixes #117227

(cherry picked from commit 03847f66e8eae50f1972511ec89421b0298dbd22)
2026-09-20 19:04:11 -07:00
teknium1
101f1c598a test(kanban): done-order board test completes tasks with result evidence (#117483) 2026-09-20 19:03:47 -07:00
teknium1
c746249861 fix(kanban): CLI and kanban_complete tool report an empty-completion refusal
EmptyCompletionError is raised by complete_task; without a handler the
CLI and the tool would surface a traceback instead of the actionable
refusal (add --result/--summary). Same shape as the HallucinatedCards gate.

Part of #117483
2026-09-20 19:03:47 -07:00
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
teknium1
af995a30aa chore: map contributor email for @mathiashuque 2026-09-20 19:02:21 -07:00
Mathias Huque
86aa9dec9a fix(mcp-oauth): classify discovery URLs by path
(cherry picked from commit a258bc20cc4f5aa602d2c5f0618893c977e2c2b6)
2026-09-20 19:02:21 -07:00
Mathias Huque
1fd7ff0f5c fix(mcp-oauth): avoid consuming MCP resource SSE as OAuth metadata
(cherry picked from commit 04a5eae070a9ae746b93c52590cf2ef8486957f5)
2026-09-20 19:02:21 -07:00
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
teknium1
16b749f61f fix(desktop): keep the chat composer mounted through transient session loaders
The chat bar was gated on `!loadingSession`, so any one-render flip of the
routed session into its loading state (periodic sidebar / live-status
refresh dropping the routed row, a hydrate swapping the transcript through
an empty frame) unmounted the composer: draft, attachments, caret and focus
vanished and came back every ~30s. Once a route has rendered with its
composer, a later transient loader for the same route now keeps it mounted;
a route change, the exhausted state and watch windows still hide it.

The unmount also fired `placeCaretEnd` / `placeCaretAtOffset` on the
detached editor from the draft repaint, throwing
`addRange(): The given range isn't in document` inside React's commit and
looping the reconciler (the visible flicker). Both caret helpers now bail
when the editor or the computed range is no longer connected.

Fixes #117375
Fixes #117285
2026-09-20 18:53:37 -07:00
Chukuwebuka-2003
7d3c0b2f94 fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.

Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.

Part of #116472 (request 4).
2026-09-20 18:53:10 -07:00
teknium1
442636bf44 fix(update): settle latest.json only when the live fleet discharges the pending restart
The catch-up rewrote the receipt's fleet matrix before checking whether that
actually cleared the debt, so an exit-1 run (marker still owes a gateway,
blocked serve_restart_pending storage) left latest.json mutated —
tests/hermes_cli/test_update_scoped_reconciliation.py pins the receipt as
byte-identical on the failure path. Decide on the in-memory settled copy and
write only when the pending restart is discharged (#117051).
2026-09-20 18:52:45 -07:00
teknium1
29c4ec8fa7 fix(update): settle a stuck failed receipt once the live fleet serves the checkout code
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).

In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.

Fixes #117051
2026-09-20 18:52:45 -07:00
teknium1
d6b0d37ece fix(update): pending fleet restart leaves gateways already on the checkout code alone
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").

Part of #117051
2026-09-20 18:52:45 -07:00
teknium1
eaef5ec717 feat(update): hermes update --list-venv-holders prints the venv guard's holders as JSON, exit 3
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].

Fixes #117246
2026-09-20 18:51:36 -07:00
teknium1
287cbb7edb fix(delegation): keep encoding= on the text=True line so the line-based footgun scanner sees it 2026-09-20 18:50:37 -07:00
teknium1
4f46bbfb92 fix(delegation): recovery-hint git probe decodes output as utf-8 (Windows footgun) 2026-09-20 18:50:37 -07:00
teknium1
a2b6f53245 feat(delegation): recovery event carries transcript tails and owner git state
When the owner dies before a background child records a result, the
recovered event now includes the last lines of each live transcript and a
one-line git snapshot of the owner cwd, and the single-delegation renderer
shows the recovery diagnostics the batch renderer already had.

Fixes #116000
2026-09-20 18:50:37 -07:00
xielevi
5d9af50aa0 fix(desktop): gate the plain-text token warning on secure storage being unavailable
The Settings -> Gateway "Token stored in plain text" banner predates the
opt-in keychain policy: it was written when a plain token could only exist
on keyring-less Linux, and it fires off the token encoding alone. Since
keychain encryption became opt-in (default off), every saved token is
plain, so the banner shows for every default user, while
probeSecureTokenStorage deliberately reports availability in that mode
"so no plain-text warning banners fire: storing plaintext is the user's
chosen (default) mode, not a degraded state."

- Compute the renderer-facing signal in a new
  hardening.resolveRemoteTokenPlainText helper: true only when the token
  is plain AND secure storage is genuinely unavailable (the keyring-less
  opt-in path the banner was written for), and never under the
  HERMES_DESKTOP_REMOTE_URL env override.
- Behavioral tests for the truth table plus a source assertion pinning
  the call site; both proven red against the ungated logic.
- Scope the Linux-only keyring names in the plain-text copy (banner and
  save-time confirm, en/zh/zh-hant/ja/ru) to Linux, since both can render
  on any platform.

Fixes #117269
2026-09-20 18:47:20 -07:00
fangliquanflq
4774bb0a1a fix(gateway): dedupe non-editable interim finals
(cherry picked from commit 9084be9bd97df053967943d6aa0b5f7b9e7e4b48)
2026-09-20 18:46:47 -07:00
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
teknium1
20fc19be8f fix(tui-gateway): an empty -wal sidecar created by our own read-only open does not count as a store change
The first read-only SessionDB open of a store that never had a WAL file creates a 0-byte state.db-wal; with it in the signature the second poll always missed. An empty WAL has no frames, so it is excluded from the store signature.
2026-09-20 18:45:07 -07:00
teknium1
2a7658989f test: trim the roster-poll memo tests to the invariants
Keep staleness (edit seen on the next poll), missing-file pickup, parsed-once and the uncached-writer contracts; the per-file mutation-poison and cross-profile cases restate the copier.
2026-09-20 18:45:07 -07:00
John Paul Soliva
0111d93178 perf(gateway): the roster's ui_meta fields ride the same memo, keyed on profile.yaml
`profiles.list` parsed each profile's profile.yaml TWICE per request: once in the listing body
(`read_profile_meta`, for description/display_name) and again in `_profile_ui_meta_fields`, whose
`_read_profile_yaml` is a plain `yaml.safe_load` with no cache. Same 5s-per-connection poll as the
session fields in the previous commit, so it reuses that module rather than adding a second one.

Measured on the registered handler with include_sessions=false, 9 profiles: 8.80ms -> 5.38ms and 9
profile.yaml parses per pass -> 1 (the default profile, which has no profile.yaml and so is never
cached). Together with the session-field memo a full poll at 8 profiles x 300 sessions is 5.1ms,
against 42.3ms before either.

The module now holds two memos over one shared core, each keyed on the file(s) its fields derive
from: the session store for one, profile.yaml — mtime, size AND inode, since the atomic writers
rename a temp file into place — for the other. Session fields are copied shallowly, ui_meta deeply:
it is a nested mapping handed to a client.

Three things stay outside the memo: `has_avatar` remains live, because it is three stats and must
react to an avatar dropped in without a profile.yaml edit; `_read_profile_yaml` keeps its uncached
contract, because its other caller `_configure_ui_meta` reads that document, mutates it and writes
it back under a CAS; and the wire key order (ui_meta_revisions, ui_meta, has_avatar) is unchanged.

Regressions drive the real handler: an edited ui_meta shows on the very next listing; the CAS
writer round-trips through it; an avatar added without touching profile.yaml is still seen; a
profile without profile.yaml is not cached and picks one up; mutating a returned ui_meta does not
poison the next listing; an unchanged file is parsed once across three polls; and the wire key
order holds. Breaking the shared signature check fails four tests across both memos; caching
has_avatar fails the avatar one.

Fixes #117383
2026-09-20 18:45:07 -07:00
John Paul Soliva
55a8f42329 perf(profiles): the listing re-reads a profile's YAML only when that file has changed
`list_profiles()` parsed three YAML files per profile on every call — config.yaml through
`read_user_config_raw` ("no cache", by contract) plus profile.yaml and distribution.yaml through
`_load_yaml_dict`. It is the shared body of `GET /api/profiles` and `profiles.list`, which the Bots
roster polls every 5s per connection, so an idle machine re-parsed every profile's YAML twelve
times a minute to produce the same handful of strings.

config.yaml is the expensive one: the installer seeds it by copying the annotated template verbatim
(`cp cli-config.yaml.example config.yaml`), 119,268 bytes / 2,267 lines of mostly comments, for
`model.default` and `model.provider`.

Measured, 9 profiles, no gateway locks so the liveness path stays out of it: 16.1ms -> 1.6ms with
the installer template, 6.4ms -> 1.6ms hand-trimmed; 27 YAML parses per pass -> 2, and those two
are the default profile's ABSENT files, correctly not cached. The pass is now independent of config
size.

The memo keys on each file's (mtime_ns, size, inode) — the atomic writers rename a temp file into
place, so a rewrite always lands a new inode even inside one mtime tick — and caches only the small
DERIVED values. `read_user_config_raw` and `_load_yaml_dict` keep their uncached contract: callers
write those documents back, and a stale read there could overwrite a newer file. `read_profile_meta`
hands out a copy so a future mutating caller cannot poison the next reader.

Regressions drive the real readers: an edited config/profile/distribution file is seen on the very
next read, `write_profile_meta` round-trips, a missing file is not cached and is picked up when
created, mutating a returned meta does not poison the next reader, an unchanged file is parsed once
across three reads, and `read_user_config_raw` still re-reads. Ignoring the signature fails four.

Fixes #117378
2026-09-20 18:45:07 -07:00
John Paul Soliva
8889060bbf perf(gateway): the roster poll reuses a profile's session fields while its store has not moved
`profiles.list` is polled every 5s per connection, and per profile it opened a fresh read-only
SessionDB, ran `list_sessions_rich(limit=20)`, `get_session_by_title("Bot Chat")`,
`get_compression_tip` and two preview scans, then tore the handle down — to produce three fields
that are a pure function of that profile's session store. An idle fleet re-derived every bot's row
17,280 times a day for the same bytes, and the cost grew with session history.

Measured on the registered handler (9 profiles, stores seeded through the real write path):
25 sessions/profile 25.9ms -> 7.4ms; 300 sessions/profile 42.3ms -> 6.7ms. The session fields were
75-81% of the poll; what remains is the `include_sessions: false` floor, and it no longer grows
with history.

The memo keys on the store's signature — the newest mtime across `state.db` and its `-wal`
sidecar, the signal the change watcher already trusts for `sessions.changed`, plus each file's
size so a write landing inside one mtime tick still invalidates. A profile with no store is never
cached: nothing to read, and a store created later must be picked up. The handler already treats
this path as poll-sensitive (`skill_count` was made lazy for the same reason); this is the
`state.db` half of that.

Regressions drive the real handler: a new session, and an append inside an existing one, both show
on the very next poll; a store-less profile is not memoised and picks one up; profiles never
answer for each other; and an unchanged store is opened once across three polls. Ignoring the
signature fails the first two.

Fixes #117257
2026-09-20 18:45:07 -07:00
fangliquanflq
c906121a46 test(kanban): cover uncapped runtime output 2026-09-20 18:43:59 -07:00
fangliquanflq
1a4c68ce97 fix(kanban): expose task runtime limit in JSON 2026-09-20 18:43:59 -07:00
chadhouser
7ff57f0cb2 test(kanban): pin the review handoff of a card assigned to its own reviewer
The suite already covers the shape where assignee != reviewer
(test_review_handoff_without_live_run_attributes_run_to_implementer). The
mirror case -- a card created already assigned to its reviewer -- was the gap.

Asserts the honest provenance (no implementer, NULL run profile) and, more to
the point, that request_changes refuses instead of routing the rejection back
to the reviewer. Fails on the parent commit with implementer='reviewer-a'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 18:43:35 -07:00
chadhouser
79c4c363c6 fix(kanban): derive review implementer from the run, not the card's assignee
request_review recorded trow["assignee"] as the implementer on the
review_requested event and on the synthesized run. That is correct while a
worker holds the card, but wrong when the card was created already assigned to
its reviewer -- `kanban create --assignee <reviewer>` followed by
`request-review`. Both provenance records then name the reviewer, and
request_changes routes a rejection back to the profile that wrote the findings.

Derive the implementer after the reviewer is resolved, from the active run's
profile, falling back to the assignee only when it is not the incoming
reviewer. When no actor can be established the field is left unset, which
request_changes already refuses on rather than misrouting.

Refs #117229. Follow-up to #111064 / #111459, which fixed the run attribution
for the distinct-reviewer case and left the derivation itself unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 18:43:35 -07:00
teknium1
31e6242fa3 fix(tests): point the CLI minting/call-site probes at the new init and TUI-runtime mixins
test_every_minting_site_imports_the_one_helper, test_session_startup_calls_maybe_pull_org_skills
and test_cli_session_store_unavailable_banner_honors_policy located their code by reading cli.py;
the phases now live in hermes_cli/cli_init_mixin.py and hermes_cli/cli_tui_runtime_mixin.py.
2026-09-20 18:43:11 -07:00
teknium1
aaafa80577 refactor(cli): move HermesCLI init phases and TUI run-loop phases into mixin siblings (#116911)
Mechanical, behaviour-neutral extraction along the existing cli_*_mixin.py
pattern: the ten _init_* constructor phases become CLIInitMixin
(hermes_cli/cli_init_mixin.py) and the fourteen _tui_* run-loop phases
(input dispatch, after-turn, startup banner/prewarm/maintenance, application
build, signal handlers, shutdown) become CLITuiRuntimeMixin
(hermes_cli/cli_tui_runtime_mixin.py). __init__ and run() stay in cli.py as
the orchestrators. Every moved body is AST-identical to the base copy; cli.py
module names are resolved lazily through cli so monkeypatch seams survive.
cli.py 4835 -> 3960 lines.
2026-09-20 18:43:11 -07:00