`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
Review follow-ups on the new `hermes sessions set-journal-mode` verb:
- The header probe used os.pread, which does not exist on Windows, while the subparser is
registered unconditionally — the command died there with an uncaught AttributeError. It now
reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
(apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
bail in the command's own error style; every open/read is guarded.
The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.
The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
Review follow-ups on the stale-code tick yield gate:
- `_gate_mocks` patches `_current_gateway_code_sha` with `raising=False` so the
base tree (where the symbol does not exist yet) fails only on the two new
invariants instead of erroring the whole module out with AttributeError.
- Note at the pid comparison that `get_running_pid(pid_path=None)` can fall back
to `get_runtime_status_running_pid()`, which derives the pid from the same
record — the equality is tautological in that branch and the real proof there
is `runtime_status_is_stale` + `code_sha`.
- Docstring says what the record actually is: `gateway_state.json` is per-HOME
and last-writer-wins, the pid equality binds it to the lock holder, and the
`--replace` takeover window fails open by design.
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).
Co-authored-by: webtecnica <webtecnica@gmail.com>
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
The new test read result.expanded (a bool) as if it were the expanded
message, so it raised TypeError in CI instead of exercising the guard.
Check the body on result.message and the refusal on result.warnings;
still red without the production allow-root and green with it.
Desktop persisted a large paste under Electron's userData dir and attached
it as `@file:<abs path>`; the tui_gateway prompt path expands that reference
with `allowed_root=cwd`, so `_resolve_path` refused it with "path is outside
the allowed workspace" whenever the chat's cwd was not an ancestor of the
paste dir (always on Linux/macOS, and on Windows for any project cwd).
Electron now writes pastes under `<HERMES_HOME>/composer-pastes`, and
`agent/context_references.py::_resolve_path` admits exactly that anchored
directory (active profile home and global root, via file_safety._hermes_dirs)
as the one root besides `allowed_root`. A sibling directory that merely
contains the substring stays refused; the credential deny-list still runs
on the admitted path.
Supersedes the substring whitelist proposed in #117150.
Fixes#117149
A Desktop whose active connection is remote (SSH/remote/cloud, including the
registry primary) owns no local messaging gateway, yet the update hand-off
always ran `hermes update --gateway`. On hosts where launchd/service recovery
fails, the updater falls back to a detached local `gateway run --replace`;
with the same Telegram bot token as the remote VPS gateway, the two processes
compete for getUpdates and Telegram rejects one consumer, taking the
production bot offline (#117529).
Pass the ownership down the hand-off: globalRemoteActive() now adds
--no-gateway (posix) / -NoGateway (windows) when the Desktop is remote-served,
and both orchestrators drop --gateway from every update invocation (initial +
retry). The local-ownership default keeps --gateway exactly as before.
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace. Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:
- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
also covers everything beneath it); a matching root never spawns, logged once at INFO. A
non-list value fails closed — WARNING naming the expected shape, every root skipped — because
silently excluding nothing would re-pay the stall the key was meant to avoid.
Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.
Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.
Part of #116472 (request 4).
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).
In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.
Fixes#117051
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").
Part of #117051
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].
Fixes#117246
When the owner dies before a background child records a result, the
recovered event now includes the last lines of each live transcript and a
one-line git snapshot of the owner cwd, and the single-delegation renderer
shows the recovery diagnostics the batch renderer already had.
Fixes#116000
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).
_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".
(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
Keep staleness (edit seen on the next poll), missing-file pickup, parsed-once and the uncached-writer contracts; the per-file mutation-poison and cross-profile cases restate the copier.
`profiles.list` parsed each profile's profile.yaml TWICE per request: once in the listing body
(`read_profile_meta`, for description/display_name) and again in `_profile_ui_meta_fields`, whose
`_read_profile_yaml` is a plain `yaml.safe_load` with no cache. Same 5s-per-connection poll as the
session fields in the previous commit, so it reuses that module rather than adding a second one.
Measured on the registered handler with include_sessions=false, 9 profiles: 8.80ms -> 5.38ms and 9
profile.yaml parses per pass -> 1 (the default profile, which has no profile.yaml and so is never
cached). Together with the session-field memo a full poll at 8 profiles x 300 sessions is 5.1ms,
against 42.3ms before either.
The module now holds two memos over one shared core, each keyed on the file(s) its fields derive
from: the session store for one, profile.yaml — mtime, size AND inode, since the atomic writers
rename a temp file into place — for the other. Session fields are copied shallowly, ui_meta deeply:
it is a nested mapping handed to a client.
Three things stay outside the memo: `has_avatar` remains live, because it is three stats and must
react to an avatar dropped in without a profile.yaml edit; `_read_profile_yaml` keeps its uncached
contract, because its other caller `_configure_ui_meta` reads that document, mutates it and writes
it back under a CAS; and the wire key order (ui_meta_revisions, ui_meta, has_avatar) is unchanged.
Regressions drive the real handler: an edited ui_meta shows on the very next listing; the CAS
writer round-trips through it; an avatar added without touching profile.yaml is still seen; a
profile without profile.yaml is not cached and picks one up; mutating a returned ui_meta does not
poison the next listing; an unchanged file is parsed once across three polls; and the wire key
order holds. Breaking the shared signature check fails four tests across both memos; caching
has_avatar fails the avatar one.
Fixes#117383
`list_profiles()` parsed three YAML files per profile on every call — config.yaml through
`read_user_config_raw` ("no cache", by contract) plus profile.yaml and distribution.yaml through
`_load_yaml_dict`. It is the shared body of `GET /api/profiles` and `profiles.list`, which the Bots
roster polls every 5s per connection, so an idle machine re-parsed every profile's YAML twelve
times a minute to produce the same handful of strings.
config.yaml is the expensive one: the installer seeds it by copying the annotated template verbatim
(`cp cli-config.yaml.example config.yaml`), 119,268 bytes / 2,267 lines of mostly comments, for
`model.default` and `model.provider`.
Measured, 9 profiles, no gateway locks so the liveness path stays out of it: 16.1ms -> 1.6ms with
the installer template, 6.4ms -> 1.6ms hand-trimmed; 27 YAML parses per pass -> 2, and those two
are the default profile's ABSENT files, correctly not cached. The pass is now independent of config
size.
The memo keys on each file's (mtime_ns, size, inode) — the atomic writers rename a temp file into
place, so a rewrite always lands a new inode even inside one mtime tick — and caches only the small
DERIVED values. `read_user_config_raw` and `_load_yaml_dict` keep their uncached contract: callers
write those documents back, and a stale read there could overwrite a newer file. `read_profile_meta`
hands out a copy so a future mutating caller cannot poison the next reader.
Regressions drive the real readers: an edited config/profile/distribution file is seen on the very
next read, `write_profile_meta` round-trips, a missing file is not cached and is picked up when
created, mutating a returned meta does not poison the next reader, an unchanged file is parsed once
across three reads, and `read_user_config_raw` still re-reads. Ignoring the signature fails four.
Fixes#117378
`profiles.list` is polled every 5s per connection, and per profile it opened a fresh read-only
SessionDB, ran `list_sessions_rich(limit=20)`, `get_session_by_title("Bot Chat")`,
`get_compression_tip` and two preview scans, then tore the handle down — to produce three fields
that are a pure function of that profile's session store. An idle fleet re-derived every bot's row
17,280 times a day for the same bytes, and the cost grew with session history.
Measured on the registered handler (9 profiles, stores seeded through the real write path):
25 sessions/profile 25.9ms -> 7.4ms; 300 sessions/profile 42.3ms -> 6.7ms. The session fields were
75-81% of the poll; what remains is the `include_sessions: false` floor, and it no longer grows
with history.
The memo keys on the store's signature — the newest mtime across `state.db` and its `-wal`
sidecar, the signal the change watcher already trusts for `sessions.changed`, plus each file's
size so a write landing inside one mtime tick still invalidates. A profile with no store is never
cached: nothing to read, and a store created later must be picked up. The handler already treats
this path as poll-sensitive (`skill_count` was made lazy for the same reason); this is the
`state.db` half of that.
Regressions drive the real handler: a new session, and an append inside an existing one, both show
on the very next poll; a store-less profile is not memoised and picks one up; profiles never
answer for each other; and an unchanged store is opened once across three polls. Ignoring the
signature fails the first two.
Fixes#117257
The suite already covers the shape where assignee != reviewer
(test_review_handoff_without_live_run_attributes_run_to_implementer). The
mirror case -- a card created already assigned to its reviewer -- was the gap.
Asserts the honest provenance (no implementer, NULL run profile) and, more to
the point, that request_changes refuses instead of routing the rejection back
to the reviewer. Fails on the parent commit with implementer='reviewer-a'.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test_every_minting_site_imports_the_one_helper, test_session_startup_calls_maybe_pull_org_skills
and test_cli_session_store_unavailable_banner_honors_policy located their code by reading cli.py;
the phases now live in hermes_cli/cli_init_mixin.py and hermes_cli/cli_tui_runtime_mixin.py.
register_plugin_provider setdefault-ed aliases regardless of provenance, so a
$HERMES_HOME plugin alias colliding with an existing row kept resolving to the
old row while providers.get_provider_profile already followed the user profile;
override_registry_row left the display name untouched on a same-name
replacement. mirror_aliases applies the provider_source ownership rule that
e5e7fbcd27 introduced for endpoints.
Fixes#116668
_Capture.read polls read_available against the stop event before every
InputStream.read, close() aborts before closing, and _halt_thread joins
outside the detector lock and keeps the handle while the thread is alive.
Fixes#117096
`_suppress_mouse_residue_early()` runs at import time, before
`_apply_profile_override()` sets HERMES_HOME, so the first read of
`display.interface` always saw the default home. The result was memoised
in a cache keyed on nothing, so every later caller — the Termux fast
paths and the TUI launch decision, which all run after the profile is
applied — was pinned to the default home's interface for the whole run.
`hermes -p coder` with `display.interface: tui` in that profile booted
the classic REPL, and the mirror case booted the TUI for a profile that
asked for cli.
Key the cache on the config path it read. The hot path still parses the
YAML once per home, and the first call after the process is re-homed
re-reads the file it should have read.
Fixes#116902
`_hermes_home_for_pid` fell back to the INSPECTING process's `Path.home()`
when the target's environment carried neither HERMES_HOME nor HOME, which
is the normal shape for a systemd/launchd unit with a scrubbed
environment. The dashboard was then attributed to whichever user ran the
command, so `hermes update` and `--stop` read the wrong profile root and
could act on the wrong backend.
A process without HOME resolves its own default home from the password
database, so resolve the target's owner the same way (psutil, then the
/proc owner) before falling back. An unreadable owner or entry keeps the
previous behaviour rather than resolving to nothing.
Fixes#116906
The binding half of _verify_reapable_browser_daemon accepted the socket-dir
basename anywhere in argv. That basename is predictable from the session
name, so a recycled PID whose argv merely mentioned it (a grep, a shell)
passed both gates and was tree-killed. Require the full normalized path as
an argv token (bare or --flag=path) or the environ match.
Fixes#116884
The 5-branch if/elif ladder in CLICommandsMixin._handle_browser_command
becomes a _BROWSER_SUBCOMMANDS word -> handler table; adding a subcommand
is one row plus a usage line. Subcommand word is case-insensitive, the
argument keeps its case (CDP URL paths are case-sensitive).
Review follow-up: _request_google_offline_access re-wrapped
context.redirect_handler on every _perform_authorization, so N
re-authorizations ran N nested wrappers. Mark the wrapper and return early
when it is already installed; it reads the issuer at call time so the
single wrapper is right for any issuer. The six test functions fold into
two — browser flow (offline access + consent once, wrap-once, other-issuer
and lookalike controls) and device flow (issuer slash-normalization +
access_type on the device request).
Google issues refresh tokens only for authorization requests carrying
access_type=offline (its idiom where OIDC servers use the offline_access
scope MCP discovery would advertise), so a Google-hosted MCP server
(Gmail/Calendar) authorized in a browser dies with the short-lived access
token: later reconnects (gateway, cron) find no refresh token and fail
back to an interactive login they cannot perform (#117510).
The device flow could not even start: the advertised authorization server
reaches issuer validation root-slash-stripped on one side only, so
Google's slash-less document issuer is rejected against a slash-terminated
expected issuer. Compare both sides normalized, the convention
_metadata_issuer and the refresh-token issuer binding already use; any
other mismatch is still rejected.
(cherry picked from commit bc08ef9572616d09ecfb9a0085b32b3582af5f5e)
Drop the duplicated _GRANT_DEAD_CODES frozenset and the relogin_required
plumbing: the pool's plugin recovery already treats exc.code in
hermes_cli.auth._OAUTH_GRANT_DEAD_CODES as terminal, so _token_http_error
imports that set late and just sets code=<json error>.
A spent refresh token returned HTTP 400 invalid_grant, and the helper mapped that to oauth_refresh_failed so the pool EXHAUSTED the row. Parse the JSON error and DEAD it. hermes auth add with an alias wrote the pool under the typed name, so status against the profile name looked logged out.
The budget-log tests drove _log_agent_budget() directly, so deleting the call in start()
stayed green. The 'unlimited' case now runs the real startup path under a temp HERMES_HOME
and asserts the line from caplog (red with the start() call removed).