`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
`container_backend_for_task` already catches everything inside `_terminal_env_type_for_task`
and `_uses_container_paths` and returns defaults, and every CheckpointManager has
`unsupported_backend_reason`: call the method directly (no getattr fallback), drop the
try/except in `unsupported_backend_reason` and the suppress() around the import+call in
turn_explainers (the one around record_agent_write stays). The rollback.restore fakes in the
server tests gain the method they now must provide.
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.
Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review follow-up (#113530): the manager no longer records the first container
backend, so a session whose terminal backend changes is answered by the backend
configured now, not by the first one seen. The checkpoint hooks simply skip
container-backed tasks; unsupported_backend_reason() classifies at call time.
The docs state what /rollback and the rollback.* RPCs do for container sessions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 2dc4c0d2b17d0f478118eca55742001eeed8bdb0)
With a container terminal backend (docker, singularity, modal, daytona,
vercel_sandbox, container plugins) file-tool paths keep container semantics,
but the checkpoint hook handed them to the host-side CheckpointManager: a path
that does not exist on the host produced a useless snapshot attempt, one that
happens to exist on the host snapshotted the wrong tree, the destructive
terminal branch did the same with the container cwd, and the post-write ledger
hashed the container path on the host so safe restore could trust unrelated
host content. Every failure was swallowed, so a docker user saw "No checkpoints
found for /home/admin" with nothing behind it.
Classify the task's backend the way the file tools do (_uses_container_paths)
and, for container-backed tasks, take no checkpoint and record no ledger entry;
/rollback prints the reason and refuses diff and restore for that session (a
host checkpoint that predates it belongs to another tree), and the
rollback.restore RPC returns the same reason as a failed restore. The refusal
classifies the session's configured backend directly (in the gateway under the
session's own identity and profile scope, as a turn binds them), so it holds
before the first mutation of the session. Local and ssh backends are untouched. This stops the
false protection; it does not add rollback support for containers (translating
bind mounts is a separate contract).
Tests: eight cases in tests/agent/test_tool_executor_checkpoint_paths.py through
the production classifier (a fake docker environment registered for the task, or
the configured backend): missing host path, colliding host tree (POSIX),
destructive terminal command, post-write ledger on a real host file, /rollback
and rollback.restore refusal in a fresh session, the local session still
restoring, and unchanged local behavior. Five fail on main on Windows, where the
collision case is skipped.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 97a5709e4980c8f85a5a640e6e5e11fe8f94affa)
Every hard interrupt reached `begin_iteration` through the same `_interrupt_requested`
flag, so the turn loop booked all of them as `interrupted_by_user` — a cron run killed
by the scheduler's inactivity watchdog, a turn aborted by the liveness watchdog, a lost
session turn lease and a gateway inactivity timeout all read as a human pressing stop,
and the investigation went to the wrong subsystem (#112647).
`interrupt()` already records a trusted category per interrupt (`_tool_interrupt_reason`,
fed by the `tool_reason` every producer can pass through `request_hard_interrupt`). The
exit reason is now derived from it: the three categories `interrupt()` itself mints for
human stops keep `interrupted_by_user` / `interrupted_during_api_call`; any other
category names its producer — `interrupted_by_system(cron_inactivity_watchdog)`,
`interrupted_during_api_call(turn_liveness_watchdog)`. The system producers that passed
no `tool_reason` (cron inactivity watchdog, turn liveness watchdog, lease loss, gateway
inactivity timeout) now name themselves, so the model-visible tool-cancellation text
says the same thing. `_publish_interrupt_state` logs ONE line naming the source so the
turn record and the log agree.
No change to WHEN anything interrupts. `interrupted_during_api_call` moves to the
prefix-matched explanation table so the parameterised form keeps its user copy.
Slim redo of #112652 by @KoNit-K, which added a parallel `issuer` attribute and
keyword; this reuses the existing `tool_reason` plumbing instead.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
gateway/run_notifications.py::_send_session_db_warning_notifications already
computed profile_arg but left `hermes doctor` (fts_index branch) and
`hermes gateway restart` (default branch) bare; agent/turn_finalizer.py had a
bare `hermes doctor` in the error fallback used when the explainer produced
no text. Same class as the previous commit. Also reflows the explainer
strings so continuation lines break at clause boundaries.
The replaced / deleted_wal / default persistence explanations printed bare
`hermes gateway stop` and `hermes doctor`, while `{home}` in the same sentence
was already profile-aware. On a multi-profile backend (Desktop serve) the
session whose state.db failed is not the process default, and a bare `hermes`
follows the sticky active_profile — so the copy-pasteable command stops or
inspects the wrong profile's database. corrupt / fts_index were pinned by #105887;
this applies the same `{profile_arg}` substitution to every cause
instead of the two-cause tuple.
Spotted via #110073 (@JoaoMarcos44), whose explainer hunk added the selector
to the since-rewritten deleted_wal runbook.
The turn-end file-mutation verifier only sees write_file/patch receipts. It
asserted "N file(s) were NOT modified this turn" whenever a call had failed,
which is wrong when the file was in fact changed afterwards through a path
that leaves no receipt (terminal redirect, execute_code) or when the
successful retry used another spelling of the same path (relative vs
absolute, separator/case variants on Windows): the state dict was keyed on
the model's raw `path` argument, so the pop never matched.
- Header now says what the recorder knows: "N file edit(s) FAILED this turn",
and asks the user to confirm what actually landed.
- Failure entries carry the task-resolved, normcase'd on-disk identity plus a
(mtime_ns, size) snapshot; a later success clears every entry with the same
identity regardless of spelling.
- At turn end `_file_mutations_still_failed` re-stats each target and drops
entries whose file changed since the failed call, so a receipt-less
mutation no longer produces a false footer.
- `tool_executor` passes the effective task id so relative paths resolve the
way the file tools resolved them.
Kept the deliberate first-error-per-path semantics (the pinned test says why);
did not add an "unverified" bucket for receipt-less non-error results, since
the built-in tools always return a receipt on success and it would only add
noise.
Co-authored-by: KoNit. <124019182+KoNit-K@users.noreply.github.com>
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
The recovery commands rendered on structural corruption — the turn explainer's
`session_persistence_failed`/corrupt body, the gateway's home-channel state.db
warning, and hermes_state_repair._persistent_repair_exhausted_error — already
interpolate the active profile's state.db path, but every `hermes ...` verb in them
was bare. A bare `hermes` follows the sticky `active_profile` file, so an operator
running the pasted `hermes doctor --fix` (or `hermes sessions recover` with a
relative source) from a named-profile incident could inspect or repair a different
profile's database (#105887).
hermes_constants.profile_cli_selector() renders `-p <name> ` for a named profile
home (default home and custom roots outside the profile tree render nothing: the
default is what a bare `hermes` already means, and a custom root is only reachable
via HERMES_HOME). Every command in the three guidance sites now carries it, and
the new `fts_index` guidance inherits the same interpolation.
Live check with HERMES_HOME=<root>/profiles/research and active_profile=other:
before `1. Run \`hermes doctor --fix\`` (targets "other"); after
`1. Run \`hermes -p research doctor --fix\`` and `hermes -p research sessions
recover --source <root>/profiles/research/state.db --inspect-only`.
Refs #105887
Reported-by: Cuttingwater
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.
One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".
Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).
Fixes#97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
The corruption explainer filled `{db_path}` from `_default_db_path()`, the
process default. A Desktop `serve` backend launched on the root home hosts
named-profile sessions whose SessionDB is `profiles/<name>/state.db`, so the
operator was told to inspect/repair a different profile's database. Pass the
agent's own `_session_db.db_path` from the turn finalizer; the process
default remains the fallback for agents without a bound store.
Reported in #105887.
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.
Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
When the redirect cap trips, the correction that cancelled the final attempt is
still sitting in _pending_redirect; finalize_turn's clear_interrupt() would drop
it silently. Drain it into the steer slot so it rides result["pending_steer"] and
becomes the next user turn on every surface that already honours that key.
Both new exit reasons get a turn-completion explanation so the user sees why the
turn stopped instead of an empty reply.
The corrupt-cause recovery guidance hardcoded `~/.hermes/backups/` while
every other path in the same message follows the active HERMES_HOME
(`{db_path}` is already interpolated). A custom-home or named-profile
deployment was told to restore from a directory that may not exist at all,
mid data-loss incident. Both sites (turn-completion explainer and gateway
startup broadcast) now interpolate `<hermes_root>/backups` via
get_default_hermes_root(), matching hermes_cli/backup.py's real backup
location.
Fixes#104250