fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide

`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
This commit is contained in:
teknium1
2026-09-20 16:02:49 -07:00
committed by Teknium
parent 12fb5eb476
commit 6ba45b0e06
13 changed files with 297 additions and 16 deletions

View File

@@ -335,6 +335,43 @@ def foreign_state_db_holders(db_path: Path) -> List[Tuple[int, str]]:
return holders
def held_store_refusal(db_path: Path, *, command: str, force_hint: Optional[str] = "--force") -> Optional[str]:
"""Operator-facing refusal for structural maintenance (VACUUM, index rebuild, bulk delete) while another
process holds ``db_path`` or a WAL sidecar; ``None`` when the store is provably quiet.
Running ``hermes sessions optimize-storage`` underneath a fleet of live gateways put every agent into
the retired-WAL refusal until all writers were stopped (#110054). Same fail-closed scan doctor and
repair use: an incomplete scan refuses too, it never reads as an all-clear.
"""
holders = foreign_state_db_holders(db_path)
if not holders:
return None
from hermes_constants import profile_cli_selector
from hermes_state_errors import STORAGE_RECOVERY_DOCS_URL
by_pid: dict[int, Set[str]] = {}
unknown: List[str] = []
for pid, target in holders:
if pid <= 0 or target.startswith("uninspectable"):
unknown.append(target)
else:
by_pid.setdefault(pid, set()).add(Path(target.removesuffix(" (deleted)")).name)
lines = [f"Refusing `hermes sessions {command}`: another process is using {db_path}."]
lines += [f" {describe_holder_pid(pid)}: {', '.join(sorted(by_pid[pid]))}" for pid in sorted(by_pid)]
if unknown:
lines.append(f" cannot prove the database is quiet (holder scan incomplete: {unknown[0][:120]})")
profile_arg = profile_cli_selector()
lines += [
"Rewriting the database under a live writer is how every agent ends up refusing turns with the "
"retired state.db-wal error. Nothing is lost.",
f"Stop them first (`hermes {profile_arg}gateway stop`, quit the Desktop app, pause cron), then re-run.",
]
if force_hint:
lines.append(f"Override with {force_hint} if you accept the risk.")
lines.append(f"Recovery guide: {STORAGE_RECOVERY_DOCS_URL}")
return "\n".join(lines)
def live_writer_holds_db(
db_path: Path,
*,