Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.
- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.
- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.
- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.
- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
Reviewer findings from the formal /simplify-code pass, all verified:
1. Scanner blind spots (HIGH): five more locked pure readers were
invisible to the v1 gate — get_compression_fallback_streak and
get_compression_ineffective_count hid behind `conn = self._conn`
aliasing; list_gateway_sessions, find_session_by_origin and
search_sessions hid behind SQL held in variables/f-strings (the
scanner required unknown == 0 to flag). All five converted to
_read_ctx(); the scanner now (a) tracks self._conn aliases and
(b) flags lock blocks with NO proven write instead of silently
skipping unprovable SQL. Sabotage self-check extended to pin all
three detection classes (literal, alias, variable-SQL) plus a
mixed variable-SQL writer that must NOT fire.
2. get_meta reverted to the writer lock: its inline comment (present
on main) documents a real read-your-writes dependency —
fts_rebuild_step reads rebuild progress that a pooled WAL reader
cannot see mid-transaction. The blanket conversion had overridden
a documented design decision; it is now the single justified
_ALLOWED_LOCKED_READERS entry, replacing the dead
_enter_fts_fail_open entry (whose lock block counts 3 writes and
never needed allowlisting).
Strengthened scanner on pre-conversion main: 43 violations
(39 pure-read + 4 no-proven-write). This branch: zero.
Suites: gate 2/2; tests/test_hermes_state.py + tests/state/ 331
passed (same 3 pre-existing FTS-rebuild reds as clean main); 433
passed across the 12 consumer suites of the five newly-converted
methods (compression anti-thrash, session search, status, scheduler).
Pattern-C architectural fix (write-lock contention), completing what
#90734 started: that PR fixed the four UNLOCKED readers racing the
writer connection; this one fixes the 39 LOCKED pure readers convoying
every concurrent turn's persistence behind the global writer lock, and
adds the gate that stops the class from re-entering.
The gateway shares ONE SessionDB across every agent. A read-only query
under `with self._lock:` blocks all concurrent writers for its
duration; under WAL, _read_ctx() serves the same read from a pooled
read-only connection with no lock at all (non-WAL falls back to the
locked writer byte-for-byte, so DELETE-journal installs are unchanged).
Converted (SELECT-only bodies, mechanical `with self._lock:` →
`with self._read_ctx() as conn:`): gateway routing loaders, session/
message counters, titles, compression tip/lineage/cooldown readers,
telegram topic bindings, prune candidate scans, meta readers, resume
resolution — 39 methods, verified per-method that every statement is a
SELECT and every conn use stays inside the with-block. Read-modify-
write methods and checkpoint/maintenance PRAGMAs stay on the writer
lock (their read is ordered against their own write).
Gate: tests/state/test_no_locked_readers_gate.py — an AST scanner over
SessionDB that fails CI when a pure-read method body takes the writer
lock, with a sabotage self-check proving the scanner fires. On
pre-conversion main it reports 38 violations; on this branch, zero.
Measured (2 writer threads + 3 reader threads, 3s, WAL):
before: 78.5k reads, p99 2.49ms, max 160.7ms (readers convoy)
after: 200.4k reads, p99 0.53ms, max 25.5ms (2.6x throughput,
4.7x better p99, 6x better tail; writes unchanged)
Honest caveat: on runtimes where WAL is refused (the currently-bundled
SQLite 3.46 trips the WAL-reset-vulnerability gate → journal=DELETE),
_read_ctx() falls back to the identical locked-writer path and this
change is behavior-neutral by construction; the win applies to WAL
installs (legacy WAL DBs, fixed runtimes, wal-configured operators).
Suites: tests/test_hermes_state.py 243 passed; combined state sweep
540 passed — the only 3 reds are the pre-existing
test_fts_runtime_rebuild failures, verified failing on clean
origin/main before this change.
`/handoff <platform>` never completes on a multiplexed gateway, and when it
does complete it can deliver through the wrong profile's bot. Three distinct
faults, all the same family: multi-profile code paths that assume a single
store / a single adapter map.
1. The watcher polls only the ROOT store.
`_handoff_watcher` resolves `self._session_db` with no profile scope, which
always yields the root `state.db`. But `/handoff` run under
`hermes -p <profile>` writes `handoff_state='pending'` into THAT profile's
store. Nothing ever reads it, so the CLI times out after 60s while the
gateway is alive and connected. The watcher now iterates
`[(None, None), *secondary_profiles]` and polls each inside
`_profile_runtime_scope`.
2. The destination session key is built without the profile namespace.
`_process_handoff` called `build_session_key()` with no `profile=`,
producing `agent:main:...` while that profile's own adapter routes organic
inbound messages on `agent:<profile>:...`. The handoff bound a key nobody
reads.
3. Delivery uses the PRIMARY profile's adapter and config.
`self.adapters` holds only the default profile's adapters (secondaries live
in `self._profile_adapters[name]`) and `self.config` only the default's home
channel. A secondary profile's handoff was therefore sent by the wrong bot,
to the wrong chat, while persisting the right session key and reporting
`handoff_state='completed'` — a false positive that looks correct in the
database and is wrong on the wire.
Two robustness fixes in the same path:
4. Head-of-line blocking between profiles. `_process_handoff` runs a full agent
turn plus delivery; awaiting it inline meant one slow handoff stopped the
watcher from even polling the other profiles. With the CLI's 60s deadline, a
valid handoff could time out purely because another profile's was ahead of
it. Dispatch is now fire-and-forget, with an in-flight guard so a row is
never claimed twice, and a bounded drain on shutdown.
5. Rows stranded in `running`. Only the watcher sets `running`, for the span of
one in-process dispatch, so a row still in that state at startup belongs to
a gateway that died mid-dispatch. It can never reach a terminal state, and
`request_handoff` refuses new requests unless the state is
NULL/completed/failed — that session could never hand off again, silently.
`reclaim_stale_running_handoffs()` now fails those rows once per store at
watcher startup. Failing (not re-queueing) is deliberate: the dead gateway
may already have switched the session key and dispatched, so a blind retry
risks double delivery.
Behaviour on single-profile installs is unchanged: the scope list degrades to
the unscoped root poll, `_resolve_profile_for_key` returns None when
multiplexing is off (byte-identical keys), and config/adapters fall back to
`self.config`/`self.adapters`.
Tests: 13 new across three files. Each was verified to FAIL against the
unpatched code (the fix was reverted and the suite re-run) so they are real
guards rather than decorative assertions. Verified end-to-end on a live
4-profile gateway: `handoff_state` goes failed -> completed, and a planted
stranded `running` row is reclaimed at startup with the reason recorded.
The write-durability contract test (test_repair_path_has_no_bare_connects)
enforces that every repair-path connection carries the macOS write
barriers; the restore rewrites the file header, so it belongs under the
same rule.
apply_durability_barriers() is the guest-connection entry point for
secondary state.db users (async delegation ledger) that must not run
journal-mode setup. database.synchronous (#90892) rides on the
journal-mode/pragma paths guests now skip, so apply it here directly —
otherwise guest connections silently run at the compile-time default.
Follow-up to the #93012 salvage.
Corruption can drop the WAL bit from the database header, and every
repair strategy rebuilds or rewrites the file in place — so a repaired
store comes back in the default journal mode (delete). The WAL-reset
gate at open time never sees the flip because it happens inside the
repair path, not at open (the open-time flip #89393 warns about is a
different door), leaving the operator no signal that a WAL store moved
to DELETE (#89674).
After a successful repair, re-apply the canonical database.journal_mode
through _set_journal_mode_no_wait (concurrent openers abort the flip
instead of sneaking it between transactions) and log a WARNING naming
the old and restored modes. Best-effort: a refused restore is logged,
never raised — the repair itself already succeeded. Probing the damaged
file for a pre-repair mode is deliberately not trusted: the corruption
itself is what drops the WAL bit, so the configured setting is the only
reliable target.
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.
Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.
Credit: @Finn763
`apply_database_pragmas()` reads five sizing pragmas from `database:` and no
durability one, and `_enforce_macos_synchronous_full()` returns early when
`sys.platform != "darwin"`. Between them, nothing in the process ever executes
`PRAGMA synchronous` against state.db on Linux or Windows.
The effective level there is therefore `SQLITE_DEFAULT_WAL_SYNCHRONOUS`, a
compile-time constant of whichever SQLite the interpreter links. Debian and
Ubuntu builds commonly ship it as NORMAL; the bundled build and a plain
source build use FULL. So the durability of state.db is decided by which
python3 the installer found, is invisible from config, and cannot be pinned.
#90837 is three weeks of corruption forensics conducted on Ubuntu under the
stated premise `synchronous=FULL`, with every other cause eliminated live.
That premise is not something the reporter could have verified from config,
because there was no config key to set and no log line to read back.
- `resolve_synchronous_level()` maps the spellings operators actually write
(OFF/NORMAL/FULL/EXTRA, any case, or 0-3) to the PRAGMA integer, and
returns None for anything else. Kept out of the sizing loop on purpose: an
unrecognised `cache_size` harmlessly falls back to a default, an
unrecognised durability level must not.
- `_apply_synchronous_pragma()` applies it, and on Darwin refuses to lower
below FULL. `_enforce_macos_synchronous_full()` runs during
`apply_wal_with_fallback()`, which is earlier than `apply_database_pragmas()`,
so without an explicit floor a configured NORMAL would silently undo #64355
by the accident of running last. Raising to EXTRA on macOS is allowed.
- Unset changes nothing, so no existing install moves.
Tests: 34, covering the parser, application, the unset path, the typo path,
a guardrail that #77630's five keys still apply, and the Darwin floor in both
directions. Removing the wiring fails 6; removing the floor alone fails 1.
Related to #90837
apply_wal_with_fallback treats an on-disk WAL database as authoritative and
says so twice: it never live-downgrades one. The mirror case had no
protection at all. When the on-disk mode is DELETE and the configured mode
is wal, the function flips the database and logs nothing.
journal_mode is a property of the FILE, so that rewrites the header and
persists after the process exits. Setting the mode directly on the file is
something operators do; it was the documented mitigation for the SQLite
3.50.4 WAL-reset bug. A config key that makes the choice durable already
exists (database.journal_mode, #68545), but nothing named it at the moment
the PRAGMA was being undone.
#89293 reports the cost: after upgrading past the vulnerable SQLite,
is_sqlite_wal_reset_vulnerable() stopped short-circuiting into
_apply_delete_for_wal_reset_bug, the flip path went live, and 4 of 5
databases silently returned to WAL with no log line anywhere.
Add a deduped WARNING at both points where the switch succeeds, decided
before the pragma runs since both inputs are only readable while the file is
still in its original state. Log-only: the flip still happens, the return
value is unchanged, and the never-live-downgrade rule is untouched.
WARNING rather than ERROR is deliberate. The reverse direction is ERROR
because dropping to DELETE costs concurrency; this direction is normally the
desirable one (managed_uv treats a database stuck on DELETE as a bug worth
repairing on update). The problem was never the change, it was that the
change was invisible.
The page_count guard is the load-bearing half. A brand-new database also
reports journal_mode=delete and is also about to be switched to WAL, and
every opener applies WAL before creating schema, so without it the warning
would fire on the first run of every install.
Refs #89293
When database.journal_mode=delete is configured but the on-disk DB is already
WAL, apply_wal_with_fallback honors the never-live-downgrade rule and keeps WAL.
That is correct (a live downgrade under open connections causes mixed-mode
corruption), but the operator's configured mode silently has no effect, and on
a WAL-incompatible filesystem (virtiofs/NFS/SMB) the DB then corrupts on the
next crash/sleep exactly what they configured to prevent.
Two code paths return WAL in this situation; both now emit a once-per-process-
per-db_label ERROR telling the operator the config did not apply and they must
convert the DB header offline (stop connections, PRAGMA journal_mode=DELETE):
1. The WAL-reset-vulnerable path (_apply_delete_for_wal_reset_bug): previously
warned only about the vulnerability with an "upgrade SQLite" remedy, which
does not help when the real cause is the filesystem. Emitted after that
warning so the actionable message is last.
2. The read-only probe path (non-vulnerable runtime): previously returned WAL
with no signal at all.
The never-live-downgrade behavior is unchanged (existing test now also asserts
the warning). New tests cover both paths, the per-db_label dedup, and the
require_wal=True edge case.
Real-world impact: a Hermes deployment with state.db on a Podman virtiofs
bind-mount (or any NFS/SMB home) that upgrades across a version where WAL was
the default, then sets journal_mode=delete, sees no corruption protection
until the DB header is converted. This makes the gap visible. See #68545.
Deduplicate the explanation across the constant comment, count_empty_sessions
docstring, and delete_empty_sessions docstring. Keep the incident refs
(#70516/#80763/#82756/#95868) and the core WHY on the constant; cross-reference
from the methods.
Precision pass on the comments added by the previous commit: prompt.submit
reaches replace_messages(archive_dropped=True) with an empty prefix on a
confirmed ordinal-0 rewind, which is the production shape that lands a
populated session on message_count = 0. archive_and_compact normally
publishes at least a summary row, so it is pinned as defense in depth
rather than claimed as an equally reachable trigger.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
`count_empty_sessions` / `delete_empty_sessions` — the dashboard's
"Delete empty (N)" affordance — defined "empty" as `sessions.message_count
= 0`. That column is a denormalized counter over the LIVE (`active = 1`)
rows only, and two production transcript-rewrite paths reset it on purpose
while keeping every dropped turn on disk as `active = 0`:
* `replace_messages(..., archive_dropped=True)` — the rewind / edit /
regenerate mode added in #82756 so a taken-back turn stays recoverable.
* `archive_and_compact` — in-place compaction, which archives the
pre-compaction transcript under the same session id (#38763).
A chat rewound to its first turn, or compacted with an empty live set,
therefore reports `message_count = 0` while still holding its entire
history — and those soft-archived rows are the only copy. A gateway reload
is what makes the row eligible: it stamps `ended_at` on every detached
session (`end_reason='ws_orphan_reap'`), satisfying the sweep's
`ended_at IS NOT NULL` gate. The next sweep then hard-deleted the session
row AND `DELETE FROM messages`, destroying the transcript silently.
Every other emptiness test in `hermes_state` already defends the counter
with a real `EXISTS (SELECT 1 FROM messages ...)` probe
(`delete_session_if_empty`, `prune_empty_ghost_sessions`,
`list_never_active_keyed_sessions`, `find_recoverable_session`). This
sweep was the only destructive path that trusted the counter alone. It now
uses the same probe, via one `_EMPTY_SESSION_WHERE` selector shared by the
count and the delete so the button's N and the sweep it triggers can never
disagree again. The counter stays as a cheap prefilter; `EXISTS` is the
authority.
Genuinely message-less rows are still swept — the feature is unchanged for
the case it was built for.
Fixes#95868
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011zTDHnWcBvQsJSX1fBLhuv
POST /api/sessions/owner-backfill stamps a store's own serving-profile
identity onto its pre-#95407 'profile_name = NULL' session rows. Single
match by construction (each profile's state.db belongs to exactly one
profile), idempotent, one-shot-per-row, never overwrites a non-NULL
owner, and reports the stamped count for logging.
Refs #94724
Bot Mode resolves the forever-chat by exact-title lookup on
(profile, 'Bot Chat'); no session-id pointer exists. A user rename
therefore orphaned the whole conversation: resolution missed, the next
click minted an empty replacement, and UNIQUE(title) then blocked ever
renaming back. Refuse the rename at SessionDB._set_session_title — the
single write path every surface funnels through (gateway session.title,
/title, CLI rename, REST). Hidden discriminates the registry row, so a
normal visible session a user happens to call 'Bot Chat' stays freely
renameable; re-asserting the same canonical title stays a no-op.
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
Close session rows left ended_at IS NULL when the in-process websocket
orphan timer dies with the process (#65194). Dual-clock staleness
(started_at AND newest message), desktop included, live in-memory
sessions excluded, scheduled once from both entry.main and the WS
sidecar so desktop/dashboard boots also run the sweep.
Reproduce the schema-btree failure where the in-place writable_schema/VACUUM ladder can reduce a 3,048-page canonical state.db to 113 pages and still return repaired=False.
Move all mutating strategies behind a complete SQLite online-backup snapshot, retain one exclusive SQLite guard from staging through transactional promotion, preserve committed WAL frames and the live inode, fail closed on environmental hazards, and add adversarial regression coverage for failed-repair preservation, post-stage writer races, interrupted copies, stale scratch, disk admission, attempt-ledger semantics, and durability routing.
Fixes#93064
Supersedes the delivery mechanics of #87409 while preserving its implementation provenance.
Co-authored-by: cervantesh <11169707+cervantesh@users.noreply.github.com>
- Clear the accidental end stamp on resurrection (at the lineage tip):
a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
archive auto-resurrect on the next lookup — the user could never retire
the canonical chat. Test pins the resurrect -> deliberate-archive ->
stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
compressed lineage carries end_reason='compression', so tip-stamped
accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
dm resolution) filtered archived rows out via list_sessions_rich and
still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
Review-pass follow-up on the load-time durability stamp:
- hermes_state.py: import the marker from agent.context_compressor instead
of a third synced literal (hermes_state already imports agent.* at module
level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
fill-empty-tail pop site with the shared constant (was outside the drift
guard).
- agent/conversation_compression.py: the no-op progress check now falls back
to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
rows are stamped at materialization time while compress() output is
marker-swept, so a semantically-identical no-op copy on a cold-resumed
session would previously compare unequal and take the progress branch.
Raw == still runs first so engine-returned list subclasses keep their
__eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
assertions; new test_noop_progress_check_is_marker_insensitive
(mutation-checked: fails when the helper is neutered).
Resumed sessions loaded message dicts from state.db WITHOUT the
_DB_PERSISTED_MARKER, so any flush that lost the identity boundary
(compression durable-snapshot adoption, incremental tool-call persists,
rotation preflight on cold resume) re-appended the ENTIRE loaded
transcript as new rows. Compression cycles then doubled the copies:
the incident session grew 998 -> 1995 -> 3990 -> 7981 rows across
three aborted rotations (15,962 active rows, only 472 distinct).
Fix at the architectural chokepoint: SessionDB._rows_to_conversation
(shared by get_messages_as_conversation and get_resume_conversations)
now stamps the marker at row materialization time - a dict built FROM
a durable row is persisted by construction, regardless of which caller
loads it or how the list is later handed to a flush.
Safety:
- Wire-safe: every transport strips underscore-prefixed keys before
the API request (chat_completion_helpers, anthropic_adapter), same
contract as the existing _row_id stamp in the same function.
- Rotation handoffs still write: compression's assembly copies strip
the marker (_fresh_compaction_message_copy + the terminal
_strip_persistence_markers sweep), so compacted transcripts still
flush to the child session (#57491 invariant preserved).
- Branch/seed copies unaffected: /branch and _persist_branch_seed
build fresh field-projected dicts and write via append_messages_batch
directly, not through the marker-gated flush.
Tests: new regression suite (marker sync, load stamping, 3-cycle
amplification repro, new-tail write guard, compaction-copy handoff);
updated the #68454 control test that asserted the old double-write
behavior and the ACP restore shape test.
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.
Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.
Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
Addresses two data-integrity gaps @andrexibiza flagged reviewing #88425.
1. Forensic dedupe no longer reuses the repair-epoch fingerprint.
_db_fingerprint masks SQLite's commit counters and samples only head/tail
so an ordinary write does not re-key the repair budget — the right
predicate for 'same damage epoch', the WRONG one for 'same recovery
image'. A live writer committing rows into an interior page (size
preserved, head/tail untouched) collided under it, so _backup_db_file
handed back a STALE backup that predates real user data. New
_backup_content_identity() digests the whole file + every sidecar; the
dedupe uses it. The O(n) read is cheaper than the O(n) copy it avoids on a
hit.
2. Backup bundle is now published atomically. The promotion loop replaced
files one at a time (main first) and cleanup unlinked only staging srcs,
so a sidecar os.replace failure after the main promotion left the
final-prefix main backup on disk — a countable-but-incomplete bundle that
passed the #69603 hard stop and deduped as legitimate next pass. Now
sidecars publish first and the main DB last (its name is the commit
marker _existing_malformed_backups counts), and cleanup rolls back every
already-published destination.
Two regressions added (both mutation-checked — each fails on pre-fix code):
- test_backup_not_deduped_after_interior_page_write
- test_publication_failure_leaves_no_countable_partial_bundle
tests/test_state_db_repair_loop_mtime.py: 28 passed.
The pre-repair copy took only -wal/-shm. In rollback-journal (DELETE) mode --
Hermes's fallback on NFS/SMB/FUSE/ZFS and on WAL-reset-vulnerable SQLite builds
-- a hot <db>-journal exists on disk whenever a transaction was open, and that
file is what rolls the damaged bytes back to a consistent state. A forensic copy
without it cannot be recovered by hand, which is the entire purpose of taking
the copy before destructive surgery.
Verified the journal is really there:
files while a txn is open: ['state.db', 'state.db-journal']
files after commit: ['state.db']
Add _DB_SIDECAR_SUFFIXES = ("-wal", "-shm", "-journal") and use it at the four
sites that must agree: the disk-guard sizing, the staging copy, the
backup-count exclusion in _existing_malformed_backups (so a copied journal is
not itself counted as a forensic backup), and _prune_malformed_backups (which
otherwise leaks one journal per pruned backup, quietly defeating the retention
cap this PR is partly about).
Matches the spelling hermes_cli/session_recovery.py:61 already uses for the
same concept.
Third self-review pass found the content fingerprint was still defeated on
rollback-journal deployments, by the same mechanism as the original mtime bug.
The head sample starts at byte 0, so it covers the database header's file
change counter (bytes 24-27) and version-valid-for (92-95). In DELETE mode a
commit writes the main file directly and bumps both. A malformed-SCHEMA DB
still accepts writes -- that is the whole premise of this PR -- so any ordinary
session write between passes re-keyed the ledger:
DELETE, 18MB db, one peer UPDATE between passes (before this commit)
pass 1..6: attempts=1 every pass, exhausted=False -> unbounded loop
after
pass 1..3: attempts=1,2,3 pass 4: BLOCKED
WAL is unaffected (commits land in -wal; the main header only moves on
checkpoint), so this was invisible on a WAL host and reproducible on every
NFS/SMB/FUSE/ZFS or WAL-reset-vulnerable host -- exactly the deployments the
earlier lock-safety commit was written for.
Mask the two volatile ranges out of the sample. Page 1's sqlite_master b-tree
sits after byte 100 and stays in, so genuine recovery still resets the budget:
verified schema rewrite, index rebuild, VACUUM and truncation all change the
key, while a bare utime and an ordinary commit do not.
Test-cost cleanup in the same file, since the new tests needed a
larger-than-sample fixture and the file was already slow:
- the two guard tests that allocated 450MB of os.urandom now use sparse
truncate (both only ever read st_size), and the new fixtures use 600 rows
rather than 40k;
- file runtime 127s -> 35s.
Self-review of the previous commit found it reintroduced the bug this PR
exists to fix, by a different route.
`_db_fingerprint` fell back to `size:mtime_ns` when a live connection made the
content read unsafe. The ledger compares keys for EQUALITY, and the two keys
have different SHAPES, so a gateway peer connecting between passes flipped the
shape and the counter reset to 1 every time:
pass 1 [offline] attempts=1 fp=8192:58c7924f0fba...
pass 2 [LIVE ] attempts=1 fp=8192:1786972039271402096
pass 3 [offline] attempts=1 fp=8192:58c7924f0fba...
... never reaches _MAX_PERSISTENT_REPAIR_ATTEMPTS
Return None instead, and teach the two ledger helpers to cope:
- `_persistent_repair_attempts_exhausted` falls back to the recorded key's
SIZE prefix (the one component both shapes share and that needs no raw
read) rather than reading as "not exhausted" — otherwise a peer connection
hides an exhausted budget on every pass, same loop.
- `_record_repair_outcome` keeps the key already on record and still
increments, rather than dropping the pass.
pass 1 [offline] attempts=1 pass 2 [LIVE] attempts=2
pass 3 [offline] attempts=3 pass 4 [LIVE] BLOCKED
Intra-pass flips were already safe (the probe and the record are both reached
with the same liveness within one `repair_state_db_schema` call); it is the
cross-pass change that desynced.
Also drops two `type: ignore` directives `ty` flagged as unused, and replaces
the `LiveConnectionError = ()` / `nullcontext()` shim with a real no-op
contextmanager + exception class so the scaffold-install path is honest.
The staging name was derived from the backup name
(`<db>.malformed-backup-<stamp>.incomplete`), which still matches the prefix
`_existing_malformed_backups` selects on -- it excludes only `-wal`/`-shm`.
Three consequences, all reproduced:
- it is COUNTED as a forensic backup;
- it sorts NEWEST (`.incomplete` > the bare stamp), so prune's
keep-3-newest slice retained partials and deleted intact copies -- the
exact inversion the staging change was meant to prevent;
- worst, the dedupe ran BEFORE the sweep, and a staging file orphaned by a
kill mid-copy is a byte-identical copy of the damaged DB, so its
fingerprint MATCHES and it was handed back as the official `backup_path`.
Repair then passed the #69603 hard-stop gate and ran destructive surgery
believing a forensic copy existed, and the next pass's sweep deleted that
very file.
Move staging outside the prefix (`<db>.backup-staging-<stamp>`) and sweep
before the dedupe. The sweep also matches the pre-merge `.incomplete`
spelling so a host that ran the earlier build does not keep prefix-matching
debris that sorts newest and survives prune forever.
Before / after on the same fixture (orphaned staging + a later pass):
before backup_path = ...malformed-backup-<stamp>.incomplete (staging!)
pass-1 forensic copy deleted by the next sweep
after backup_path = ...malformed-backup-<stamp> (real copy)
debris swept, pass-1 forensic copy preserved
The content fingerprint takes a raw descriptor, and close() on ANY descriptor
cancels every POSIX advisory lock the process holds on that file. The
exhaustion probe runs before _backup_db_file's has_live_connection guard, so
the read happened even when a peer SessionDB held a write lock.
Verified end-to-end (journal_mode=DELETE, gateway mid-turn write, peer in a
subprocess):
before peer BLOCKED -> repair -> peer BLOCKED, holder COMMIT ok
unfixed peer BLOCKED -> repair -> peer STOLE the lock,
holder COMMIT: disk I/O error
WAL is immune (it coordinates through -shm), but DELETE is what Hermes falls
back to on NFS/SMB/FUSE/ZFS and on SQLite builds vulnerable to the WAL-reset
bug, so this is a real deployment shape.
Run the read under offline_file_access and fall back to size:mtime_ns when a
connection is live. That keeps the ledger counting instead of returning None
(which reads as "not exhausted" and would restore the unbounded loop), and the
content key stays load-bearing on the offline repair path -- the only path
where surgery actually runs.
Also fail the free-space guard CLOSED: a nearly-full volume is exactly where
statvfs is likeliest to fail, and proceeding is the multi-GB copy that finishes
off the disk.
Follow-up to adversarial review of the first commit. Three findings, two
confirmed by test and fixed here, one disproven and left alone.
CONFIRMED — the free-space guard was a threshold, not cleanup. Prune runs
only on the success path, so any copy that failed partway (ENOSPC, sidecar
copy failure, kill mid-copy) left a file matching the `malformed-backup-`
prefix that nothing ever removed. Measured on the unpatched tree: backups
capped at 3 while copies succeed, but 13+ and climbing once copy2 raises —
self-reinforcing, since each partial consumes the space that guarantees the
next failure. Worse, partials sort newest-by-name, so a later successful
prune KEPT the garbage and deleted the intact forensic copies.
Fix: copy to a `.incomplete` staging name that does not match the backup
prefix, os.replace into place only after every copy succeeds, unlink staging
on failure, and sweep stale staging debris on entry.
CONFIRMED — the 2GiB floor was a small-volume regression. A 50MB DB on a
10GB volume with 1.5GB free (30x headroom) was refused, and since a refused
backup is a HARD STOP (#69603) that silently converts "repair loops" into
"repair never runs". Fix: require the copy itself (now including its
-wal/-shm sidecars, which the old check ignored) plus proportional headroom
— max(256MiB, 2% of volume).
DISPROVEN — the review claimed a refused backup skips _record_repair_outcome
so the loop never terminates. It does not: repair_state_db_schema records the
outcome on the result returned by _repair_state_db_schema_locked, which is
where the hard stop returns. Verified on a simulated low-disk host: terminal
at pass 4 with zero backups written. No change made.
Tests: 5 new (small-volume allow, proportional headroom, sidecar accounting,
failed-copy leaves no countable debris + staging swept). 23 pass with the
#86747 suite; test_hermes_state.py 252 passed. Pre-existing unrelated
failures unchanged.
A malformed-schema state.db sent Hermes into a repair loop that wrote a
fresh full-size forensic backup every ~10s: 31 copies / 2.3GB in 20
minutes, free space heading to zero on a host running an agent fleet.
The #86747 guards for exactly this were already present and did not hold.
Both keyed on `size:mtime_ns`:
* `_db_fingerprint` -> the ledger's attempt counter reset to 1 on every
pass, so `_MAX_PERSISTENT_REPAIR_ATTEMPTS` was never reached and the
loop never terminated;
* `_backup_db_file`'s dedupe compared mtime, so it never matched and
each pass wrote another full-size copy.
The assumption behind that key -- "nothing can successfully write to a
damaged file" -- holds for the b-tree damage of #86747 but not for the
malformed-SCHEMA class: the DB still opens and accepts writes (only
sqlite_master is unreadable), so live writers, WAL checkpoints and the
in-place repair strategies themselves all move mtime between passes.
Fixes:
* fingerprint on size + a bounded head/tail content sample instead of
mtime. Stable across passes that merely touch the file, still changes
on genuine repair/truncation/restore (so recovery resets the budget),
and stays O(1) on a multi-GB DB.
* dedupe the forensic backup on that same fingerprint.
* add the missing free-space guard: refuse the pre-repair copy when it
would leave under 2GiB free, with an actionable error. The backup is a
full raw copy of the damaged DB, so a repair loop is a disk amplifier
that can take down every process on the host -- and the refusal path
already hard-stops the repair (#69603) rather than mutating the only
remaining copy.
Tests fail on the unfixed tree and pass here; the pre-existing failures in
test_state_db_malformed_repair.py and TestFTS5Search are unrelated and
reproduce on the base commit.
Follow-up to the salvaged repair-durability commit. Scope corrections so this
PR ships only the reachable, non-competing, WAL-mode-correct half:
- Drop verify_state_db_integrity() + its 4 tests. Zero production callers here
(dead code); the caller lives in the follow-up that wires it into
SessionStore._open_session_db_for_active_scope() (PR #91754). The function
moves with its wiring.
- Drop the _db_fingerprint change (size:mtime_ns -> dev:ino:size) + its 3
ledger tests. This is competing work: PR #88425 (salvage of @jirathip-k's
#88224) already fixes the same size:mtime_ns budget-reset bug with a
content-sample + volatile-header-mask that also handles the DELETE-mode
commit-counter case, and carries @jirathip-k's diagnosis/credit. Landing a
second, divergent fingerprint contract would stomp that lineage. Fingerprint
stays with #88425; this PR reverts _db_fingerprint to main's form.
- Mark test_repair_refuses_while_another_connection_holds_the_db requires_wal.
_live_writer_holds_db detects an out-of-process holder via the WAL-index
exclusive lock, absent in journal_mode=DELETE (used on WAL-reset-vulnerable
SQLite <3.51.3 incl. CI's 3.50.4, and on NFS/SMB). The test failed there;
the conftest requires_wal gate auto-skips it. DELETE-mode limitation is now
documented on the guard docstring: repair is serialised only by the
cross-process repairer lock there. The reported incident was in WAL mode.
- Map dhanesh@users.noreply.github.com -> dhanesh (contributors/emails) so the
attribution CI gate passes.
Net: this PR is repair-connection durability barriers + the live-writer guard.
addresses @andrexibiza's #90747 review (dead-code verifier + fingerprint
interlock with #88425).
state.db corrupted twice in two days with the torn-b-tree signature —
repeated "2nd reference to page", "Rowid out of order", and long runs of
"never used" pages in messages (rootpage 5) and idx_messages_session.
macOS fsync() guarantees neither data-on-platter nor write ordering, which
_enforce_macos_synchronous_full already documents: a rewrite interrupted by
process or OS termination leaves half-written b-tree pages. The mitigation
is per-connection (synchronous=FULL + checkpoint_fullfsync=1) and was
applied only through apply_wal_with_fallback(). The repair path opened
state.db with a bare sqlite3.connect() six times and then ran REINDEX,
VACUUM and writable_schema surgery through it — the operations that rewrite
nearly every page of the file — with no barrier at all.
- _connect_repair_durable() routes every repair/probe connection through the
barriers. Applying them is best-effort by necessity: SQLite loads the
schema before any statement, so on a malformed schema even
PRAGMA synchronous=FULL raises DatabaseError, and a malformed database is
precisely this helper's input. _reapply_durability_barriers() retakes them
before REINDEX and VACUUM, once the schema parses and they can stick.
- verify_state_db_integrity() adds the proactive check that was missing.
Repair only ever ran reactively, after a caller already hit a malformed
error, so a database torn in pages no query happened to touch stayed live
and kept accepting writes. On 2026-08-19 that gap was 11 hours across two
restarts that both reported a clean start. Size-aware: degrades to an O(1)
probe above 2 GiB rather than pegging a CPU at startup.
Also restores two fixes lost when `hermes update` reset the tree to
origin/main before they were committed:
- _db_fingerprint keys the repair ledger on dev+inode+size instead of
size+mtime_ns. The old form was justified as "stable for a file nothing
can successfully write to"; that premise is false, because on FTS
corruption this module deliberately keeps canonical writes enabled with
FTS detached. mtime churned on every write, so each pass re-keyed the
ledger and reset the counter to 1 — the cap could never be reached and the
damaging surgery could retry forever.
- _live_writer_holds_db() refuses surgery while another connection holds the
database. The cross-process lock only serialises repairers against each
other; it says nothing about the gateway, Desktop or a CLI. Rewriting
b-tree pages under a concurrent writer is what spread the 2026-08-18/19
damage out of the FTS shadow tables and into the canonical ones. Fails
open, so it cannot strand the self-heal path it protects.
The guard's own tests built a two-table toy schema, so every repair aborted
on "no such table: sessions" before reaching the guards under test — the
assertions were passing over a code path that never ran. They now build
through a real SessionDB.
Targeted state/repair suites: 330 passed, 1 pre-existing unrelated failure.
Broader sweep: 50 failed/1221 passed -> 46 failed/1225 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Dhanesh Purohit <dhanesh@users.noreply.github.com>
The cmdline fallback was matching every system daemon with an
unreadable fd table (init, systemd-journald, dockerd, etc.), causing
FTS rebuilds to be skipped on every Linux system. Add _looks_like_hermes
filter so only processes whose cmdline contains Hermes markers are
flagged — matching @jackulau's suggestion of 'uninspectable AND
identifiable as another Hermes process.'
Address review feedback from @jackulau on PR #90871:
1. psutil.open_files() silently drops '(deleted)' WAL sidecar entries
on Linux because isfile_strict() stats the literal path including
the suffix and fails. Switch to direct /proc/<pid>/fd readlinks
which preserve the '(deleted)' suffix so _canonical can match.
2. psutil.process_iter() converts AccessDenied to None, which
or-() skips silently — the fail-closed branch never runs. For the
root-gateway vs user-desktop topology in the issue, the fd table is
unreadable but /proc/<pid>/cmdline is world-readable. Add a cmdline
fallback that flags uninspectable processes.
Also keep the psutil path for macOS/BSD (no '(deleted)' convention).
Add the foreign-holder guard to gateway/session.py::_rebuild_fts_once(),
the third FTS rebuild path that was not covered by the original fix.
Also add a comment explaining why _fts_runtime_rebuild_attempted is set
before the foreign-holder check: the fail-open path that follows
persists FTS_STALE_KEY so the next startup retries via _recover_stale_fts.