classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.
One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".
Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).
Fixes#97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
`sqlite3 state.db .recover` re-emits the FTS5 shadow tables (messages_fts_data,
_idx, _docsize, _config, _content) as ordinary tables but cannot re-emit the
CREATE VIRTUAL TABLE row. The next SessionDB open ran the FTS DDL in
_ensure_fts_schema and died with "fts5: error creating shadow table
messages_fts_data: table 'messages_fts_data' already exists", so a recovered
database was unusable until someone hand-dropped the shadows.
_init_fts now runs _drop_orphan_fts_shadow_tables before any FTS DDL. It is
per-family and exact-name scoped: a family's shadows are dropped only when its
own vtable row is absent from sqlite_master (type='table' AND sql LIKE
'CREATE VIRTUAL TABLE%'), so a healthy messages_fts_trigram survives a
base-family repair untouched. A repaired base/trigram family is then treated
like a missing-trigger repair and rebuilt from the canonical messages table
under the cross-process rebuild admission; the shadows are derived index
state, nothing is lost.
Live repro: real `sqlite3 x.db .recover | sqlite3 y.db` on sqlite 3.50.4
keeps the vtable rows (the shell emits CREATE VIRTUAL TABLE), so the
deterministic fixture removes the vtable row via writable_schema leaving the
shadows behind: BEFORE OperationalError on open; AFTER opens, fts_enabled,
MATCH returns every message, trigram sqlite_master rowids unchanged.
Salvaged from PR #56824 (intent applied onto the current hermes_state_fts /
hermes_state_schema siblings). The ownership-safety point (never touch a live
family's shadows) was raised by @ggoldani in #103868 / #103840.
Refs #103840
Refs #56815
Co-authored-by: ggoldani <ggoldani@users.noreply.github.com>
Keep the two invariant tests from #105428: a genuine second instance
whose argv is scoped to another HERMES_HOME is not a holder of our
state.db; argv that names our state.db stays flagged. The helper-shape
tests were change-detectors on `_argv_scoped_to_other_home` internals.
The salvaged commit is re-authored to TaoMasterCoder's GitHub noreply
identity: the original `devops@77hub.com` is a shared org address
(misconfigured local git, not malice).
Refs #107440
Field evidence (2026-09-07, production host): the host runs two independent
Hermes instances - a main gateway (user ubuntu, HERMES_HOME=/home/ubuntu/.hermes)
and a demo gateway (user demo, HERMES_HOME=/home/demo/.hermes). Because the
demo process is owned by another user, its /proc/<pid>/fd table is unreadable
and foreign_state_db_holders() falls back to cmdline + _looks_like_hermes().
The demo argv matches Hermes patterns exactly, so it was flagged as an
uninspectable holder of the MAIN instance's state.db even though lsof proved
0 open handles on it. Result: _recover_stale_fts was deferred 42 times across
6 gateway restarts, the fts_stale breadcrumb never cleared, and FTS
self-repair stayed permanently disabled.
#92419 removed substring false positives (journalctl/grep mentioning hermes);
a genuine second instance with a DIFFERENT HERMES_HOME was still misjudged.
Add _argv_scoped_to_other_home(): when the argv of an uninspectable Hermes
process proves it lives under a different /.hermes home (or a state.db
sidecar under a different parent) AND no token references our state.db,
sidecars, or home, do not count it as our holder. Applied to all three
uninspectable branches (holder + two descriptor paths). Ambiguous argv
without absolute-path tokens remains fail-closed, preserving the
conservative intent.
References #92401
_pin_connection took the shared connection via self._conn, so the lexical
lock audit flagged it (then its caller) as unlocked use. Pass the conn in
from the lock-held close paths and whitelist _settle_lost_generation_locked
as a close()-only lifecycle helper.,
Slim redo of the mechanism from #106410 on top of its pick (no wrappers, no
persisted "kind" enum, no process-local flag that dies with the process):
- Futility = the SAME holder PID set has blocked >= _FTS_HOLDER_FUTILE_ATTEMPTS
(10) deferrals over >= _FTS_HOLDER_FUTILE_SECONDS (30 min); tracked as
holders_since/holders_attempts in the persisted fts_rebuild_deferral record
and reset whenever the holder set changes. The 3-deferral/60 s escalate
window is the orphan-reap gate and stays as is.
- ONE escalated ERROR line names each holder pid + cmdline and the remedy that
can actually be followed from inside a gateway session: stop ONLY the other
holder; this process's own retry admits the rebuild within 60 s. The old
"with the gateway stopped" advice was unrunnable from a gateway-hosted
session (the gateway is the session) and is gone from both log and doctor.
- hermes doctor renders the futile record distinctly.
- retry_deferred_fts_recovery: a capped backoff earned by holder set X no
longer applies once the live holder set differs from X, so stopping the
other service is followed by a retry on the next tick, not up to an hour
later (the issue's 16-min wait).
- Tests trimmed from 5 to 2 invariants (futile line + doctor entry after N
same-holder deferrals; backoff reset when the holder set changes); the
contributor's control tests for changing PIDs / orphan reap are covered by
the existing test_repeated_deferrals_reap_inactive_orphan_then_rebuild.
The "canonical writes and LIKE search remain available" WARNING is kept
because it is true on origin/main: a stale open drops every FTS trigger, so
the messages INSERT succeeds (probed live with a real state.db + a second
process holding it). Writes fail only when a peer re-publishes triggers over
the corrupt index — a separate class, not this diagnostic.
Refs #106393
A supervised peer never satisfies the orphan reap, so stale-FTS repair retried forever with a misleading "canonical writes remain available" warning.
(cherry picked from commit e57f3a975d311aa44da1e92c5e727eba7c8cff70)
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
Composition bug between the salvaged #101266 (in-place v29 startup
migration: swap the trigram view/triggers, FTS5 'rebuild') and #88217
(FTS_STORAGE_VERSION 2 drops the tool_calls column from the trigram
vtable, opt-in via optimize-storage). On an install still carrying the v1
vtable, the startup migration replaced the view with one that has no
tool_calls, then 'rebuild' failed with 'no such column: T.tool_calls' and
SessionDB.__init__ raised — reproduced by opening a real main-built DB.
Gate the in-place migration on the vtable not projecting tool_calls; such
installs are already offered optimize-storage, which recreates the vtable
from FTS_TRIGRAM_SQL (cron-filtered view included). E2E: main-built DB ->
opens on this branch, optimize_fts_storage() yields v2 columns and purges
the cron row. Test fixture now builds a real external-content vtable for
both layouts; new test mutation-checked against the missing guard.
Replace the requires_wal-gated unit assertion (which never runs on WAL-reset-
vulnerable runtimes such as macOS 3.46) with an end-to-end test through
repair_state_db_schema that traces every _connect_repair_durable call against
the guard's enter/exit and fails on main's shape (a connect after guard-exit).
Runs on every platform. Docstring now names the actual hazard.
apply_wal_with_fallback keeps a store in DELETE where the linked SQLite has
the WAL-reset bug; the connection-reuse contract the test guards is unaffected
but its final journal_mode assertion is not satisfiable there.
The seed put the retry deadline 900s in the future, so on unguarded code the
method short-circuited on the backoff check and the protection assertions
passed anyway; only the field-reset assertions failed. Seed the deadline in the
past so a rebuild is due, and condense the guard comment.
retry_deferred_fts_recovery() is called unconditionally on every
housekeeping tick for the life of a long-running gateway process
(#100108). It checks _fts_stale, read_only, and _conn is None, but never
_db_corrupt (bcc2e65818, #101095/#101224): a handle that observed
structural corruption is supposed to stop being touched entirely (see
_try_wal_checkpoint's identical guard, and close()'s skip of the
checkpoint), but this method has no such check.
If a handle both has a deferred stale-FTS breadcrumb AND later trips
quarantine (both are plausible on the same corrupted file — the field
incidents motivating the quarantine feature describe corruption touching
FTS shadow tables and canonical btrees together), every subsequent
housekeeping tick runs a real FTS rebuild (DROP TABLE / CREATE VIRTUAL
TABLE / bulk INSERT) against the file the code has explicitly decided to
stop touching — exactly what quarantine exists to prevent.
Fixed by returning False immediately when _db_corrupt is set, mirroring
_try_wal_checkpoint's "quarantined: never touch a damaged image" guard.
The method's own contract ("never raises") is preserved — no
StateDbCorruptError is raised here, this is a quiet skip like the other
corrupt-aware call sites.
Also resets the backoff bookkeeping (_fts_stale_retry_after,
_fts_stale_retry_interval) in the same early-return, mirroring the
success path's own reset a few lines down (review feedback from
Baophan00 on the PR). Verified empirically before making this change:
_db_corrupt is set to False nowhere in the codebase outside __init__, and
the shared registry's file-replace path always constructs a genuinely new
SessionDB instance rather than clearing the flag on a live one — so no
code path today revives a quarantined handle in place, and leaving the
backoff fields untouched is inert in practice. The reset is still cheap,
harmless, and closes a real footgun for whoever adds an un-quarantine
path later: without it, a handle quarantined mid-backoff would carry a
doubled multi-minute interval into any future retry instead of starting
from the default.
Added a regression test that marks a handle stale, forces the open-time
recovery to defer via a real held rebuild lock (so _fts_stale survives
construction), sets _db_corrupt plus a pre-existing multi-minute backoff,
and asserts the retry is a no-op with both backoff fields reset to 0.0.
Mutation-verified: reverting hermes_state_schema.py makes the retry
actually run the rebuild and return True, and separately makes the
backoff-reset assertions fail with the stale pre-quarantine values still
in place.
(cherry picked from commit 3445da1d98d84bb60cb3799ef59e1fa4c619100f)
The salvaged #101266 gated its rebuild on current_version < 27; main had
already reached SCHEMA_VERSION 28 (column-reconciliation bumps), so on any
existing install the gate would never fire and cron rows would stay in the
trigram index forever. The cherry-pick resolution renumbers to v29; this
test seeds a v28 database and asserts the migration runs. Mutation-checked
against the original < 27 gate.
test_no_locked_readers_gate.py (#97676) parses hermes_state.py's SessionDB
class body with ast and flags any method that holds the writer lock
around a pure-read query — Pattern C, where every concurrent turn's
persistence convoys behind an unrelated read. #97676 converted 39 such
methods and closed detection blind spots for alias/variable-SQL readers.
But SessionDB is declared as
`class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)`,
and the gate only ever opened hermes_state.py — it never parsed the three
mixin files those base classes are defined in, so a locked reader
declared there was structurally invisible to the scanner regardless of
how good the alias/variable-SQL detection got.
Applying the gate's exact scanning logic to the three mixin files
directly turns up 9 genuine pure-read methods still holding the writer
lock, none in #97676's converted list:
- hermes_state_search.py: _fts_teardown_trash_step, fts_optimize_available,
optimize_fts_storage, list_recent_user_messages
- hermes_state_portability.py: distinct_session_cwds, list_cron_job_runs,
_get_session_rich_rows_batch, list_skill_scaffolded_sessions,
get_first_assistant_text
_get_session_rich_rows_batch is a hot path: it backs list_sessions_rich's
compression-tip resolution and the web server's session-search hydration
across every gateway install — its own docstring already claims "same
read-your-writes guarantee as list_sessions_rich", but list_sessions_rich
was already using _read_ctx() (its guarantee comes from flush_token_counts()
before the read, not from holding the writer lock) while this method's
implementation never caught up to match.
Converted all 9 to `with self._read_ctx() as conn:`, the exact pattern
#97676 used, verified each is a genuine pure read with no hidden writes
by tracing every helper call it makes.
Extended the gate itself (_ALL_STATE_SOURCES) to scan all three mixin
files under their own class names, plus hermes_state.py, so this blind
spot can't silently reopen. Added a regression test
(test_scan_all_state_sources_visits_every_mixin_file) that plants a
synthetic violation in a mixin-shaped file and asserts the scanner still
catches it — a change that reverts the file list back to one file passes
the existing sabotage test but fails this one.
Mutation-verified: with the gate's new scope but the old (unconverted)
mixin sources, test_no_locked_pure_readers fails and names all 9 real
violations with correct file/line. Restored the fix; it passes clean.
Move live-holder inspection into a bounded helper and make repair fail closed when a process still owns the state database, including deleted WAL/SHM descriptors and ambiguous procfs reads.
Preserve the contributor lineage from the original four-commit review train while presenting one coherent release object on current main.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: ciabata <296402666+ciabata-git@users.noreply.github.com>
A bare SQLITE_CORRUPT/NOTADB on a live write (not FTS-scoped, not a
replaced file) now sets a sticky per-instance flag: later writes fail
fast with StateDbCorruptError, the handle never reopens after close(),
and close() skips its explicit PASSIVE WAL checkpoint. Gateway and agent
flush paths divert pending transcripts to JSONL/spool like the replaced
case instead of retrying forever.
Field evidence: a handle that kept writing for ~50 minutes after the
first structural error checkpointed 15 pages under the wrong page
numbers on shutdown (page 1 <- messages_fts_trigram_data leaf), turning
"malformed" into "file is not a database".
Refs #90837, #90950, #97940, #89332, #45383
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CNX8rNYHqA5pT4tAGSzXtb
Regression coverage for the #100130 salvage, all against real SessionDB
files and a real child process holding the flock:
* errno table for `is_advisory_lock_contention` (EAGAIN/EWOULDBLOCK/EACCES
contend; ESTALE/ENOTSUP/ENOLCK/EIO fail fast); no misleading "held by
another process" line on the fast-fail path; `_cross_process_repair_lock`
shares the filter (sibling site).
* `retry_deferred_fts_recovery`: open under a live holder -> stale; retry
returns in <2s with a 30s admission budget (timeout=0); rate limit +
60s->120s backoff engaged; holder dies -> same instance recovers, triggers
restored, breadcrumb cleared; no-op when not stale / read-only.
* `_start_gateway_housekeeping` tick (real loop, 50ms interval) recovers a
stale shared-registry SessionDB with no direct call and no extra thread.
Backoff floor: a monkeypatched 0s base interval must not zero the doubled
interval (min 1s), so the cap math is testable.
Sabotage run (source at origin/main, these tests): 16 failed / 35 passed,
including 30s timeouts on the fast-fail tests.
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.
Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.
Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).
Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
Drives a real unopenable lock path — a directory where the code expects a
regular file, so open() raises a genuine kernel OSError — rather than
monkeypatching the helpers, standing in for the ENOSPC/EMFILE the field reports
hit without needing to fill a disk.
Both authorities are covered at the primitive and the behavior level:
fts_rebuild_admission refuses admission and rebuild_fts() reports no progress
(asserted against a preceding successful rebuild, so the 0 is the deferral and
not an unrelated no-op); _cross_process_repair_lock refuses the authority and
repair_state_db_schema runs no writable_schema surgery, takes no forensic
backup, and leaves the damaged image byte-identical for the next authorised
pass. A guardrail test pins that a pathless in-memory store is still admitted,
so the fix cannot turn that legitimate no-op into a permanent deferral.
All four deferral assertions fail on the pre-fix code, where the repair test
shows the surgery really did proceed without cross-process authority.
Real byte-flip fixtures (no mocks) proving the #96038/#98090-class fix
end to end, closing the acceptance gate on issue #97940:
- test_canonical_btree_corruption_fails_closed: checkpoint the WAL,
clobber every messages-table B-tree leaf page header, then assert a
live append raises the genuine bare SQLITE_CORRUPT, the classifier
refuses the FTS route, no rebuild/detach/stale-marker side effects
occur, and the field incident's misdiagnosis log line ('canonical
message rows are preserved') never appears.
- test_fts_only_corruption_still_self_heals: contrast case — a real
messages_fts_data shadow-table stomp raises SQLITE_CORRUPT_VTAB (267),
is classified as FTS-scoped, and the write path still self-heals with
canonical rows intact.
Sabotage-verified: reverting the classifier fix (96739033c4) makes the
canonical-corruption test fail by entering the FTS self-heal route.
Credits @fangliquanflq (PR #98090) for the production timeline analysis
and @diatche (PR #96038) for the classifier fix these tests gate on.
Refs #97940, #98077.
Subagent/cron sessions died mid-run with "Session DB append_message
failed: 'NoneType' object has no attribute 'execute'": a teardown owner
(cron run_job finally, delegate timeout owner, agent close()) called
SessionDB.close() — nulling _conn — while a still-unwinding worker had
one more transcript flush to land. The flush then hit None.execute, the
turn force-ended as session_persistence_failed, and the session tail was
silently dropped while cron delivery reported last_status: ok.
Fix at the shared persistence boundary: _execute_write and the _read_ctx
writer-lock fallback detect the closed handle under self._lock and
reopen a connection to the same database file with a loud WARNING naming
the race. Read-only handles never reopen — they raise an explicit
'was closed' error. A failed reopen raises an OperationalError naming
the teardown race so classify_persistence_error gets a real cause.
Closes#94736
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
(including the bare SystemError CPython's sqlite3 layer raises under
cross-thread errmsg scrambling) are contained and logged instead of
escaping and making the caller replay an ambiguous, possibly-durable
write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
hammer + AST sweep freezing the no-unlocked-writer-conn invariant.
On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):
1. publish_compression_child gains require_lease_refresh: the lease is
extended inside the same transaction as the expiry check (same conn, no
TOCTOU), giving a worker whose refresher thread died from transient DB
failures one final chance to keep its completed work.
2. A failed compression split now records a 60s failure cooldown, so the
next turn cannot immediately re-trigger the same compression.
Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
Reviewer findings from the formal /simplify-code pass, all verified:
1. Scanner blind spots (HIGH): five more locked pure readers were
invisible to the v1 gate — get_compression_fallback_streak and
get_compression_ineffective_count hid behind `conn = self._conn`
aliasing; list_gateway_sessions, find_session_by_origin and
search_sessions hid behind SQL held in variables/f-strings (the
scanner required unknown == 0 to flag). All five converted to
_read_ctx(); the scanner now (a) tracks self._conn aliases and
(b) flags lock blocks with NO proven write instead of silently
skipping unprovable SQL. Sabotage self-check extended to pin all
three detection classes (literal, alias, variable-SQL) plus a
mixed variable-SQL writer that must NOT fire.
2. get_meta reverted to the writer lock: its inline comment (present
on main) documents a real read-your-writes dependency —
fts_rebuild_step reads rebuild progress that a pooled WAL reader
cannot see mid-transaction. The blanket conversion had overridden
a documented design decision; it is now the single justified
_ALLOWED_LOCKED_READERS entry, replacing the dead
_enter_fts_fail_open entry (whose lock block counts 3 writes and
never needed allowlisting).
Strengthened scanner on pre-conversion main: 43 violations
(39 pure-read + 4 no-proven-write). This branch: zero.
Suites: gate 2/2; tests/test_hermes_state.py + tests/state/ 331
passed (same 3 pre-existing FTS-rebuild reds as clean main); 433
passed across the 12 consumer suites of the five newly-converted
methods (compression anti-thrash, session search, status, scheduler).
Pattern-C architectural fix (write-lock contention), completing what
#90734 started: that PR fixed the four UNLOCKED readers racing the
writer connection; this one fixes the 39 LOCKED pure readers convoying
every concurrent turn's persistence behind the global writer lock, and
adds the gate that stops the class from re-entering.
The gateway shares ONE SessionDB across every agent. A read-only query
under `with self._lock:` blocks all concurrent writers for its
duration; under WAL, _read_ctx() serves the same read from a pooled
read-only connection with no lock at all (non-WAL falls back to the
locked writer byte-for-byte, so DELETE-journal installs are unchanged).
Converted (SELECT-only bodies, mechanical `with self._lock:` →
`with self._read_ctx() as conn:`): gateway routing loaders, session/
message counters, titles, compression tip/lineage/cooldown readers,
telegram topic bindings, prune candidate scans, meta readers, resume
resolution — 39 methods, verified per-method that every statement is a
SELECT and every conn use stays inside the with-block. Read-modify-
write methods and checkpoint/maintenance PRAGMAs stay on the writer
lock (their read is ordered against their own write).
Gate: tests/state/test_no_locked_readers_gate.py — an AST scanner over
SessionDB that fails CI when a pure-read method body takes the writer
lock, with a sabotage self-check proving the scanner fires. On
pre-conversion main it reports 38 violations; on this branch, zero.
Measured (2 writer threads + 3 reader threads, 3s, WAL):
before: 78.5k reads, p99 2.49ms, max 160.7ms (readers convoy)
after: 200.4k reads, p99 0.53ms, max 25.5ms (2.6x throughput,
4.7x better p99, 6x better tail; writes unchanged)
Honest caveat: on runtimes where WAL is refused (the currently-bundled
SQLite 3.46 trips the WAL-reset-vulnerability gate → journal=DELETE),
_read_ctx() falls back to the identical locked-writer path and this
change is behavior-neutral by construction; the win applies to WAL
installs (legacy WAL DBs, fixed runtimes, wal-configured operators).
Suites: tests/test_hermes_state.py 243 passed; combined state sweep
540 passed — the only 3 reds are the pre-existing
test_fts_runtime_rebuild failures, verified failing on clean
origin/main before this change.
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:
- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
_rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()
Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.
Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
The PR narrowed _is_fts_write_corruption_error to only match FTS5-specific
'fts5: corrupt structure record' errors, dropping the generic 'database disk
image is malformed' match. But FTS shadow table corruption (the common case)
raises the generic error on SQLite < 3.53, not the FTS5-specific one. This
broke FTS self-heal for 10 existing tests and for users on older SQLite.
Restore the generic match via is_malformed_db_error. Safety is preserved
because the FTS rebuild only touches derived indexes — if the damage is
actually in a canonical B-tree, the rebuild itself fails and the write
propagates.
Also restore the original test assertion and remove the
test_generic_malformed_write_fails_closed test whose premise (generic
corruption should not trigger FTS rebuild) was wrong for the FTS self-heal
path.
Address review feedback from @jackulau on PR #90871:
1. psutil.open_files() silently drops '(deleted)' WAL sidecar entries
on Linux because isfile_strict() stats the literal path including
the suffix and fails. Switch to direct /proc/<pid>/fd readlinks
which preserve the '(deleted)' suffix so _canonical can match.
2. psutil.process_iter() converts AccessDenied to None, which
or-() skips silently — the fail-closed branch never runs. For the
root-gateway vs user-desktop topology in the issue, the fd table is
unreadable but /proc/<pid>/cmdline is world-readable. Add a cmdline
fallback that flags uninspectable processes.
Also keep the psutil path for macOS/BSD (no '(deleted)' convention).
Only the initial SELECT of _dedupe_legacy_system_prompts was guarded;
a 'database is locked' on any per-row write propagated out, aborted
schema init, left the schema version below 25, and made every later
SessionDB.__init__ re-enter the same migration against the same
contended DB - the second half of the enterprise crash-loop report.
The per-row loop now catches OperationalError, logs once, and returns.
Partial migration is safe by design: the legacy system_prompt column
is the documented read fallback for unmigrated rows, and the next
schema init resumes where the contention stopped. Tests prove rows
migrated before the failure stay migrated, the remainder stays
readable, and a later run completes it.
CI caught the sibling site the in-place fix missed: legacy (non-in-place)
compression rotates via publish_compression_child, where a mid-summary
append previously stranded in the closed parent. Same watermark + pure-SQL
column clone as archive_and_compact, with session_id rewritten to the child.
Lineage-guard test flipped to pin the appends-flow-freely contract; rotation
watermark tests added (tail follows the child; None = historical behavior).