The /proc fd monitor appends to one list for the module-scoped chamber,
and each episode asserted the whole history. So one hit failed its own
episode and also every later one: CI showed 6 WAL episodes red for a
single hit during lock_cancellation. Each episode now marks the list at
its start and only asserts the hits recorded after that mark.
Sabotage check (fix reverted, close-gap injection on): 1 of 10 WAL tests
fails, the one that caused the hit. The old harness failed 7.
Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.
- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
raises=AssertionError and RpcError subclasses it, so a gateway crash counted
as the expected failure. Every invariant is now asserted normally; only the
known leftovers (surviving tool tree and the tool_call it leaves without a
result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
the xfail accepts. The DB check used to sit behind the orphan assert and
never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
first prefix break, before usage and integrity. Messages/system prompt,
usage and integrity are asserted first; the tools-array drift is checked
last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
the fake provider's handler thread, where an assert only dropped the
connection. Script records those failures and the test re-raises them after
every turn; steer must land and the interrupted turn must report
interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
Readiness now requires our child's pid from authenticated /health/detailed
and retries on a fresh port when the child reports it in use. The fixture
guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
constructor failed: messages_fts"; the delete arm's busy tolerance keys on
the result code. A failed episode's roles are stopped so the shared chamber
and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
it). The passive update check made a GitHub round-trip from every test
surface, and on a blobless clone whose objects lag upstream its
`git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
call counts and the agent's stale-kill log, so a cross-turn breaker trip
can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
retried into green; own uv cache entry (cache-suffix: e2e).
Both suites skipped outright on a WAL-reset-vulnerable SQLite, so the DELETE-mode population
(every user whose Python bundles SQLite 3.7.0-3.51.2, e.g. uv CPython 3.11.14's 3.50.4, and the
network/FUSE homes Hermes also falls back to DELETE on) had no multi-process integrity coverage.
- Parametrize both module fixtures over journal mode [wal, delete]. The delete arm goes through the
production decision path: each child (_roles.py, including the `hermes` CLI run as __main__)
pins the version hermes_state_wal.is_sqlite_wal_reset_vulnerable() reports to 3.50.4, and
apply_wal_with_fallback picks DELETE itself; the harness never issues a journal-mode pragma.
The wal arm skips only where Hermes would not run WAL; the delete arm runs everywhere.
- Every invariant that holds in both modes stays on in both: integrity_check, acknowledged writes
exactly once, monotonic counts, repair never lowers rows, FTS parity, fd bounds,
persisted == replayed. New in both: the store is still in the arm's mode and every SessionDB
role reported the matching _wal_active (vacuity guard for the seam). WAL-only: the deleted
-wal/-shm fd scan (the deleted main-file scan runs in both).
- DELETE mode blocks readers on writes and has no writer fairness: in that arm a SQLITE_BUSY
refusal of a read/open/FTS pass is waited out and counted (reported in failure context), and
plain writers are paced by 20 ms; in the wal arm busy still fails the role.
- chmod_flip runs 3-5 flips instead of 6-10 (same fault, process start-up dominated the time).
Red-proof (delete arm red, wal arm green): repair without live-writer checks/exclusion ->
"acked ... stored 0x"; a DELETE-mode commit misreported busy so the retry re-runs the insert ->
"acked ... stored 2x" (torture and compaction); default config enabling WAL on a vulnerable
SQLite -> "journal_mode is 'wal' (arm is delete)" + "_wal_active=True in the delete arm".
Class C1, "state.db corrupting from compaction": real AIAgent processes run
real turns (user -> terminal tool call -> answer) through the loopback fake
provider on one shared WAL state.db while a plain writer, a reader and an
open/close churner work the same file. A gateway-like agent micro-compacts
every turn; a TUI-like agent runs /compress here 2; faults are clean, kill -9
mid-turn and kill -9 timed into the micro-compaction summary/commit window;
a fresh process then resumes the session. Invariants: every acked user turn
live exactly once, its tool result and answer recoverable (live or
compacted history) and never live twice, the /compress-kept exchanges live
exactly once, canonical rows never shrink (external sampler + reader),
integrity ok, FTS == canonical, resumed request carries every acked user
turn exactly once, plain-writer appends exactly once.
Red-proof: reverting 91df54184d (micro-compaction tool rows) turns every
absorbed tool result/answer into active=0,compacted=0 debris; reverting
526d135a96 (/compress here tail) leaves the kept exchanges with no live row.
Class C1: state.db corruption, WAL generations unlinked under a live holder,
lost/duplicated rows, repair destroying data, fd leaks. Every role is a real
OS process on one WAL state.db through the production SessionDB: gateway- and
TUI-like writers with intent/ack journals, a dashboard reader living across
episodes, short-lived openers (SessionDB, bare sqlite3, real `hermes sessions
list/stats`), FTS rebuild/optimize and repair_state_db_schema. Nine seeded
episodes inject kill -9 mid-write, SIGTERM graceful close, POSIX lock
cancellation by a stray in-process open/close, chmod flips, concurrent and
killed FTS rebuilds, repair against a live and an offline store, FTS
corruption and a whole-fleet SIGKILL, then assert the same invariants:
integrity_check ok and still WAL, no child holding a (deleted) db/-wal/-shm
(/proc fd monitor), every acked append stored exactly once, counts only grow,
repair never lowers them, FTS == canonical rows and search finds acked rows
once, no role errors, bounded fds for the long-lived reader and churner.
Red-proof: disabling the OFD WAL lock guard (75e155ab09, whose revert no
longer applies cleanly) makes the lock_cancellation episode kill the gateway
writer with SIGBUS / fail sibling openers, 3/3 runs.