Commit Graph

5 Commits

Author SHA1 Message Date
teknium1
c451cd1bba test(e2e): torture chamber counts deleted-sidecar hits per episode
The /proc fd monitor appends to one list for the module-scoped chamber,
and each episode asserted the whole history. So one hit failed its own
episode and also every later one: CI showed 6 WAL episodes red for a
single hit during lock_cancellation. Each episode now marks the list at
its start and only asserts the hits recorded after that mark.

Sabotage check (fix reverted, close-gap injection on): 1 of 10 WAL tests
fails, the one that caused the hit. The old harness failed 7.
2026-09-26 11:32:10 -07:00
teknium1
23366b44f3 test(e2e): close the core-suite review findings (narrow xfails, surfaced handler errors, no-retry CI)
Independent review of #120171 found checks that could not fail. Each is now
proven red by a mutation that the old version reported as XFAIL or pass.

- chaos/test_tui_gateway_turn_liveness: the orphaned-tool xfail used
  raises=AssertionError and RpcError subclasses it, so a gateway crash counted
  as the expected failure. Every invariant is now asserted normally; only the
  known leftovers (surviving tool tree and the tool_call it leaves without a
  result, both fixed by #120306) raise ToolOutlivedGateway, the only exception
  the xfail accepts. The DB check used to sit behind the orphan assert and
  never ran; running it exposed the dangling tool_call half of the same bug.
- history/test_prefix_stability: surface_switch's strict xfail tripped at the
  first prefix break, before usage and integrity. Messages/system prompt,
  usage and integrity are asserted first; the tools-array drift is checked
  last and raises ToolsArrayDrift, the only exception the xfail accepts.
- history/test_transcript_ledger: scripted steer/interrupt callables run on
  the fake provider's handler thread, where an assert only dropped the
  connection. Script records those failures and the test re-raises them after
  every turn; steer must land and the interrupted turn must report
  interrupted=True within 30 s.
- fakes/fake_llm_provider: Hang drops the connection at its deadline instead
  of leaving a kept-alive client waiting past it.
- parity: the API server port was picked, released, then bound by the child.
  Readiness now requires our child's pid from authenticated /health/detailed
  and retries on a fresh port when the child reports it in use. The fixture
  guard refused any HERMES_HOME under ~/.hermes, failing all parity tests
  whenever TMPDIR is Hermes's scratch dir; it now refuses only the live root
  or a real profile.
- chaos/_gateway_harness: the gateway stays in pytest's process group, so
  the runner's kill of a timed-out file reaches it.
- sqlite: a DELETE-mode open can fail with SQLITE_BUSY reported as "vtable
  constructor failed: messages_fts"; the delete arm's busy tolerance keys on
  the result code. A failed episode's roles are stopped so the shared chamber
  and rig no longer fail every later episode.
- chaos, compaction, parity homes: updates.check=false (history already had
  it). The passive update check made a GitHub round-trip from every test
  surface, and on a blobless clone whose objects lag upstream its
  `git merge-base --is-ancestor <upstream tip> HEAD` starts a lazy fetch that
  the 5 s timeout orphans; the orphan scans then failed on git processes.
- chaos/test_agent_turn_liveness: a PROBE failure now carries the provider
  call counts and the agent's stale-kill log, so a cross-turn breaker trip
  can be told apart from a slow probe.
- tests.yml e2e: HERMES_TEST_FILE_RETRIES=0 so a race detector's red is never
  retried into green; own uv cache entry (cache-suffix: e2e).
2026-09-23 14:54:36 -07:00
teknium1
84a4b33834 test(e2e): run the state.db torture chamber and compaction suites in DELETE journal mode too
Both suites skipped outright on a WAL-reset-vulnerable SQLite, so the DELETE-mode population
(every user whose Python bundles SQLite 3.7.0-3.51.2, e.g. uv CPython 3.11.14's 3.50.4, and the
network/FUSE homes Hermes also falls back to DELETE on) had no multi-process integrity coverage.

- Parametrize both module fixtures over journal mode [wal, delete]. The delete arm goes through the
  production decision path: each child (_roles.py, including the `hermes` CLI run as __main__)
  pins the version hermes_state_wal.is_sqlite_wal_reset_vulnerable() reports to 3.50.4, and
  apply_wal_with_fallback picks DELETE itself; the harness never issues a journal-mode pragma.
  The wal arm skips only where Hermes would not run WAL; the delete arm runs everywhere.
- Every invariant that holds in both modes stays on in both: integrity_check, acknowledged writes
  exactly once, monotonic counts, repair never lowers rows, FTS parity, fd bounds,
  persisted == replayed. New in both: the store is still in the arm's mode and every SessionDB
  role reported the matching _wal_active (vacuity guard for the seam). WAL-only: the deleted
  -wal/-shm fd scan (the deleted main-file scan runs in both).
- DELETE mode blocks readers on writes and has no writer fairness: in that arm a SQLITE_BUSY
  refusal of a read/open/FTS pass is waited out and counted (reported in failure context), and
  plain writers are paced by 20 ms; in the wal arm busy still fails the role.
- chmod_flip runs 3-5 flips instead of 6-10 (same fault, process start-up dominated the time).

Red-proof (delete arm red, wal arm green): repair without live-writer checks/exclusion ->
"acked ... stored 0x"; a DELETE-mode commit misreported busy so the retry re-runs the insert ->
"acked ... stored 2x" (torture and compaction); default config enabling WAL on a vulnerable
SQLite -> "journal_mode is 'wal' (arm is delete)" + "_wal_active=True in the delete arm".
2026-09-23 14:54:36 -07:00
teknium1
485e25ce68 test: compaction persistence under multi-process contention (C1)
Class C1, "state.db corrupting from compaction": real AIAgent processes run
real turns (user -> terminal tool call -> answer) through the loopback fake
provider on one shared WAL state.db while a plain writer, a reader and an
open/close churner work the same file. A gateway-like agent micro-compacts
every turn; a TUI-like agent runs /compress here 2; faults are clean, kill -9
mid-turn and kill -9 timed into the micro-compaction summary/commit window;
a fresh process then resumes the session. Invariants: every acked user turn
live exactly once, its tool result and answer recoverable (live or
compacted history) and never live twice, the /compress-kept exchanges live
exactly once, canonical rows never shrink (external sampler + reader),
integrity ok, FTS == canonical, resumed request carries every acked user
turn exactly once, plain-writer appends exactly once.

Red-proof: reverting 91df54184d (micro-compaction tool rows) turns every
absorbed tool result/answer into active=0,compacted=0 debris; reverting
526d135a96 (/compress here tail) leaves the kept exchanges with no live row.
2026-09-23 14:54:36 -07:00
teknium1
adc35881ed test: multi-process SQLite torture chamber for state.db integrity (C1)
Class C1: state.db corruption, WAL generations unlinked under a live holder,
lost/duplicated rows, repair destroying data, fd leaks. Every role is a real
OS process on one WAL state.db through the production SessionDB: gateway- and
TUI-like writers with intent/ack journals, a dashboard reader living across
episodes, short-lived openers (SessionDB, bare sqlite3, real `hermes sessions
list/stats`), FTS rebuild/optimize and repair_state_db_schema. Nine seeded
episodes inject kill -9 mid-write, SIGTERM graceful close, POSIX lock
cancellation by a stray in-process open/close, chmod flips, concurrent and
killed FTS rebuilds, repair against a live and an offline store, FTS
corruption and a whole-fleet SIGKILL, then assert the same invariants:
integrity_check ok and still WAL, no child holding a (deleted) db/-wal/-shm
(/proc fd monitor), every acked append stored exactly once, counts only grow,
repair never lowers them, FTS == canonical rows and search finds acked rows
once, no role errors, bounded fds for the long-lived reader and churner.

Red-proof: disabling the OFD WAL lock guard (75e155ab09, whose revert no
longer applies cleanly) makes the lock_cancellation episode kill the gateway
writer with SIGBUS / fail sibling openers, 3/3 runs.
2026-09-23 14:54:36 -07:00