Commit Graph

10 Commits

Author SHA1 Message Date
kshitijk4poor
000f3e0d5d refactor(state): use sqlite3.SQLITE_CONSTRAINT_FOREIGNKEY directly
The getattr(..., 787) fallback guarded a constant every supported
Python (>=3.11) always defines; reference it directly and drop the
module-level alias.
2026-09-26 22:39:42 +05:30
kshitijk4poor
d3faaee370 refactor(agent): key the session-row heal on the classified cause
The heal predicate (errno 787 or a 'foreign key constraint' substring)
and the classifier phrase were two separate definitions of the same
failure. Move the SQLITE_CONSTRAINT_FOREIGNKEY code check into
classify_persistence_error, document the bucket, and branch the heal on
agent._last_persistence_error_cause == 'session_row_missing' so the two
can't drift.
2026-09-26 22:39:42 +05:30
赵桂雄
6e4e0638e8 fix(agent): heal a session row deleted under a live agent (#123583)
`_ensure_db_session` trusts the cached `_session_db_created` flag as proof
the row exists, and the flush path only retries row creation while that
flag is False. Any store-side removal of the row under a live agent —
`hermes sessions delete`, the Desktop/web delete, bulk prune, a
profile-repair move, an in-place store rebuild — leaves the flag stale, so
every later turn's append fails the FK and is dropped with one WARNING per
turn. The agent keeps answering; the durable transcript silently stops
growing, and because the failed transaction leaves no rows behind there is
no post-hoc trace in the store.

Classification (`hermes_state_errors.py`): add the `session_row_missing`
cause — matched by `SQLITE_CONSTRAINT_FOREIGNKEY` (787) when the code
survives, else the RPC-wrapped phrase — so the turn-end explanation names
the real failure instead of "unknown".

Heal (`agent/session_persistence.py::_db_flush_failed`): on that cause,
drop the stale flag, reset the flush markers, call `_ensure_db_session()`,
and replay once within the flush's existing `_adoption_budget`. Because the
deletion erased the session's message rows too, the replay clears the
per-message persisted markers so the FULL in-memory transcript lands on the
recreated row, not just the current tail (mirrors
`_db_flush_adopt_compression_tip`). If row creation also fails, return
False without appending into a guaranteed rollback — fail-open, batch stays
unmarked for the next flush. No new fail-closed path.

(cherry picked from commit ad97047da0f9801a4fdca4478bfd68085f02c3a8)
2026-09-26 22:39:42 +05:30
teknium1
8ac45786bf fix(state): SessionDB open waits out a lock lost inside the FTS constructor
In rollback-journal (DELETE) mode a sibling process can take the write lock
between schema load and the messages_fts probe. FTS5's xConnect then fails its
%_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor
failed: messages_fts". Every state.db lock classifier matched on the words
"locked"/"busy", so:

- a writable SessionDB() failed after 1s instead of waiting out the lock with
  _WRITE_PATIENCE_S, and callers disabled persistence for the run;
- a read-only open (dashboard, `hermes sessions list`, cross-profile readers)
  failed on the first busy timeout with no retry at all;
- the error read as not transient (dashboard 500, not 503) and as persistence
  cause "unknown" instead of "locked".

Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by
result code when SQLite supplies one, text only when it does not (our own
re-raised messages, RPC-wrapped strings). Route the writer open patience loop,
the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the
maintenance holder probe, is_transient_sqlite_error and
classify_persistence_error through it. The read-only open retries a lock
inside its existing bounded retry budget, next to the transient IOERR case.
2026-09-23 11:35:07 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
Sulthan Zahran
38adfe90a4 fix(state): classify FTS-scoped corruption as fts_index, never whole-file damage
classify_persistence_error bucketed every _DB_CORRUPTION_MARKERS hit as "corrupt",
so an error SQLite itself scoped to the FTS5 index layer (SQLITE_CORRUPT_VTAB, or an
`fts5: corrupt structure record for table "messages_fts"` report) that escaped the
write path — the detach in _enter_fts_fail_open refused (generation/lock check),
or a read/search path with no fail-open at all — reached the turn boundary and the
gateway startup notice as structural corruption: the turn ended with `.recover` /
restore-backup advice on a file whose canonical tables were provably healthy.

One provenance rule, hermes_state_errors.is_fts_scoped_corruption_error, now feeds
both the write-repair gate (SessionDB._is_fts_write_corruption_error delegates to it,
so the gateway transcript retry inherits it) and the classifier: a known result code
outranks prose (only SQLITE_CORRUPT_VTAB is FTS-scoped; bare SQLITE_CORRUPT/NOTADB
and any contradictory code fail closed), and without a code the text must both carry
a corruption marker and name a messages_fts* object. The new "fts_index" cause
renders index-scoped guidance (doctor --fix / restart, do not run recovery) in the
turn explainer and the home-channel notice. The structural fail-close is untouched:
bare malformed / not-a-database still quarantine and still classify "corrupt".

Salvaged from PR #97843 (SulthanZahran1), trimmed: the quick_check-backed
"corrupt_unconfirmed" tier is dropped — on a live handle that just observed an
unscoped SQLITE_CORRUPT, PRAGMA quick_check on a damaged shadow b-tree raises rather
than reports on 3.53.1, so the probe could never downgrade the exact shape it was
built for, and a verdict that softens quarantine guidance on prose alone weakens the
fail-close. #97841 (Finn763) reached the same fts_index cause via text markers
only; its LIKE-degradation intent already lives in _search_messages_impl (_fts_stale).

Fixes #97794
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
2026-09-11 06:37:27 -07:00
gaoanze
031d3760d0 fix(state): distinguish retired WAL recovery guidance
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.

Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
ff3ebf509f refactor(state): dict-dispatch persistence-error classifier, unify WHERE/placeholder builders, compact docstrings across state modules 2026-09-02 20:18:08 -07:00
Teknium
e527b0699d refactor(state): extract error types/classifiers to hermes_state_errors and the live-DB guard to hermes_state_guard 2026-09-02 18:44:39 -07:00