fix(state): SessionDB open waits out a lock lost inside the FTS constructor
In rollback-journal (DELETE) mode a sibling process can take the write lock between schema load and the messages_fts probe. FTS5's xConnect then fails its %_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor failed: messages_fts". Every state.db lock classifier matched on the words "locked"/"busy", so: - a writable SessionDB() failed after 1s instead of waiting out the lock with _WRITE_PATIENCE_S, and callers disabled persistence for the run; - a read-only open (dashboard, `hermes sessions list`, cross-profile readers) failed on the first busy timeout with no retry at all; - the error read as not transient (dashboard 500, not 503) and as persistence cause "unknown" instead of "locked". Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by result code when SQLite supplies one, text only when it does not (our own re-raised messages, RPC-wrapped strings). Route the writer open patience loop, the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the maintenance holder probe, is_transient_sqlite_error and classify_persistence_error through it. The read-only open retries a lock inside its existing bounded retry budget, next to the transient IOERR case.
This commit is contained in:
@@ -40,7 +40,7 @@ from hermes_state_errors import (
|
||||
_DELETED_WAL_GENERATION_MSG, _DISK_IO_ERROR_MARKER, _STATE_DB_CORRUPT_MSG, _STATE_DB_GENERATION_KEY,
|
||||
_STATE_DB_REPLACED_MSG, DeletedWalGenerationError, SessionCompressionInProgressError, StateDbCorruptError,
|
||||
StateDbReplacedError, _is_no_more_rows, classify_persistence_error, is_malformed_db_error,
|
||||
is_malformed_schema_error,
|
||||
is_malformed_schema_error, is_sqlite_lock_error,
|
||||
)
|
||||
from hermes_state_guard import (
|
||||
_STATE_DB_GUARD_BYPASS_ENV, _in_test_context, _is_production_state_db, _real_platform_state_root,
|
||||
@@ -711,7 +711,9 @@ class SessionDB(
|
||||
# SQLITE_IOERR to a mode=ro reader (it can't do the -shm recovery the read
|
||||
# needs). Closes in milliseconds: retry a bounded number of times before
|
||||
# classifying the store as failed (#100436; see _READ_ONLY_IOERR_RETRY_ATTEMPTS).
|
||||
transient = _DISK_IO_ERROR_MARKER in str(ioerr).lower()
|
||||
# A DELETE-mode writer's commit outlasting the busy timeout is the same "busy,
|
||||
# not broken" class; each retry waits the busy timeout again.
|
||||
transient = is_sqlite_lock_error(ioerr) or _DISK_IO_ERROR_MARKER in str(ioerr).lower()
|
||||
if attempt >= _READ_ONLY_IOERR_RETRY_ATTEMPTS or not transient:
|
||||
raise
|
||||
time.sleep(_READ_ONLY_IOERR_RETRY_BACKOFF_S)
|
||||
@@ -801,8 +803,7 @@ class SessionDB(
|
||||
self._connect_and_init()
|
||||
return
|
||||
except sqlite3.OperationalError as exc:
|
||||
err = str(exc).lower()
|
||||
if "locked" not in err and "busy" not in err:
|
||||
if not is_sqlite_lock_error(exc):
|
||||
raise
|
||||
self._close_connection_quietly(self._conn)
|
||||
now = time.monotonic()
|
||||
@@ -1041,7 +1042,7 @@ class SessionDB(
|
||||
continue
|
||||
err_msg = str(exc).lower()
|
||||
if isinstance(exc, sqlite3.OperationalError):
|
||||
if "locked" in err_msg or "busy" in err_msg:
|
||||
if is_sqlite_lock_error(exc):
|
||||
if self._sleep_before_write_retry(deadline, patience_s):
|
||||
continue
|
||||
# Say what actually happened, not disk/permission damage. The holder goes to
|
||||
|
||||
Reference in New Issue
Block a user