Commit Graph

576 Commits

Author SHA1 Message Date
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
daa2f96d8c refactor(state): drop the registered flag; budget unregister is idempotent
_PathReadBudget.unregister() discards from a WeakSet, so close() can call it
unconditionally; a handle whose constructor raised was never registered and the
call is a no-op. Same fix as PR #113857, filed independently.

Co-authored-by: ClintonEmok <54935030+ClintonEmok@users.noreply.github.com>
2026-09-18 10:35:03 -07:00
KoNit-K
a7ed109847 fix(state): exclude closed SessionDB writers 2026-09-18 10:35:03 -07:00
teknium1
a834988165 docs(state): name recovered placeholders in the stale-open sweep; exercise the real recovery path
sessions.md lists the state-owned sources the auto-prune sweep closes; add
`recovered`. The second test drives `_reconstruct_missing_sessions` itself so
the placeholder shape the sweep must age out is the one recovery writes, and
keeps a fresh placeholder open as the control.

Fixes #114730
2026-09-18 10:30:20 -07:00
chelsealong
9ec089bf34 fix(state): auto-prune recovered session placeholders
`sessions recover` synthesizes placeholder rows with source='recovered'
and no ended_at (hermes_cli/session_recovery.py::_reconstruct_missing_sessions).
_AUTO_PRUNE_STALE_OPEN_SOURCES did not include "recovered", so the
stale-open sweep never closed them and prune_sessions (which only
deletes ended rows) could never reach them either — salvaged
placeholders accumulated forever regardless of auto_prune/retention_days.

Add "recovered" to _AUTO_PRUNE_STALE_OPEN_SOURCES so idle placeholders
are closed after the retention window like other state-owned sources
(cli/cron/kanban/acp/api_server/subagent/tool), then pruned after a
further window.

Fixes #114730
2026-09-18 10:30:20 -07:00
EvanZhu0721
e088a092f8 fix(state): identify duplicate SessionDB holders 2026-09-17 19:07:11 -04:00
kshitijk4poor
532b6d1cb1 fix(state): pin the session-db-unavailable fallbacks to the failing profile
format_session_db_unavailable consumes the now-pinned cause table but its own
two fallbacks (no recorded cause; network-drive gloss) still printed a bare
`hermes doctor`. Same class; one test covers both fallbacks.

Also: spell the corrupt action as plain concatenation instead of an f-string
with escaped braces, and fix the finalizer comment to state the invariant
(final_response is never rebound) rather than describe a local that no longer
exists.
2026-09-16 21:15:30 +05:30
outpoints
44ba32a565 fix(honcho): preserve deferred routing invariants
Thread logical session cwd through deferred Desktop/TUI builds, normalize absent cwd during construction, and share title provenance constants between SessionDB and Honcho.

(cherry picked from commit 2693f4f27c776ac819d92c9b52e8a03ad2a985d8)
2026-09-15 22:30:11 -07:00
teknium1
cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
KoNit-K
6a8862404d fix(state): preserve replies after consecutive review harness prompts 2026-09-15 05:41:36 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
d12cec1833 fix(tests): keep the leak registry in hermes_state_guard and skip shared handles
Rebase follow-up. hermes_state.py is a facade now; the test-isolation
guard code the registry sits beside moved to hermes_state_guard.py, so
the WeakSet and _register_test_instance live there (gated on the same
_TEST_ISOLATION_MARKER_ENV the guard already owns) and the facade only
calls the helper from __init__.

The sweep now skips instances flagged _shared_registry_owned: since
#90837 close() on a hermes_state_registry.acquire() handle releases a
refcount instead of closing, so sweeping them would retire a shared
generation that a wider-scoped fixture still holds. The registry owns
that lifecycle (close_all()).

Per Enough1122's nit, the registry comment states explicitly that
production never populates it and that the gate must not be removed.
2026-09-15 03:44:50 -07:00
Teknium
d070e480a3 fix(tests): close leaked SessionDB handles suite-wide and cap pytest memory
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.

Fix the class, not the sites:

* hermes_state: register every successfully constructed SessionDB in a
  test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
  i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
  everything left in the registry after each test. close() is idempotent
  and unregisters the pinning atexit hook, so instances become
  collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
  defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
  leak fails fast with MemoryError instead of eating the box.
  Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
  scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
  for registration, idempotent close, and the cross-test sweep.

Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.

Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
2026-09-15 03:44:50 -07:00
kshitijk4poor
30a299180c refactor(state): stdlib-only home for read_only_db_uri; reuse the doctor/repair helpers
hermes_state_common pulls in agent.* at import, so the URI builder moves to
hermes_state_holders (errno/os/sqlite3/pathlib only) where the gateway
readiness probe and backup can adopt it in a follow-up sweep. The doctor
structural-damage branch is one helper instead of two copies, the holder
scan goes through hermes_state_repair._live_writer_holds_db, the migration
hint uses _schema_not_built (the startswith("no such ") check also matched
"no such module: fts5"), and the hermes_state import is hoisted so an import
failure cannot mask itself as UnboundLocalError.
2026-09-15 12:51:27 +05:30
kshitijk4poor
55d9a49c1d refactor(state): one read-only URI builder; probe a held store via snapshot only
read_only_db_uri() replaces four inline mode=ro URI sites (two of which
still used the raw f-string that truncates on ?/# in the home path:
state_db_has_structural_damage and collect_state_db_stats). The doctor
write probe now applies the live-holder gate in both modes: a quiet store
is probed in place as on main, a held store is probed through a read-only
snapshot, and a held store over 1 GB is skipped with an info line unless
--fix is given (the unconditional copy cost one full DB write per plain
doctor run). Connect/backup failures propagate to the existing
classification instead of being reported as FTS write-health failures.
Observational sessions commands print a migration hint instead of a raw
traceback when a read-only opener meets an older schema.

Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
2026-09-15 12:51:27 +05:30
kshitijk4poor
4a268c9e4a fix(cli): keep doctor's read-only opens URI-safe and holder-gated
- _session_count: back to main's raw sqlite mode=ro COUNT(*) via as_uri() — routing it through SessionDB(read_only=True) both re-introduced the raw f-string URI ('?'/'#' in the home path truncate it) and queries columns (s.archived) an unmigrated store lacks, so doctor would report a healthy DB as broken.
- _write_health_reason: snapshot source URI built with as_uri() for the same reason; the --fix live probe (_db_opens_cleanly runs BEGIN IMMEDIATE) now falls back to the snapshot unless live_writer_holds_db proves the store quiet, matching _state_db_wal — hermes doctor --fix never becomes a second writer against a gateway's state.db (#103339).
- SessionDB._connect_read_only: same as_uri() form so every read-only opener is safe in a home containing '?' or '#'.
- test_sessions_export_output_dir: fixture accepts the read_only kwarg the PR introduced.
- Drop the two doctor tests that pinned the SessionDB factory kwargs; main's URI-reserved-chars test covers _session_count.

Co-authored-by: Ahmett101 <Ahmett101@users.noreply.github.com>
2026-09-15 12:51:27 +05:30
kshitijk4poor
efca6efdc8 refactor(hermes_state): one canonical_sqlite_path
`hermes_state_dbfile._canonical_sqlite_path` was a byte-identical copy of
`hermes_state_holders.canonical_sqlite_path`; keep the public one and repoint
the two hermes_state call sites. No import cycle: holders is stdlib+psutil.
2026-09-14 21:13:54 +05:30
teknium1
274fd56dca fix(state): WAL lock guard follows the handle's lifecycle
Three gaps in the #110544 guard, all reported in its review and reproduced:

- A writer reopened by _reopen_after_close_locked (teardown/worker race,
  #94736) came back with no guard: the next stray close + foreign close
  deleted its WAL again.
- _try_wal_checkpoint refreshed the guard outside self._lock; landing after
  close() it pinned an OFD lock with no connection behind it, so a foreign
  `PRAGMA journal_mode=DELETE` saw `database is locked` forever.
- Refcounts keyed on (fd, inode) treated a recycled fd number as a surviving
  lock: A+B live, close A, C reuses A's fd, close B left C recorded as guarded
  while a foreign EXCLUSIVE succeeded.

The guard now counts handles per inode, re-locks every matching descriptor on
each hold (OFD re-lock is idempotent), and unlocks on the last handle only;
the reopen path holds it; the checkpoint refresh runs under self._lock and
skips a closed handle. The macOS holder scan folds case so a case-only alias
of the sidecar path on APFS still matches.
2026-09-14 06:54:07 -07:00
teknium1
beb546b0f2 fix(state): lock guard rides SQLite's own descriptors
OFD locks on the connection's fds die with the connection: no private
descriptors to track, retire, or exclude from holder scans.
2026-09-14 05:28:22 -07:00
teknium1
75e155ab09 fix(state): a live writer's WAL generation survives lock cancellation and sibling closes
SQLite protects a WAL generation with per-PROCESS POSIX locks (SHARED on
state.db, DMS byte on -shm). Any in-process open()/close() of either file
cancels both (sqlite.org/howtocorrupt.html §2.2); the next last-connection
close in ANY process then checkpoints and unlinks -wal/-shm, and the holder
sticky-halts with DeletedWalGenerationError. #109841 removed one such
close (mode tightening) but the class is open-ended: raw header probes,
plugins, tool reads of ~/.hermes, any library that touches the files.

hermes_state_lockguard re-holds the same two ranges as OFD locks
(F_OFD_SETLK) on private descriptors for as long as a writer handle is
open. OFD locks belong to the open file description, so a stray close()
cannot cancel them, and they conflict with the EXCLUSIVE a sibling needs
for the close-time reset exactly like SQLite's own. Released before the
handle's own close so a true last close still ends the generation; the
descriptors are closed only once no connection to the path remains, so a
holder scan from another process never counts them. Works on Python 3.11
(where sqlite3 cannot arm SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE) and on macOS
(F_OFD_SETLK=90 per XNU bsd/sys/fcntl.h); no-op on Windows.

Live repro (Linux, Python 3.11.15, SQLite 3.53.1): holder = SessionDB
writer; in-process os.open/os.close of state.db and -shm; then a foreign
sqlite3.connect()+close(). Before: -wal unlinked, holder write raises
DeletedWalGenerationError. After: -wal keeps its inode, holder writes.
2026-09-14 05:28:22 -07:00
teknium1
c3de99bbf0 refactor(state): one carrier-aware user-turn rewind behind CLI /undo, /retry, gateway and TUI
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).

The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
2026-09-13 05:20:26 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
e16f686706 fix(state): never open+close an existing state.db inode while tightening modes
create_main used O_WRONLY|O_CREAT on the main file and closed the fd, which drops this process's POSIX locks whenever state.db already exists (the gateway's own async_delegation import path). O_EXCL restricts the descriptor to a brand-new inode; existing files take the chmod(2) path.

Refs #109786 #109687
2026-09-13 04:29:06 -07:00
Konstantin Khlopkov
ccd360e94f fix(state): tighten existing db files by chmod(2), not an open/fchmod/close cycle
POSIX fcntl locks are owned per (process, inode): closing any descriptor
for state.db releases every lock the process holds on that inode,
including the locks of an already-open SQLite connection. The
owner-only hardening cycle opened the live database and its -wal/-shm
read-only, fchmod'ed, and closed, so any process that already held a
connection (gateway, desktop hermes serve, dashboard share one) dropped
its live locks on every SessionDB init. A sibling process then took the
shared-memory DMS exclusively at its own close, checkpointed, and
unlinked the sidecars while long-lived holders kept the deleted inodes
open, tripping the deleted-WAL generation guard.

chmod(2) on the path never opens the file, so it cannot disturb locks.
The descriptor path remains only for first-time main-db creation, where
no locks can exist yet.
2026-09-13 04:29:06 -07:00
teknium1
d21913b09c fix(state): let sqlite raise the canonical error for a directory db_path
The owner-only pre-create helper ran before sqlite3.connect() and turned
a directory-as-state.db misconfiguration into IsADirectoryError instead
of the sqlite OperationalError the open path (and its lock-patience
classifier) expects. A directory leaks no row data, so skip it and let
sqlite fail canonically. Also map the salvage carry-commit author email
for the attribution gate.
2026-09-12 20:43:42 -07:00
Hermes fleet-fix
1916cb249d security: make state databases and snapshots owner-only 2026-09-12 20:43:42 -07:00
Teknium
76d21f93d2 fix(profiles): a served turn never recreates an archived profile home (#94590)
SessionDB._open_writer did db_path.parent.mkdir(parents=True), so a multiplexer
or Desktop backend still holding a moved-away profile's route re-scaffolded
profiles/<name> (then cache/, cron/, logs/, SOUL.md...) on its next turn. Route
the mkdir through mkdir_under_hermes_home, which refuses a deleted or missing
named profile home — the same guard config/logging/cron already use.
2026-09-11 19:50:46 -07:00
teknium1
e00a21cd39 fix(state): retry a transient SQLITE_IOERR on the pooled read path before surfacing it
Since 0.21.0 reads go through mode=ro pooled connections. A read-only OPEN already
rides out the millisecond WAL transition window (checkpoint / WAL reset / frame flush
by a sibling process; the ro reader cannot rewrite the -shm index) with a bounded retry
(#100436), but a WARM pooled reader hitting the same window while its SELECT executes
propagated `disk I/O error` straight out of get_session(): 37 identical tracebacks on a
multi-process WSL2 ext4-on-vhdx install, each followed by "compression session recovery
failed", with quick_check=ok (#100871). The reporter's A/B shows the operator
workaround (journal_mode=delete) collapses read throughput ~30000x, so the flake has to
be absorbed on the read path.

_read_one/_read_all now replay the idempotent statement within the existing read-only
IOERR budget (3 x 50 ms) on the SAME connection -- close+reopen would cancel this
process's POSIX locks for every sibling connection -- and a persistent IOERR still
propagates. No quarantine: EIO on a read is busy, not broken. Every SELECT in the
SessionDB siblings (63 call sites) reaches the pool through these two helpers, so the
class is covered without a wrapper type.

Same-connection retry per #100882's analysis (@fangliquanflq); #100883
(@Sahilvishnaliya) diagnosed the missing recovery in the 0.21.0 read pool.

Fixes #100871.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Sahilvishnaliya <222165401+Sahilvishnaliya@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
teknium1
295edf2557 fix(state): gate the lock-free read pool on a confirmed WAL header, not the assumed mode
apply_wal_with_fallback() reports "wal" in two indeterminate cases -- the vulnerable-
SQLite gate (_apply_delete_for_wal_reset_bug) and the non-vulnerable probe-unknown path
(a7f2a593d1) -- meaning "touched nothing, the connection inherits the header's mode".
SessionDB turned that assumption into `_wal_active=True`, which enables the mode=ro read
pool that skips `self._lock`. On a file that is really in rollback-journal mode those
readers race the writer with a 5s busy timeout and no retry: random SQLITE_BUSY read
failures for the instance's lifetime (#86515).

Confirm the header on the freshly opened connection before enabling the pool. When the
probe is still blocked, reads queue on the writer connection under the lock -- slower,
never wrong. Every other apply_wal_with_fallback caller ignores the return value, so the
consumer is the right place to gate; changing the return contract to Optional across
15 call sites (#87044's shape) is not needed.

Live repro: DELETE-mode file, sibling holding BEGIN EXCLUSIVE during open ->
before: _wal_active=True and _checkout_read_conn() hands out a pooled mode=ro conn;
after: _wal_active=False, reads take the locked writer path.

Fixes #86515. Based on the analysis in #87044.
Co-authored-by: QDung210 <dqdung205@gmail.com>
2026-09-11 06:23:23 -07:00
ca-shrimp
3a3ec6eeaf fix(state): serialize replaced/generation probe with close() in _execute_write
A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (#105567).

Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.

Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.

Fixes #105567
2026-09-11 06:23:23 -07:00
chelsealong
d6b036b726 fix(state): compare fd identity against the watched path, not st_nlink
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
2026-09-11 06:23:23 -07:00
chelsealong
84a3c4de74 fix(state): require nlink==0 before treating a /proc fd as an unlinked WAL sidecar
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.

Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
2026-09-11 06:23:23 -07:00
kshitijk4poor
c4c7e102aa refactor(state): tighten generation-stamp comment to the why 2026-09-11 06:21:00 -07:00
John Paul Soliva
ef8d682200 perf(state): stop taking the state.db write lock to open a database that needs no writes
Every read-write SessionDB open issued three writes that usually change
nothing, and a write statement takes the database write lock even when it
matches no rows:

- _ensure_db_file_generation's INSERT OR IGNORE into state_meta. The stamp
  is minted once per FILE, so every open after the first inserted nothing.
- the NULL-`active` heal, UPDATE messages SET active = 1 WHERE active IS
  NULL, which matches nothing on a healthy database.
- the fts_storage_version stamp, which re-wrote the same value on every
  open of an already-optimized database.

The connection is opened with timeout=1.0, so each blocked write costs a
full busy timeout while a sibling process holds the write lock, and the
open path's patience loop can ultimately give up and raise.

Gate all three on a read. Measured on a real 99 MB state.db (239
sessions, 6470 messages) with a sibling holding the write lock: 2117-2136
ms -> 3.6-7.1 ms. On an already-optimized database the unpatched open does
not merely stall, it raises `database is locked`; patched it completes in
3.7-6.8 ms. With a sibling running 200 ms write transactions in a loop
(n=20 opens): p50 894.6 -> 9.6 ms, p90 1094.7 -> 13.9 ms. A settled
database now issues zero main-database write statements to open.

The reads cost nothing measurable: the state_meta probe is a primary-key
seek (2.0 us), the messages probe is 1.7 us on the modern NOT NULL column
(unsatisfiable constraint, short-circuited) and 0.4-0.6 us at 300k rows on
a legacy default-less column via the partial index that already exists for
exactly this predicate. An uncontended open is unchanged.

Semantics are preserved. INSERT OR IGNORE still resolves the first-opener
race inside SQLite and racers still converge on the winner's token via the
re-read; the application_id gate and the PASSIVE-only checkpoint are
untouched; the heal is still considered on every startup, as #60108
deliberately made it, with only the write now conditional on a read
proving there is something to repair.

Read-first also fixes a correctness bug. Under contention the generation
block was abandoned by its `except sqlite3.Error` handler, so a process
ended up with no generation token at all even though the value was already
on disk and a plain read would have returned it -- and that token feeds
the deleted-WAL and replaced-file guards added by #101221. The heal's
`except OperationalError: pass` likewise skipped the repair silently, so
the unconditional form did not even deliver the unconditional repair it
advertised whenever it mattered most.

The probe deliberately does not use INDEXED BY: that hint raises
OperationalError("no query solution") against the modern NOT NULL column,
and the existing handler would swallow it, disabling the repair forever.
2026-09-11 06:21:00 -07:00
kshitijk4poor
6ac77111af refactor(sessions): simplify the retired-generation capture after review
- capture_retired_wal_generation: drop the main_image_max_bytes kwarg (no caller passes it; read the module constant).
- replace _write_json_durably with utils.atomic_json_write (late import keeps the capture dependency-light).
- clean the abandoned .partial staging dir on capture failure instead of leaving it for the retry to work around.
- _disable_close_time_checkpoint: document that it is the per-instance twin of _close_time_checkpoint_configurable and must agree with __init__'s capability decision.
- _close_quietly kept on the late import (module-level import from hermes_state_dbfile risks future cycles).
2026-09-11 11:14:56 +05:30
kshitijk4poor
64fe13a647 fix(sessions): retire-capture integrity - backups skip capture dirs whole; short reads fail the capture
- hermes backup excluded *.db-wal by suffix but not the retired-wal capture dirs, so it would ship the capture's main-image copy while dropping the captured -wal that is the artifact's point; exclude <name>.retired-wal-* dirs whole (they must move as manifest+image+wal unit).
- _copy_range treated a short read as success, yielding a truncated copy with a valid manifest while the unlinked inode still dies at exit; raise RetiredGenerationCaptureError and clean the .part file.
- _quarantine_reason docstring no longer claims close() checks replaced before generation loss (close evaluates loss first and skips quarantine when lost).
2026-09-11 11:14:56 +05:30
kshitijk4poor
5cb9f5c347 fix(sessions): keep the writer-conn lock audit green for the lost-generation close paths
_pin_connection took the shared connection via self._conn, so the lexical
lock audit flagged it (then its caller) as unlocked use. Pass the conn in
from the lock-held close paths and whitelist _settle_lost_generation_locked
as a close()-only lifecycle helper.,
2026-09-11 11:14:56 +05:30
kshitijk4poor
d74a88c51e fix(sessions): a runtime setconfig failure on a lost-generation handle logs at ERROR
Where SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE exists no retention capability is bound, so a
setconfig() that raises means close() will run SQLite's close-time checkpoint over the
newer generation. That fallthrough was a DEBUG line; an operator has to see it.
2026-09-11 11:14:56 +05:30
kshitijk4poor
d87bf0d07e fix(sessions): a failed retired-generation capture still pins the 3.11 handle and logs at ERROR
Every production close of a SessionDB goes through hermes_state_registry.release_or_close,
whose teardown swallows any exception from close() at DEBUG. With the capture in place that
meant a RetiredGenerationCaptureError (disk full, permissions) left the handle open silently
and, on Python 3.11, skipped the retention pin: the raise happened before the pin, so the
interpreter's exit still ran sqlite3_close's implicit checkpoint and wrote the stale frames
over the newer generation (probe: 302 rows -> 4 after exit, worse than main).

- close() now decides retention first and takes the pin BEFORE attempting the capture on
  runtimes without setconfig; the capture failure is logged at ERROR and re-raised, so a
  later close() retries it while the pin already protects the newer generation.
- The registry logs RetiredGenerationCaptureError at ERROR instead of DEBUG (other teardown
  errors stay quiet). Same treatment on the release_or_close fallback path.
- The lost-generation settlement moves out of close() into _settle_lost_generation_locked();
  the sticky _close_checkpoint_disabled attribute and the dead "setconfig exists but failed"
  late-bind block are gone (_disable_close_time_checkpoint returns the call's outcome).

Regression test drives release_or_close with a failing capture: no raise, error text in the
log, pin taken exactly once (3.11), retry closes cleanly. Red on the salvaged head on both
3.11 and 3.12, green here.
2026-09-11 11:14:56 +05:30
Totoro-qaq
d93f72c460 fix(sessions): retire lost-generation writers unclosed where the close-time checkpoint cannot be switched off
With the retired generation captured durably, closing a lost-generation
handle is safe wherever SQLITE_DBCONFIG_NO_CKPT_ON_CLOSE took effect: the
frames are preserved and sqlite3_close no longer checkpoints them into the
newer generation. On Python 3.11, where sqlite3 has no setconfig, closing
still runs SQLite's internal checkpoint over the newer main file, so only
there the exact quarantined connection is retained instead of closed: one
public Py_IncRef reference via ctypes.pythonapi (no struct-layout access),
bound before the writer opens, taken in close() after the capture. Runtimes
with setconfig never touch ctypes; a writable SessionDB requires CPython
with ctypes only where retention is the sole guard.

Read-only handles and every other quarantine reason close as before.
Regressions cover both branches: where retention applies, the retired inode
stays readable after close() and GC and an independent deleted WAL under
the same pathname is untouched; elsewhere the capture is the surviving copy.
A separate process's newer generation survives close, GC and normal exit
in both cases. The mock-based close-time regression originally written for

Refs #105670

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-11 11:14:56 +05:30
Totoro-qaq
d616424571 fix(sessions): capture a lost WAL generation durably before shutdown settles
After DeletedWalGenerationError the retired frames exist only in an
unlinked -wal inode that this process keeps open. #106315 stops them from
being checkpointed under wrong page numbers, but the canonical remediation
("stop the gateway, dashboard and cron writers, then reopen") lets the
kernel drop the inode with them: committed transactions whose disposition
is unknown were destroyed by the prescribed recovery path itself.

Capture the exact retired generation next to the database at the first
halt, or at close() if that sees the loss first. The WAL is located by the
(st_dev, st_ino) recorded at open among this process's own descriptors,
never by pathname, so a sibling's deleted WAL or a newer sidecar minted at
the same path cannot be mistaken for it; it is read with pread and no
descriptor is closed, moved or truncated. The artifact holds the WAL, the
-shm when still ours, the main image (or its header past a size cap) and a
manifest with identities, digests and the sidecar generation found at the
path at capture time. Nothing is merged: whether the frames belong on top
of the file now at the path stays an operator decision. close() refuses to
settle without the capture: it raises and leaves the handle open.

Regressions: capture at halt and at close, refusal to settle on capture
failure, inode-not-pathname selection with a second deleted WAL under the
same name, refusal to guess without a recorded identity, and a subprocess
control where rows committed only in the retired WAL are recovered from
the capture alone after the writer process has exited.

Refs #105670
2026-09-11 11:14:56 +05:30
kshitijk4poor
9e84ce5bae refactor(state): one pre-open header reader for both probes
`is_zeroed_state_db` and `has_invalid_sqlite_header_preopen` shared the
is_file / stat / live-connection / read_header preamble; `_preopen_header`
owns it now (zeroed is the NUL subset of "no SQLite header"). The
quarantine message also names `hermes sessions recover --source <bak>` so
the preserved bytes are actionable, not just parked.
2026-09-09 18:25:11 +05:30
Ahmett101
dc8044c086 fix(state): quarantine a state.db whose page 0 is not SQLite, not just a zeroed one
Startup only quarantined a 0-byte / all-NUL state.db. A file whose first page
was clobbered with record bytes (#102198) went straight to sqlite3.connect,
which raised "file is not a database" and deleted the -wal sidecar — the one
piece of evidence that could have been recovered.

`has_invalid_sqlite_header_preopen` generalises the zeroed probe (zeroed is a
subset of "no SQLite header"; same live-connection contract, never raises).
`quarantine_invalid_state_db` moves the file AND its -wal/-shm aside as
`state.db.<zeroed|notadb>-<ts>-<pid>.bak` before anything opens it; a fresh
DB is created as before.

Re-authored on current main (the quarantine helpers moved to
hermes_state_dbfile.py in d15c61b5dc); one regression test proves the
notadb case is quarantined with its sidecars and the new DB passes
integrity_check.

Refs #102198 (the write-after-SIGTERM that clobbers page 0 is not addressed
here; this preserves the evidence instead of destroying it).
2026-09-09 18:25:11 +05:30
kshitijk4poor
f290946e4e fix(state): the deferred FTS rebuild retry is quarantined by the same rule as the checkpoints
retry_deferred_fts_recovery gated only on _db_corrupt ("mirrors _try_wal_checkpoint /
close") — after this PR it no longer mirrored them: on a replaced/lost-generation handle
the periodic housekeeping tick still ran FTS DDL/DML + commit, the same split-brain write
class as the #105670 checkpoint. One SessionDB._quarantine_reason() now decides for the
periodic checkpoint, close(), and the FTS retry, with the halt path's precedence
(replaced before generation loss) and the operator wording in one place.

Test: the periodic-checkpoint case folds into the close test (same setup), which now
also proves the FTS retry returns False without touching the file; the mutation with
main's schema sibling swapped in returns True (a rebuild ran).
2026-09-09 12:21:19 +05:30
kshitijk4poor
7a4985a814 refactor(state): drop the ruff-format reflow from the checkpoint guard, keep the ~20 semantic lines
The cherry-picked commit re-wrapped hermes_state.py wholesale (+527/-131 for a
fix of about twenty lines). Restore main's layout and re-apply only the fix:
disable the close-time checkpoint on both generation-loss halts, gate the
periodic checkpoint on the sticky generation flags, and name the quarantine
reason at close.
2026-09-09 12:21:19 +05:30
Konstantin Khlopkov
68d9365aa7 fix(state): guard close-time checkpoint for replaced/deleted-generation handles (#105670)
- close() and _try_wal_checkpoint() now skip when _db_replaced or _db_wal_generation_lost
  (previously only _db_corrupt was checked) — prevents checkpointing stale-generation frames
  into the main DB, which is the shutdown-time damage reported in #105670
- _halt_if_db_generation_changed() calls _disable_close_time_checkpoint() alongside the flag
  set (3.12+: disables SQLite internal last-connection checkpoint too)
- Regression tests: halted handle must not run explicit PRAGMA checkpoint on close(),
  halt must call setconfig(NO_CKPT_ON_CLOSE), periodic _try_wal_checkpoint() must skip
2026-09-09 12:21:19 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00