close_all_under returned after the last release dropped the generation
and before the physical close finished, so rmtree still saw the open
handle. Wait directory-matching teardown barriers the same way close_all
does.
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.
Co-authored-by: Cursor <cursoragent@cursor.com>
Every production close of a SessionDB goes through hermes_state_registry.release_or_close,
whose teardown swallows any exception from close() at DEBUG. With the capture in place that
meant a RetiredGenerationCaptureError (disk full, permissions) left the handle open silently
and, on Python 3.11, skipped the retention pin: the raise happened before the pin, so the
interpreter's exit still ran sqlite3_close's implicit checkpoint and wrote the stale frames
over the newer generation (probe: 302 rows -> 4 after exit, worse than main).
- close() now decides retention first and takes the pin BEFORE attempting the capture on
runtimes without setconfig; the capture failure is logged at ERROR and re-raised, so a
later close() retries it while the pin already protects the newer generation.
- The registry logs RetiredGenerationCaptureError at ERROR instead of DEBUG (other teardown
errors stay quiet). Same treatment on the release_or_close fallback path.
- The lost-generation settlement moves out of close() into _settle_lost_generation_locked();
the sticky _close_checkpoint_disabled attribute and the dead "setconfig exists but failed"
late-bind block are gone (_disable_close_time_checkpoint returns the call's outcome).
Regression test drives release_or_close with a failing capture: no raise, error text in the
log, pin taken exactly once (3.11), retry closes cleanly. Red on the salvaged head on both
3.11 and 3.12, green here.
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
those files carry no windows_only marker so they never run on Windows, and the monkeypatched
predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
db.close() already releases a registry-shared handle, so the described NameError leak only
existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
teardown; replacement not published before the last close settles) plus the raising-close
settlement; comments say the WHY once.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:
* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
environment failures that polling cannot fix — `_acquire_db_flock` and
both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
now defer immediately with the real errno instead of burning the full
120s / holder timeout and then logging a fake "held by another process".
* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
lock) stayed `_fts_stale` — LIKE-only search — until the process
reopened state.db. Short-lived CLIs reopen every run; the gateway opens
once and stays up for days, so the deferral was effectively permanent
(#100108). The retry runs from the EXISTING gateway housekeeping tick
(`_start_gateway_housekeeping`, 60s) against the shared SessionDB
instances via `hermes_state_registry.live_shared_session_dbs()`:
non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
bounded backoff 60s -> 1h, no new thread, still fails closed on live
holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.
* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
line can be matched to the interpreter that actually linked it
(#100108 point 3).
Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Post-merge review follow-ups on #100201:
- acquire(): drop the never-iterating 'while True' and the redundant
'existing is not generation' half of the race check — after the retire
path, generation is always None on the fresh-open leg, so 'existing is
not None' is the complete condition. Same behavior, flat flow.
- test_mirror: the conversion to the shared registry dropped the
cleanup assertion entirely; restore it by patching
hermes_state.release_or_close and asserting the handle is released
exactly once after _append_to_sqlite.
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).