Files
hermes-agent/tests/hermes_state/test_session_db_leak_sweep.py
teknium1 4a757f25dc test: join auto-title threads at teardown; stop titling in the sidecar replay test
tests/gateway/test_timestamp_sidecar_replay.py crashed the interpreter on CI
(native fault, green on rerun) on unrelated PRs. Root cause: every
run_conversation turn in its fixture spawns the auto-title upgrade daemon
thread (title_generator.maybe_auto_title). That thread outlives the test,
fails its model call (no provider under CI), and then writes the derived
title into the fixture's SessionDB after the fixture closed it, which
reopens sqlite on the daemon thread (_reopen_after_close_locked) and prints
the auxiliary-failure warning after pytest capture teardown. At the end of
the file the threads are still in native sqlite while the interpreter
finalizes: the check_same_thread=False-at-shutdown SIGSEGV shape of
#113186. Locally the thread finishes in ~250 ms so the race never shows;
on a loaded runner it lands on finalization.

Fix the class, not the file:
- tests/conftest.py: the autouse SessionDB leak sweep now joins the
  auto-title upgrade threads (bounded, agent.title_generator.
  wait_for_title_upgrades) before closing stores, so no title worker
  outlives its test in any file (5 other files spawn them today).
- tests/gateway/test_timestamp_sidecar_replay.py: titling is not under
  test; the fixture no-ops maybe_auto_title (same as
  tests/agent/test_tool_call_incremental_persistence.py), so its own
  db.close() no longer races a worker either.
- tests/hermes_state/test_session_db_leak_sweep.py: handoff pair pinning
  the invariant (a slow upgrade thread started in one test is dead by the
  next); red on base, green with the fix.

Proof (scratch plugin delaying the title model call by 1 s):
base: 2 auto-title threads alive at interpreter exit, every thread
"SessionDB reopened after close() on thread auto-title"; fixed: no thread
spawned / none alive at exit in this file and the other five.
2026-09-20 11:40:16 -07:00

85 lines
3.3 KiB
Python

"""Suite-wide SessionDB leak-closing contract (OOM incident 20260816).
A raw single-process ``pytest tests/hermes_cli/`` used to accumulate every
SessionDB a test constructed and forgot to close — writer connection,
pooled read connections, and (once token accounting ran) an ``atexit``
registration pinning the instance alive — ballooning to 16-25 GB RSS.
The fix is two-sided:
* ``hermes_state_guard._register_test_instance`` adds every successfully
constructed SessionDB to a WeakSet registry when the
``HERMES_TEST_ISOLATION`` marker is set (test-isolation runs only).
* the autouse ``_close_leaked_session_dbs`` fixture in ``tests/conftest.py``
closes everything in the registry at each test's teardown.
These tests pin the *behavior contract*: instances register under pytest,
close() empties them idempotently, and a leaked instance from an earlier
test is actually closed by the suite-level sweep.
"""
from __future__ import annotations
import threading
import time
import hermes_state_guard
from hermes_state import SessionDB
# Deliberate cross-test handoff: test_leaked_instance_* leaks an instance;
# the later test (pytest runs file order deterministically without a
# randomizer plugin, and the sanctioned runner executes whole files in one
# process) asserts the autouse teardown sweep closed it.
_leaked: list[SessionDB] = []
def test_constructed_sessiondb_is_registered(tmp_path):
db = SessionDB(db_path=tmp_path / "state.db")
try:
assert db in hermes_state_guard._test_instance_registry
finally:
db.close()
# close() must fully release the writer connection…
assert db._conn is None
# …and be idempotent: a second close (the suite sweep will call it
# again at teardown) must not raise.
db.close()
def test_leaked_instance_stays_open_within_the_test(tmp_path):
db = SessionDB(db_path=tmp_path / "state.db")
db.create_session(session_id="leak-probe", source="cli", model="m")
# Intentionally NOT closed — the suite-level sweep owns cleanup.
assert db._conn is not None
_leaked.append(db)
def test_previously_leaked_instance_was_closed_by_the_sweep():
assert _leaked, "expected the previous test to have leaked an instance"
db = _leaked.pop()
# The autouse _close_leaked_session_dbs teardown between the two tests
# must have closed the leaked instance (writer conn released), which is
# what bounds fd/RSS growth in single-process runs.
assert db._conn is None
# Same handoff shape for the auto-title upgrade thread a turn leaves behind: it holds the
# turn's SessionDB, so the sweep must join it before closing stores (else the daemon thread
# reopens the closed store and races interpreter finalization — the SIGSEGV shape of #113186).
_upgrade_thread: list[threading.Thread] = []
def test_slow_title_upgrade_thread_is_left_running_within_the_test():
from agent.title_generator import _UPGRADE_THREADS
thread = threading.Thread(target=time.sleep, args=(1.5,), name="auto-title", daemon=True)
_UPGRADE_THREADS.add(thread)
thread.start()
assert thread.is_alive()
_upgrade_thread.append(thread)
def test_title_upgrade_thread_was_joined_before_the_sweep_closed_stores():
assert _upgrade_thread, "expected the previous test to have started an upgrade thread"
assert not _upgrade_thread.pop().is_alive()