Commit Graph

19 Commits

Author SHA1 Message Date
teknium1
f2755aee20 fix(gateway): stalled-session spool replay probe no longer warns on every append
Once a stalled session's backlog is spooled (#114266), the order-preserving
replay attempt before each write hit the still-dead DB and logged a
gateway.shutdown_flush 'Replay of spooled transcript message ... failed'
WARNING per append, on top of the per-append ERROR escalation that already
reports the outage. When the session already has recorded append failures the
replay failure is expected and now logs at DEBUG; the first drain (no recorded
failure yet) still warns, and a recovered DB still replays the spool in order.
2026-09-18 10:06:25 -07:00
teknium1
d0d004bce1 fix(gateway): count no-DB transcript appends and spool a stalled session's backlog to disk
A SessionStore whose state.db is unavailable (_db is None) early-returned from
append_to_transcript: no failure counter, no log, and the turn was gone. The
missing store now flows through the same serialized path (RuntimeError
'no owning session store ... deferring' is counted), so the WARNING->ERROR
escalation covers the reporter's outage shape.

Once a session hits the escalation threshold its in-memory backlog is moved to
the on-disk pending spool (recover_pending_to_db replays it at boot) instead of
sitting in memory until the 200-message cap or a crash. The spool is drained
BEFORE the next live write so recovery preserves transcript order.

Part of the 'best-effort durable write' atom of #114266.
2026-09-18 10:06:25 -07:00
Rook-CodeVolt
997d109289 fix(gateway): stamp FTS rebuild cooldown only on real attempts
Addresses review finding on PR #114301: _rebuild_fts_once was stamping
_fts_rebuild_last_attempt_at before the 'db is None or no rebuild_fts'
early-return guard, so a call with no usable DB burned the 5-minute
cooldown window without attempting anything. Move the stamp below the
guard so only real rebuild attempts start the cooldown.

Adds a regression test covering: no DB, DB without rebuild_fts, and the
follow-up real attempt succeeding immediately once a usable DB appears
(not blocked by a phantom cooldown). Full gateway/test_session.py suite:
67 passed.
2026-09-18 10:06:25 -07:00
Rook-CodeVolt
caba409bb8 fix(gateway): retry FTS index rebuild after cooldown instead of giving up
When the session FTS (full-text search) index rebuild fails, the
gateway previously entered a cooldown and then permanently stopped
retrying once the cooldown expired, leaving search silently broken
for the rest of the process lifetime. This adds a proper retry after
the cooldown window elapses, and escalates to ERROR-level logging
(previously missing 'import logging' meant this path could not even
log the failure) so repeated rebuild failures are visible to
operators instead of being swallowed.

Related but out of scope: while testing the fallback paths we
observed that the JSONL transcript fallback writer can write zero
bytes when the primary DB write path fails partway through - flagging
this for maintainers as a separate follow-up, not fixed here.

Fixes #114266
2026-09-18 10:06:25 -07:00
teknium1
c3de99bbf0 refactor(state): one carrier-aware user-turn rewind behind CLI /undo, /retry, gateway and TUI
CLI `_rewind_persisted_user_turn`, TUI `_rewind_active_session_history` and gateway
`rewind_session` each re-ran get_active_message_ids -> get_messages_as_conversation ->
split_user_originated_turn -> rewind_to_message with their own warm/durable comparison
helpers and three different out-of-range contracts (RuntimeError / ValueError / None).

The durable transcript is the authority for a rewind, so the implementation now lives
with the data: `SessionDB.rewind_user_turn` (hermes_state_rewind.py) with one typed
out-of-range error (`RewindTargetUnavailableError`). Surfaces keep only lock, eviction
and rendering glue and map that error to their own message.
2026-09-13 05:20:26 -07:00
kshitij
044a77b3b6 fix(gateway): failed-turn boundary keyed on the durable conversation tail; exception path classifies overflow before writing
Review findings on the salvage (all reproduced with a real SessionStore):

1. Primary persisted-agent path skipped the boundary. The agent's turn-start
   flush already persists the user row stamped with the inbound platform id, so
   `has_platform_message_id` saw THIS turn's own row, took the "duplicate" branch
   and skipped the whole block — including the new assistant boundary. The
   transcript stayed `[..., 'user']`, exactly the open tail #107070 is about.
   Fresh sessions hid it a second way: `session_meta` is appended after the
   agent-flushed user row, so a naive "newest row" tail read sees `session_meta`.
2. The exception fallback appended the boundary unconditionally; a redelivery of
   an already-closed turn produced `['user', 'assistant', 'assistant']`.
3. The exception fallback wrote the user row + boundary before classifying a
   400/500-on-long-session as overflow, growing a session that is already too
   large (the #1630 no-grow rule the persist path honours).

Fix: `SessionDB.latest_conversation_role()` (newest active row excluding the
`session_meta`/`system` bookkeeping rows the model never sees) behind
`SessionStore.transcript_tail_role()`, which resolves the same route
`load_transcript` reads via the existing `_compression_tip_for_session_id`.
One `_hmwa_close_failed_turn()` appends the boundary iff that tail is an open
user row; both the persist path and the exception fallback call it, so the
user-row dedupe no longer gates the boundary and a redelivery never stacks.
The overflow verdict in `_hmwa_agent_error_reply` is an early return ahead of
every transcript write. The `failed_turn_notice` kwarg and its dead
`or _hmwa_failed_turn_notice(...)` fallback are gone; the notice is derived
where each consumer needs it.

Tests (each red with the production change reverted, green here): boundary
keyed on the durable tail with the user write deduped (agent-flushed row →
closed; redelivery → nothing); fresh-session agent-flushed failed first turn
closed despite `session_meta` (real store); exception-path redelivery adds no
second boundary (real store, every lineage location, contract asserted from
store state); exception-path overflow persists nothing. Live E2E:
`evals/gateway_failure_ownership/probe.py` (real AIAgent + fixture provider)
20/20; the two `failed provider input` turns that previously left an open user
tail now close with the "not processed" row.
2026-09-12 13:00:29 +05:30
Teknium
3114916ee4 fix(gateway): carry accepted-input ownership through persistence
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.

Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
2026-09-07 14:11:18 -07:00
Teknium
136d80d040 fix(gateway): retain one durable owner for failed input turns 2026-09-07 14:11:18 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
e0a146c6b0 refactor(gateway): session_transcript — reuse _spool_dropped, assistant-only key table, suppress-based tip lookup 2026-09-02 23:59:14 -07:00
Teknium
f6c2c1a80a refactor(gateway): session mixins — compact docstrings/comments, fold short call sites (AST-identical) 2026-09-02 20:52:08 -07:00
Teknium
0f438d9e37 refactor(gateway): session — merge split log-format literals 2026-09-02 19:40:01 -07:00
Teknium
46258d89c5 refactor(gateway): session — rewrap comment/docstring paragraphs to 100 cols 2026-09-02 19:38:43 -07:00
Teknium
75fdd85316 refactor(gateway): session — pack exploded argument lists 2026-09-02 19:36:48 -07:00
Teknium
4f943a1cdd refactor(gateway): session — hug short parameter/argument lists 2026-09-02 19:33:00 -07:00
Teknium
e890c57a2b refactor(gateway): session — fold short multi-line signatures/calls 2026-09-02 19:29:01 -07:00
Teknium
ef4f0f4f61 refactor(gateway): session — checkpoint: phase helpers for get_or_create_session, shared clock/id helpers in lifecycle, context/state/stall folding 2026-09-02 19:21:38 -07:00
Teknium
d7bdf2788d refactor(gateway/session): split SessionStore into persistence/recovery/lifecycle/transcript mixins by call-graph cohesion; compact wire helpers 2026-09-02 16:15:42 -07:00