Commit Graph

66 Commits

Author SHA1 Message Date
liuhao1024
0bb27d37cf fix(state): declare origin_session_id in the canonical async_delegations schema
The delegation tool carries its own CREATE TABLE for async_delegations
(tools/async_delegation.py _initialize_schema) plus a lazy ALTER TABLE
ADD COLUMN for the tables it finds already existing. Its column list
had drifted ahead of the canonical SCHEMA_SQL: origin_session_id
(raw api_server session id of the originating request, the wake
self-post target) existed only through the tool's lazy path, so two
databases at the same schema_version had different
async_delegations shapes depending solely on whether the delegation
tool had ever run. Rebuild/replay pipelines that reconstruct state.db
from the canonical schema then hit the column with no version gate to
explain it (#94691).

Declare the column in SCHEMA_SQL with the same TEXT NOT NULL DEFAULT ''
shape the tool uses. Fresh installs now carry it canonically; the
declarative _reconcile_columns backfills it into legacy databases on
the next writable open (same pattern as earlier additive columns); the
tool's lazy ALTER keeps serving pre-reconciliation databases. The two
schema authorities now agree, pinned by a test that runs the tool's
initializer over a canonical database and asserts the shape is
unchanged.

Fixes #94691
2026-09-11 06:24:54 -07:00
Efe Büken
a239c4f811 fix(state): tolerate malformed session marker JSON 2026-09-11 06:24:54 -07:00
Siddharth Balyan
c22a8d8e3f Seeded sessions survive a gateway restart and store their seed once (tui_gateway) (#107549)
* fix(tui-gateway): a seeded session is durable at create, and its seed is written once

session.create accepts opening messages. Three defects sat in that path:

- A seeded session without a parent was never persisted at create, so a
  restart before the first prompt lost it and session.resume answered
  4007. Only branch children (#93959) were persisted up front. The
  same rationale applies to any seeded create: seeded content is
  intent, not an abandoned draft. Parentless seeds now persist their
  row, transcript and client title at create; empty drafts stay lazy.
- _coerce_seed_history dropped display_kind, so a seeded row tagged
  "hidden" (model-facing scaffolding) rendered as a user bubble. The
  coercion keeps "hidden" and only "hidden"; every other kind is
  stamped by the gateway at turn time and is not accepted from the wire.
- A branch child's seed was written twice: _seed_branch_row copied it at
  create but never marked it persisted, so the first prompt's
  _persist_branch_seed appended the copy again. The create path now
  sets _branch_seed_persisted, and the gate is a create-time `seeded`
  stamp instead of parent_session_id, so a resumed session (whose
  history comes from the DB) can never re-append its transcript.

Two invariant tests, both red on main: a parentless seed survives a
gateway restart with the hidden row kept out of the wire transcript and
not re-written by the first-submit path; a branch child's seed is stored
exactly once. The reasoning-fields fixture stamps `seeded`, the flag
session.create sets.

* fix(tui-gateway): a hidden seed row stays out of the list preview and the create count

Live-testing the seeded create on every surface showed two places where
the newly durable hidden row (display_kind="hidden") still surfaced:

- session.list built a session's preview from its first user row with no
  display_kind filter, so a hidden opening row (model-facing scaffolding
  the gateway never paints) became the sidebar preview. The preview
  predicate now skips hidden rows, in every listing query that shares it.
- session.create reported message_count as the raw seed length while its
  messages array already filtered the hidden row (2 vs 1). It now counts
  what is on the wire, the same rule session.resume applies.

Both are covered by the existing seeded-create test: the create count
equals the wire transcript, and the preview of a session whose first
user row is hidden is its first visible user row.

* fix(tui-gateway): a live unpersisted resume counts the wire transcript

session.resume on a live session that has no row yet reported message_count as
the raw history length while its messages array was already filtered, the same
mismatch the previous commit fixed on session.create. Count the wire, as the
cold, deferred and reuse-live resume paths already do.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-10 17:55:24 +00:00
Totoro-qaq
9cd1453f26 fix(sessions): a stale explicit-close stamp no longer makes a session permanently uncompressible
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
2026-09-10 03:04:51 -07:00
Xipong
1c6683e8e0 fix: index compacted display identity writes 2026-09-09 10:05:59 -07:00
Xipong
49e6d661a0 fix: bound compacted display history paging 2026-09-09 10:05:59 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
2b55ded1ac perf(state): keep delegate-child transcripts out of the trigram FTS index (schema v30)
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.

Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.

The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
2026-09-03 02:35:37 -07:00
Teknium
c36d8a692e refactor(state): drop redundant isinstance guard in _lock_holder_provably_dead (except path already fails closed) 2026-09-02 23:10:39 -07:00
Andrew Wikel
5a7bee0fa8 fix(state): exclude tool calls from trigram FTS
Keep structured tool_calls searchable through the standard FTS index while
removing their repetitive JSON from the trigram projection. Reuse the
existing optimize-storage rebuild path for deployed v1 layouts.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-03 11:35:13 +05:30
Teknium
5462f592a3 refactor(state): fold lock-release and /proc parsing in hermes_state_common; trim restating comments 2026-09-02 23:04:23 -07:00
Teknium
c4a6c8fbc6 refactor(state): compact hermes_state_common comments and docstrings, keeping every invariant and failure mode 2026-09-02 22:56:57 -07:00
Teknium
7f24f2f987 refactor(state): collapse multi-line SQL assembly and log calls in hermes_state_common to one-liners 2026-09-02 22:48:44 -07:00
Teknium
36b7fe14ee refactor(state): hand-compact remaining long docstrings (all invariants kept) 2026-09-02 20:12:31 -07:00
Teknium
28eb24eb3a refactor(state): reflow full-line comment paragraphs to the 108-col budget (word-preserving) 2026-09-02 19:59:11 -07:00
Teknium
f731c63e89 refactor(state): drop blank separators around nested _do txn closures 2026-09-02 19:47:36 -07:00
Teknium
3c615c48e9 refactor(state): AST-neutral bracket/string-literal layout pass on the six mixin modules 2026-09-02 19:45:18 -07:00
Teknium
4de8710b74 refactor(state): unify _placeholders/_ended_by_compression/row-probe SQL into hermes_state_common 2026-09-02 19:43:49 -07:00
Teknium
7d48a84acf refactor(state): reflow prose docstrings to the 108-col budget (word-preserving) 2026-09-02 19:39:50 -07:00
Teknium
511c7fb9b7 refactor(state): compact common preview/lineage helpers and comments 2026-09-02 18:49:10 -07:00
Teknium
b8bc47032b refactor(state): compact common lock helpers (shared lock-file rewrite, merged holder-record parse) 2026-09-02 18:43:04 -07:00
fangliquanflq
ea65fcd980 perf(state): exclude cron sessions from trigram FTS 2026-09-03 05:08:22 +05:30
fangliquanflq
57162d0cc1 fix(state): bound FTS indexing for large tool results 2026-09-03 05:03:04 +05:30
Teknium
a138533157 refactor(state): fold docstring closers (whitespace only) 2026-09-02 16:21:39 -07:00
Teknium
b422113061 refactor(state): compact long rationale comments (all rules kept) 2026-09-02 16:20:35 -07:00
Teknium
acb12608e4 refactor(state): squeeze blank-line runs between constants (AST-neutral) 2026-09-02 16:19:09 -07:00
Teknium
18946a781f refactor(state): join short adjacent string literals (AST-neutral) 2026-09-02 16:16:37 -07:00
Teknium
95714f6d93 refactor(state): AST-neutral line packing; derive v22 session_model_usage DDL from the heal DDL 2026-09-02 16:13:05 -07:00
Teknium
fc2972b299 refactor(state): common — shared freshest-of/after-marker SQL builders, msvcrt lock helper, compact preview constants 2026-09-02 16:11:04 -07:00
Jason Wu
98a6e4a90e fix(session-search): bound recent-session browsing
Preselect indexed recent candidates before rich hydration, cap and deduplicate compression lineage traversal, and interrupt sustained SQLite work through a cooperative progress deadline. Fail closed when the bounded browse API is unavailable and cover legacy-schema reconciliation plus malformed lineage cases.
2026-09-03 04:23:35 +05:30
Teknium
d15c61b5dc refactor(state): split SessionDB into domain mixins and free-function modules; unify SQL boilerplate
hermes_state.py 17,220 -> 6,442 LOC. Behavior-neutral: every moved body is
AST-identical to the original, verified per extraction.

SessionDB core
- _write_sql / _write_rowcount / _read_one / _read_all replace ~120 copies of
  the `def _do(conn): conn.execute(...)` + `_execute_write(_do)` and
  `with self._read_ctx() as conn: row = conn.execute(...).fetchone()` shapes.
- _set_lineage_column replaces four copies of the recursive compression-lineage
  UPDATE (archived / pinned / hidden / last_read_at).
- _read_session_number unifies the three compression counter readers.
- Dead (zero refs repo-wide): restore_rewound, delete_gateway_routing_entries,
  _is_duplicate_replayed_user_message, SessionPortabilityMixin.get_first_assistant_text.

New mixins bound onto SessionDB via the MRO (logger name stays "hermes_state"):
  hermes_state_messages    SessionMessagesMixin       48 methods
  hermes_state_compression SessionCompressionMixin    30
  hermes_state_gateway     SessionGatewayMixin        26
  hermes_state_maintenance SessionMaintenanceMixin    13
  hermes_state_usage       SessionUsageMixin          12
  hermes_state_titles      SessionTitlesMixin         13
  hermes_state_telegram    SessionTelegramTopicsMixin 11
Origin-internal symbols resolve through a lazy `from hermes_state import ...`
inside the few methods that need them (no import cycle).

New free-function modules, every name re-imported into hermes_state so
`hermes_state.<name>` (and test monkeypatches on it) keep working; intra-module
calls to patched helpers go through the lazy origin import:
  hermes_state_repair   repair/backup/preflight (43 defs)
  hermes_state_wal      journal-mode / PRAGMA policy (33 defs)
  hermes_state_dbfile   header probes, zeroed-db quarantine, stats, holders (21 defs)

Existing mixins: search — shared FTS MATCH/LIKE builders, unified rebuild
status/step/finish engines, state_meta helpers; schema — one legacy/v23 FTS init
branch, shared _live_pk_columns, Row/tuple dual access dropped; portability —
shared _PREVIEW_RAW_SUBQUERY_SQL and _rich_row; common — single
stat_db_file_identity (was 3 copies), AUTO_VACUUM_MIN_FREELIST_RATIO.

Docstrings/comments hand-compacted (AST-identical) keeping every invariant,
ordering rule, failure mode and WHY. Schema SQL, migration order and PRAGMAs
untouched. test_repair_path_has_no_bare_connects repointed to hermes_state_repair.
2026-09-02 13:32:13 -07:00
Teknium
8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
teknium1
fd05029430 fix(state): fail fast on non-contention flock errors and retry deferred FTS rebuilds in-process (salvage #100130)
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:

* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
  EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
  ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
  environment failures that polling cannot fix — `_acquire_db_flock` and
  both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
  now defer immediately with the real errno instead of burning the full
  120s / holder timeout and then logging a fake "held by another process".

* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
  open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
  lock) stayed `_fts_stale` — LIKE-only search — until the process
  reopened state.db. Short-lived CLIs reopen every run; the gateway opens
  once and stays up for days, so the deferral was effectively permanent
  (#100108). The retry runs from the EXISTING gateway housekeeping tick
  (`_start_gateway_housekeeping`, 60s) against the shared SessionDB
  instances via `hermes_state_registry.live_shared_session_dbs()`:
  non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
  bounded backoff 60s -> 1h, no new thread, still fails closed on live
  holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.

* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
  line can be matched to the interpreter that actually linked it
  (#100108 point 3).

Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.

Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-02 04:15:02 -07:00
Teknium
238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
HexLab98
f5f4cefc3d fix(state): fail closed when a state.db admission lock file cannot be opened
state.db has two cross-process admission authorities gating destructive work
on a file several Hermes processes share: fts_rebuild_admission for full
structural FTS rebuilds, and _cross_process_repair_lock for writable_schema
surgery / VACUUM. Both document themselves as fail-closed, and both honoured
that only for a timed-out acquire. When the lock file could not be open()ed at
all they yielded True and proceeded "with in-process serialisation only" —
which is no cross-process authority whatsoever.

That inversion is reachable exactly when it does the most damage. Creating the
lock file needs a directory entry and an inode, so on a full disk open() raises
ENOSPC — while a sibling that opened ITS handle before the disk filled is still
mid-rebuild or mid-surgery. Every process then ran concurrent destructive work
on the same live DB: precisely the interleaving PR #93200 added these locks to
prevent, and the shape reported in #100368 (disk-full trigger, then a fresh
corruption on every boot with other writers alive, and no re-corruption on a
boot with zero other writers).

Both helpers now yield False on OSError. This routes the error into the
outcome the locks already define and every caller already handles: rebuild_fts
returns 0, _recover_stale_fts leaves canonical writes plus LIKE search
available behind the retryable stale breadcrumb, the startup path detaches FTS
triggers, and repair_state_db_schema re-probes and reports. Nothing reachable
is lost — on a read-only directory the rebuild's and the repair's own writes
could not have committed either. The repair report's error string now names
both ways the authority can be missing, since operators read it directly.
2026-09-01 23:58:51 -07:00
Teknium
894fc35337 fix(state): break provably-orphaned repair/FTS-rebuild locks left by dead holders (#100108) 2026-09-01 10:52:06 -07:00
joaomarcos
6b9b3e0145 chore(cache): take the pre-merge cleanups on the declared conversation scope
@teknium1's maintainer-side review found no blocking defect on 09004753c9 and
listed five cleanups. All five are here.

1. scratch/repro_96811.py is deleted. It would have landed on main as a
   tracked file: scratch/ is not gitignored and has never existed on main, so
   this PR was creating the directory. Nothing referenced the probe, and
   TestConversationGenerationRotates / TestGenerationSurvivesPruning /
   TestPeerIdentityIsSourceQualified already carry all four of its stages, so
   it is dropped rather than parked under tests/.

2. Upgrade notes are written into this commit body (below) and the PR body.
   There is no committed changelog to add them to: scripts/release.py
   generates .release_notes.md from commit SUBJECTS at release time, and
   .gitignore keeps that file out of the tree.

3. declared_conversation_scope() now reads the sessions row ONCE. The fork
   verdict and the source the peer queries match on both live on that row, and
   asking for them separately read it twice per resolution. The new
   SessionDB.declared_scope_identity() returns the pair and keeps the marker
   rules beside is_explicit_fork_child() instead of re-implementing them in the
   caller. A SessionDB that does not expose the combined view keeps the
   original two-call path, so nothing that predates it changes behaviour --
   including the three doubles that certify the fail-closed contract, which are
   untouched. TestOneIdentityReadPerResolution pins the single read, the
   two-call fallback, the fail-closed degrade and the fork refusal; removing
   the fold turns the first of those red.

   The third read stays: the generation lives in conversation_generations, a
   different table, and cannot be folded into a sessions lookup.

4. _declared_conversation_session() documents the concurrent first-turn race.
   Two simultaneous first requests on one declared key can each miss the
   lookup, mint a row and both bind, because each row is unkeyed at bind time
   and the mismatch guard does not fire. That converges rather than crossing:
   both rows carry the same key under the same source, so the lookup returns
   the later one for every subsequent reply and the earlier row is an abandoned
   transcript, never another conversation's identity.

   The same docstring still claimed the generation was durable in
   sessions.end_reason and that "nothing here needs a counter". That stopped
   being true in 09004753c9, which moved the generation into
   conversation_generations precisely because deriving it from prunable session
   rows was ABA. Corrected, along with the same stale sentence on
   TestConversationBoundariesRotate.

5. conversation_generations rows are now documented as deliberately never
   collected, rather than merely uncollected. Dropping one resets that peer to
   "no generation", so its next boundary writes 1 again and re-issues a gwk_
   scope a retired conversation already used -- the exact ABA the table exists
   to close. Worth stating because the repo already carries both patterns a
   maintainer would extend: delete_session() cascades to messages, and
   gateway_hygiene_state is already swept by session_key.

Upgrade notes, one-time on merge:

- One cold prompt-cache bucket per keyed conversation. Every gateway platform
  declares gateway_session_key, so each keyed conversation's affinity scope
  moves once from its compression-lineage root session id to the gwk_ hash.
  One cache miss per live conversation, on its next turn only.
- hermes status counts more sessions. A declared API conversation is now
  recorded as a keyed row and appears in "Active: N session(s)" where it was
  invisible. Those sessions already existed; only their visibility changes.
- A database upgraded mid-conversation starts with no generation and takes its
  first from the next boundary written, so a conversation that reset before the
  upgrade shares its predecessor's scope once. One warm bucket, never a crossed
  identity.

Verified on this head: 55 in test_declared_conversation_scope.py (51 + 4 new),
33 in test_prompt_cache_scope.py, 49 in test_api_server_declared_conversation.py,
25 in test_api_server_runs.py, 109 in test_api_server.py, 12 in
test_cross_process_turn_lease.py, and 526 across test_hermes_state.py +
tests/hermes_state/ + tests/state/. ruff clean.

Found in review by @teknium1.

Refs #96811
2026-09-01 02:14:35 -07:00
joaomarcos
832d68aba4 fix(cache): repair settlement, and make the generation unprunable
Four blockers from @andrexibiza's reviews of 28a2d7f0ee and dc7865765c. The
first two are defects I introduced in 99f2d4394f by replacing the wrong
occurrence of an identical call site.

1. _run_agent raised NameError on every opted-in declared bind. Its worker
   finally evaluated `if _declared_selected:`, a local of _handle_responses /
   _handle_runs that is neither a parameter nor an enclosing binding here, so
   the successful declared-key paths failed at settlement after the agent run.
   bind_declared_conversation already IS the gate; the inner name is gone.

2. /v1/runs never received the gate at all -- it landed on _run_agent instead.
   _run_sync bound unconditionally, so an explicit body session_id that existed
   with an empty session_key was adopted by the header key even though the
   header lost precedence. It now carries the same gate.

3. COUNT(*) + MAX(ended_at) over session rows cannot prove non-reuse.
   delete_session() deletes the selected row and bulk prune selects ended rows,
   so the aggregate can return a pair it already emitted:
   (1,T1) -> (2,T2) -> delete boundary B -> (1,T1), handing a new conversation
   a retired affinity identity. The backwards-clock shape needs no pruning at
   all. The generation now lives in a conversation_generations table keyed by
   (source, session_key), advanced by _bump_conversation_generation inside the
   same transaction that writes each boundary -- outside prunable session
   history, wall-clock-free, and increment-only. end_session() and
   promote_to_session_reset() both advance it, and only when they actually
   wrote a boundary, so a repeated end cannot double-count.

4. The carrier could be memoized under the wrong source. _agent_source() fell
   back to agent.platform before the row landed while persistence uses
   _session_source_for_agent(), which honors HERMES_SESSION_SOURCE. Because a
   declared scope is non-None immediately, resolve_prompt_cache_scope memoizes
   it and never re-resolves once the authoritative row appears, so under an
   override both sides of a /new read the platform domain and hashed the same
   scope. The pre-row path now uses the persistence resolver itself.

Coverage answers the review's specific objection that mocked tests proved the
mock rather than the path. TestRealRunAgentSettlement stubs _create_agent and
lets the real _run_agent settle; the /v1/runs case persists an unkeyed explicit
row and waits for the worker to retire before asserting. Both were verified by
mutation: reinstating the inner name fails two of them, and removing the
/v1/runs gate fails the explicit-session one. The first version of that test
passed with the gate removed -- it asserted before settlement -- and would have
been the same empty proof the review called out.

TestGenerationSurvivesPruning covers deleting the newest boundary, deleting
every boundary, the backwards-clock-then-prune shape, compression and
accidental ends not advancing it, repeated ends not double-counting, promotion
advancing it, unkeyed rows advancing nothing, and peer scoping.
TestSourceOverrideDomain covers the override across a reset.

Found in review by @andrexibiza, whose analysis located each of these
defects and specified what a correct fix had to prove.

Refs #96811

Co-Authored-By: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-01 02:14:35 -07:00
Teknium
c0667439ec fix(compression): rotation heals stale automatic ended_at stamps instead of wedging (#88197)
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).

Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.

TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
  rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
  session_reset (deliberate boundary) => still aborts BEFORE the #47202
  flush, preserving #88411's no-growth contract where an abort remains
  correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
2026-08-31 09:58:11 -07:00
Finn763
cae58be1f5 fix(state.db): cross-backend heartbeat gates orphan sweep
A live sibling serve sharing state.db is no longer treated as a dead process by the startup orphan sweep.

Covers the sweep half of #94895. The launchd Errno 48 KeepAlive loop is not addressed here.

Credit: @Finn763
2026-08-27 11:31:59 -05:00
Jan-Stefan Janetzky
1104ffe0b9 feat(memory): opt-in fail-closed pre-compress checkpoint contract (API v1)
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.

This adds an opt-in, provider-agnostic checkpoint contract:

- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
  by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
  historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
  on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
  failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
  (default false, documented in cli-config.yaml.example). When enabled,
  compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
  uncompressed transcript is preserved) unless a checkpoint-capable provider
  confirms the durable checkpoint. Providers receive normalized direct
  user/assistant evidence: tool rows, system messages, tool-call wrappers,
  and prior compaction summaries are filtered host-side into one stable
  contract. codex_app_server compaction is rejected under the gate because
  it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
  migration via _reconcile_columns) so summary provenance survives process
  restarts; only the resume model history carries the marker, keeping
  get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
  (skip_memory=False) so a required checkpoint also guards those rewrites.

The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.

Refs #93986
2026-08-25 03:55:55 -07:00
fangliquanflq
8f3a82f96a fix(state): recover FTS after orphan holder deferrals 2026-08-23 20:01:41 -07:00
Teknium
9d0727d49b fix(state): single fail-closed cross-process authority for all full FTS rebuilds
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:

- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
  30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
  _rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()

Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.

Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
2026-08-23 19:00:36 -07:00
Teknium
47d6ce78a2 fix(tui): make startup_orphan_reap recoverable and move its config onto dashboard.*
Follow-up to the #65422 salvage:

- startup_orphan_reap joins _RECOVERABLE_END_REASONS (kept distinct from
  ws_orphan_reap for forensics): every recovery fence
  (find_latest_gateway_session_for_peer, unarchive_recoverable_session,
  promote_to_session_reset) now treats a startup-swept row as an
  accidental end, so a sweep never makes a session unresumable.
- Config key moves from sessions.orphan_reaper to
  dashboard.startup_orphan_sweep in DEFAULT_CONFIG, next to its siblings
  ws_ping_interval / ws_ping_timeout / ws_orphan_reap_grace_s; the raw
  loader in tui_gateway.server reads the new key (fail-open on missing).
  cli-config.yaml.example and website/docs/user-guide/configuration.md
  follow the dashboard.* documentation pattern.
- New regression test: a stranded 'active' row (ended_at NULL, no live
  runtime) is swept AND still recoverable via peer-keyed lookup and fully
  revivable via reopen_session afterward.
2026-08-23 18:58:40 -07:00
Teknium
4aa162b30d fix(gateway): cancel pending WS-orphan reaps on resume and supersede stale runtimes quietly
Storm killer for the reap->broadcast->auto-re-resume feedback loop:

- New _pending_ws_reaps registry (sid -> Timer): _schedule_ws_orphan_reap
  registers, _reap pops, and _cancel_ws_orphan_reap(sid) is called from
  every resume/reuse/rebind path — the session.resume fast-path reuse
  (methods_session.py), _claim_or_reuse_live winners, and the
  _live_session_payload live-transport rebind.
- When a resume mints a fresh runtime for stored session id S, any prior
  runtime for S still parked on the detached-WS sentinel is claimed under
  the resume lock, its reap Timer cancelled, and the record finalized
  quietly with end_reason superseded_by_resume — NOT in
  _RECLAIM_END_REASONS, so no session.reclaimed broadcast fires and the
  client's auto-re-resume can't storm.
- superseded_by_resume added to _RECOVERABLE_END_REASONS in
  hermes_state_common.py so canonical Bot Chat resurrection still applies.

Unit tests: resume cancels the reap timer, superseded runtimes finalize
without a reclaimed broadcast, and the normal orphan reap still fires
when nobody re-resumes.
2026-08-23 17:43:39 -07:00
kshitij
4865194772 fix(bot-mode): review follow-ups for recoverable-archive resurrection
- Clear the accidental end stamp on resurrection (at the lineage tip):
  a surviving ws_orphan_reap/agent_close reason made a LATER deliberate
  archive auto-resurrect on the next lookup — the user could never retire
  the canonical chat. Test pins the resurrect -> deliberate-archive ->
  stays-archived cycle.
- Judge recoverability at the compression TIP: the registry row of a
  compressed lineage carries end_reason='compression', so tip-stamped
  accidents were unrecoverable through the registry row. Lineage test.
- Heal the third lookup: the api_server exact-title listing (hermes peer
  dm resolution) filtered archived rows out via list_sessions_rich and
  still failed for reap-archived canonical chats.
- Single source of truth for the recoverable set: tuple moved to
  hermes_state_common (mirroring _RESET_END_REASONS_SQL) and interpolated
  into all three recovery SQL sites — literals cannot drift.
- methods_session gate uses BOT_CHAT_TITLE (not a literal) and re-fetches
  by id after resurrection (title has no DB-level UNIQUE).
- Idempotence pinned: two consecutive profiles.list calls both resolve.
2026-08-24 03:27:11 +05:30
abundantbeing
a2a23a8f7e fix(clients): hide compaction carriers across surfaces 2026-08-21 21:53:11 +05:30
fangliquanflq
39d2d858fd fix(gateway): persist hygiene failure cooldown rung 2026-08-16 02:02:06 -07:00