The why (retry can win vs. deferred compression_exhausted wipe, and why the
streak is durable) was spelled out in full at the constant, in
_on_summary_failure, in record_completed_compaction and in the durable
getter docstring. Keep it once on _CONSECUTIVE_OVERLOAD_ABORT_ESCALATION
and cut the other copies to one line, matching the sibling fallback-streak
getter. Comment-only.
The durable overload budget was a read-modify-write (memory += 1, then
set the column), so two agents compressing the same session concurrently
could each write N+1 and lose a strike, delaying the fallback escalation.
Bump the column with one UPDATE ... RETURNING inside _execute_write and
take the returned row value as authoritative; memory-only counting is
kept for unbound compressors. No try/except or legacy-setter shim: the
existing _durable_read plumbing already degrades on unsupported DBs.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Review P1 on #123186 (ehz0ah): the N=3 sustained-overload budget was
object-local. The gateway binds a fresh compressor to the same session on
every turn / cache eviction, and restart / API-server requests construct
one too, so each fresh instance restarted the budget at zero and a
sustained summary-provider outage walked the session back into
compression_exhausted and auto-reset — the exact wipe the PR exists to
prevent. Reproduced by the reviewer with three fresh compressors bound to
one session: every attempt ended at counter=1.
Persist the streak as sessions.compression_overload_streak through the
same durable channel the fallback streak and recovery deadline already
use (#100185): SessionCompressionMixin get/set pair over the sessions
row, loaded in the compressor's durable-load block, written on every
change (overload abort, successful summary, committed boundary, runtime
model switch).
- Compression rotation carries the streak to the child row at the
boundary, same as the fallback streak: the parent's value is read
before the bind, re-applied, and persisted onto the fresh child row.
- A completed compaction boundary settles the budget to 0 — including
the committed degraded fallback, so a recovered provider regains a
full budget instead of the session staying degraded forever.
- Schema event appended to SCHEMA_HISTORY["sessions"] (seq 28, after
compression_recovery_deadline) so salvage replay maps the new column.
- Fresh-instance regression: two aborts on one compressor, then a fresh
compressor bound to the same session inherits streak=2 and the third
session-wide attempt commits the fallback; plus a rotation carry-over
test. Both fail red against the memory-only implementation.
(cherry picked from commit 3742fee8d03d7077d23b0dac3a3c79fe48e102ac)
Review follow-up on the byte-identical tools[] pin.
- The pin records the code identity that built it (checkout/build sha, else
the release version). Written by the same code, every pinned tool that is
still available keeps its pinned bytes, including tools whose parameters
are derived per surface (delegate_task, text_to_speech, memory, patch).
The per-tool "parameters differ -> take current" rule replaced those bytes
on every surface hop and rewrote the ~44KB pin each time. A pin from other
code (`hermes update`, legacy name lists) takes the current definitions
once and is re-pinned.
- A pinned tool this process did not build is carried forward only while
this agent's toolset selection allows it (enabled minus disabled toolsets
and role reservations, before check_fn). It must also pass the session
schema gates on the merged array, so browser_exec never comes back once
terminal is gone. Client-surface toolsets (desktop_ui, project) still
carry across hops: no config choice removed them there.
- The rotation compaction child inherits the parent's pin in the publish
transaction.
- `hermes sessions recover` keeps pin rows in its system_prompts sweep and
clears dangling pin hashes, as lost-and-found now does too. Profile moves
carry the pin like the prompt. A continuing session whose pin is missing
or unreadable (a row swept by an older build) pins the tools it sends on
that turn, so later hops stay stable.
_NON_CONTINUATION_CHILD_FILTER_SQL now excludes a child whose
_reset_from marker is bound to the queried parent, like the branch and
delegate markers, so find_live_compression_child never recovers to (or
fails closed because of) a reset fork and a lone reset child no longer
blocks reopen_orphaned_compression_session (#114271).
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.
Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.
Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.
Part of #114271
A reset fork child (model_config._reset_from, or the legacy same-key
heuristic) is a separate user-visible conversation that already lists
as its own row, but the compression chain step only excluded
branch/delegate/tool children. When a reset sibling ended later than
the real continuation it won the last_active tiebreak, so the lineage
tip projection landed on it: the newest session became invisible in
every list and the reset sibling showed twice (#114271).
Exclude _RESET_CHILD_SQL children in both forward-chain walkers
(get_compression_chain's _CHAIN_STEP_SQL and the order_by_last_active
CTE in list_sessions_rich), mirroring the resume path's exclusion.
The publish-side taxonomy (_end_stamp_class, _compression_parent_obstacle,
compression_parent_deliberately_ended) and the agent-guard delegation produced
exactly main's verdict -- every non-automatic stamp still fails closed -- so
they only changed an error string. Inline the "explicit close with no
continuation" test into reopen_if_explicitly_closed(), restore main's publish
branch and agent guard untouched, and keep two tests: the field shape rotates
after the host clears the stamp; boundary/compression/automatic stamps and a
session already claimed for teardown are never cleared.
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
Review follow-up for #106276: the compensating UPDATE can succeed and the
session row can still be retired before the verification read-back runs.
That window raised the same RuntimeError as a genuine restore mismatch.
Accept the absent session (warn + return) exactly like the pre-update
deletion window, keep strict verification for a surviving row, and add a
regression test that deletes the session after the UPDATE commits.
[salvage: test hunk dropped from this pick, see previous commit]
The exact-row rollback in restore_compression_failure_cooldown_row raised
RuntimeError when its UPDATE hit rowcount == 0, so a session row retired or
expired mid-attempt (e.g. by the maintenance sweep) crashed the turn
dispatcher during the compression summary failure path (#106271). A missing
session row means the cooldown died with it: nothing is left to restore, so
treat rowcount == 0 as a tolerated no-op (warn + return, skipping read-back
verification), mirroring the existing early-return for snapshots taken while
the session already did not exist. Write and verification failures still
propagate.
[salvage: contributor test file dropped from this pick; regression tests are carried from #106277 and trimmed to two invariants]
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.