A gateway session whose summary model keeps timing out no longer retries
compaction on the same fixed interval forever.
The in-agent compressor already escalates repeat summary timeouts
60 -> 300 -> 900s (ContextCompressor.record_timeout_failure), but that ladder
reads the in-memory _consecutive_timeout_failures counter and
bind_session_state() zeroes it (context_compressor.py:1645). Session hygiene
constructs a FRESH AIAgent for every run (gateway/run.py:16820) and re-binds
state each time, so from the gateway that streak is structurally always 0 --
only the flat hygiene_failure_cooldown_seconds (300s) could ever be recorded.
Issue #79624 reported exactly that steady state: an oversized session
(1053 messages, ~119.5k tokens) whose aux model always timed out, re-attempting
compaction every 300s across five days until the reporter deleted the session
by hand.
Track the streak on PersistentState instead, which outlives the per-run agent
and is not cleared by turn/boundary resets, so consecutive hygiene failures
climb 300 -> 900 -> 2700s and then saturate. Both failure sites (progress
timeout and aborted compression) feed it; a real compression resets it, so a
session that recovers starts from the first rung again. The ladder multiplies
the configured base, so operators who tuned
hygiene_failure_cooldown_seconds keep their first rung. Per-session, so one
wedged chat cannot penalize other conversations.
Deliberately NOT changed, since each is a maintainer policy call rather than a
defect (all three are written up on #79624):
- no durable failure-streak column, so escalation still resets on restart
- the gateway 30s / in-agent 120s / aux-client 300s-floor timeout mismatch
- no `hermes doctor` check or `hermes sessions list` marker for a session
stuck in a compression-failure cooldown
Note the reported exit(1) is NOT a crash: it is the deliberate
_signal_initiated_shutdown path (gateway/run.py:26746-26751, #5646) that lets
systemd Restart=on-failure revive the gateway after a bare SIGTERM, and it
fires on every `systemctl restart` independently of compaction. The compaction
log lines appear after the shutdown line because the gateway-owned executor is
torn down with shutdown(wait=False, cancel_futures=True) (run.py:21164), so an
in-flight turn keeps logging during teardown. Full analysis on the issue.
Post-review hardening (Phase 2c + /simplify-code found five real defects in the
first cut):
- the recovery gate hand-rolled `_new_tokens < _approx_tokens` when a canonical
predicate already existed: `compression_made_progress` (agent/turn_context.py,
#39548). They disagree on 3 of 5 cases -- the hand-rolled form misses a
row-count win when the summary keeps the token estimate flat, misses one
where the summary is slightly MORE verbose (so a genuinely recovered session
would keep escalating forever), and counts a sub-5% wobble as recovery. Now
reuses the shared predicate, promoted from `_compression_made_progress` to a
public name with the old private name kept as a back-compat alias so the
existing importer (tests/agent/test_protected_tail_pressure_61932.py) and any
patcher of that symbol keep working.
- the reset was gated on "not aborted", but the degenerate "did not rotate or
compact in place" branch (#21301) is NOT aborted and yields zero reduction,
so a session wedged there reset its streak every run and could never
escalate -- silently defeating the fix. Now gated on real progress.
- no absolute ceiling: base * 9 reaches 9h at an operator base of 3600s,
indistinguishable from "compaction switched off". Added
_HYGIENE_COOLDOWN_MAX_SECONDS = 3600, mirroring the in-file
_RECONNECT_BACKOFF_CAP precedent.
- the reset used the get-or-create accessor to write a 0 that was already 0,
materialising a _sessions entry (never evicted). Now peeks.
- the abort verdict was probed twice, leaving the reset/record mutual
exclusion implicit; a future await between the probes would have broken it
silently. Computed once into _hyg_aborted.
Tests: 19 new in tests/gateway/test_hygiene_failure_cooldown_ladder.py --
ladder escalation, saturation, the absolute cap, per-session isolation,
reset-on-recovery, custom/zero base, PersistentState scoping (a mutation moving
the field to TurnState fails), degraded runners, the progress gate, the exact
progress-predicate semantics the gate depends on, and end-to-end that the
escalated value is what reaches the state DB. All 12 mutations caught, including
ones that restore the flat cooldown (the original bug), ungate the reset, swap
the canonical predicate back for the hand-rolled comparison, remove the cap, and
share the streak globally; the harness hard-errors when a mutation cannot be
applied, since a silently no-op mutation check is worse than none -- an earlier
version of it WAS silently no-opping after a refactor. The gate's contract test
slices by AST node span rather than a fixed character count, which had already
truncated once as the block grew. gateway hygiene + session-state + the three
touched agent compression suites: 50 passed; ruff clean.
E2E with real imports demonstrates the premise rather than asserting it:
bind_session_state zeroes the in-agent counter, and the recorded deadlines go
300 -> 900 -> 2700 -> 2700 -> 2700s where they were previously a flat 300s.
Reported by @yucezerey (#79624), whose state.db column dump and
"deleting the session fixed it" datapoint made the real mechanism findable.