Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
The #116420 retry loop retried on any unchanged transcript, and every retry adds
another stall to `_consecutive_timeout_failures`, so a regression that raises the
escalation threshold satisfied itself on a later iteration and the suite stayed
green. Re-assert the state the phase tests (one stall on the record) before each
attempt: retries stay legal, accumulated history can no longer reach a raised
threshold.
Also correct the comment naming the writer of the surviving backoff: on the
deterministic-commit path `on_timeout` never runs, so the asserted value is the
cancelled primary worker's stall_interrupted record, not a synchronous host arm.
`test_second_consecutive_stall_commits_the_deterministic_fallback_summary`
failed roughly a quarter of the time under parallel load, on current main and
on any branch: the second compression attempt never ran, so no fallback
summary was committed.
Instrumented order of events for one stall cycle:
MainThread record_timeout_failure('host compress_context timeout…', kind='stalled')
test clears the in-memory cooldown + the durable row
compress-ctx-timeout_0 record_timeout_failure('stall_interrupted:…', kind='stall_interrupted')
The host's arm is synchronous — it fires inside `_compress_context` before it
returns — but the fence-cancelled worker keeps unwinding on its own thread and
re-arms the same cooldown a few milliseconds AFTER the test clears it. The next
attempt then hits the transient guard, is refused, and the transcript is never
compacted. Widening the old 3s poll only moved the window.
The test no longer depends on which write lands last: it clears and retries
(bounded) until an attempt actually runs, and every assertion now keys on the
synchronous host arm. The durable-row assertion became the in-memory arm for
the same reason — the row is written by that asynchronous unwind, and one
unwind path (`rollback:<Exc>`) never writes it at all.
Two supporting hardenings in the same file: the stall stub polls its cancel
flag coarsely behind a 15s backstop instead of a 1ms spin behind a 6s one, and
the fence-ladder worker blocks on an Event instead of `sleep(0.3)`, which could
outrun the idle timer on a loaded box.
Test-only; no production file changes.
Measured with the same bursty load that reproduces it (the compressor suite run
repeatedly):
before (file as on main): 6/15 failing runs
after: 0/15 failing runs
second round, same load: 6/15 before, 0/15 after
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.