A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request (even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the suggested recovery reproduced the same loop. - request_exceeds_model_window(agent, tokens): one predicate, two consumers. - Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does exactly this every turn); the summary-failure cooldown stops the retry from repeating. - Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline). compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript. - Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars, prompt build ms) so a stalled attempt is distinguishable from a slow prompt build. - Deterministic-rung wording no longer claims "again after a stall backoff". Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True, compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed (main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
40 lines
1.5 KiB
Python
40 lines
1.5 KiB
Python
"""Regression coverage for stalled preflight compression."""
|
|
|
|
from types import SimpleNamespace
|
|
|
|
import pytest
|
|
|
|
from agent.turn_context import (
|
|
PreflightCompressionTimedOut,
|
|
_fail_closed_after_preflight_timeout,
|
|
)
|
|
|
|
|
|
def test_preflight_timeout_blocks_unchanged_provider_payload():
|
|
"""Unknown window (no compressor): the conservative #98424 default — never send blind."""
|
|
agent = SimpleNamespace(_last_compression_timed_out=True)
|
|
|
|
with pytest.raises(PreflightCompressionTimedOut, match="provider call was not sent"):
|
|
_fail_closed_after_preflight_timeout(agent, 190_035)
|
|
|
|
|
|
def test_structural_noop_keeps_existing_preflight_behavior():
|
|
agent = SimpleNamespace(_last_compression_timed_out=False)
|
|
|
|
_fail_closed_after_preflight_timeout(agent, 190_035)
|
|
|
|
|
|
def test_preflight_timeout_sends_a_request_that_fits_the_model_window():
|
|
"""Over the compression threshold but under the window (#113646): the turn runs uncompressed, exactly
|
|
like the cooldown-blocked path sends it; above the window it still fails closed."""
|
|
fits = SimpleNamespace(
|
|
_last_compression_timed_out=True, context_compressor=SimpleNamespace(context_length=120_000),
|
|
)
|
|
_fail_closed_after_preflight_timeout(fits, 99_347)
|
|
|
|
over = SimpleNamespace(
|
|
_last_compression_timed_out=True, context_compressor=SimpleNamespace(context_length=1_000_000),
|
|
)
|
|
with pytest.raises(PreflightCompressionTimedOut):
|
|
_fail_closed_after_preflight_timeout(over, 1_313_423)
|