A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request (even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the suggested recovery reproduced the same loop. - request_exceeds_model_window(agent, tokens): one predicate, two consumers. - Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does exactly this every turn); the summary-failure cooldown stops the retry from repeating. - Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline). compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript. - Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars, prompt build ms) so a stalled attempt is distinguishable from a slow prompt build. - Deterministic-rung wording no longer claims "again after a stall backoff". Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True, compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed (main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
61 lines
2.8 KiB
Python
61 lines
2.8 KiB
Python
"""Summary-hook dispatch and cancellation rollback for context compression."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import inspect
|
|
from typing import TYPE_CHECKING, Any, Dict, List, Optional
|
|
|
|
from agent.auxiliary_client import AuxiliaryExplicitCancellation
|
|
|
|
if TYPE_CHECKING:
|
|
from agent.context_compressor import _HandoffScan
|
|
|
|
|
|
def _accepts_keyword_argument(callable_obj: Any, name: str) -> bool:
|
|
"""Return whether an inspectable callable accepts ``name`` as a keyword."""
|
|
try:
|
|
parameters = inspect.signature(callable_obj).parameters
|
|
except (TypeError, ValueError):
|
|
return False
|
|
if any(parameter.kind is inspect.Parameter.VAR_KEYWORD for parameter in parameters.values()):
|
|
return True
|
|
parameter = parameters.get(name)
|
|
return parameter is not None and parameter.kind in (
|
|
inspect.Parameter.POSITIONAL_OR_KEYWORD,
|
|
inspect.Parameter.KEYWORD_ONLY,
|
|
)
|
|
|
|
|
|
class SummaryDispatchMixin:
|
|
def _summarize_window(
|
|
self, messages: List[Dict[str, Any]], turns_to_summarize: List[Dict[str, Any]], scan: "_HandoffScan",
|
|
focus_topic: Optional[str], memory_context: str, bypass_cooldown: bool,
|
|
) -> Optional[str]:
|
|
"""Run the summary LLM; a cancellation rolls back the handoff scan's self-heal mutation first.
|
|
A deterministic pin (repeated stall, #112420) skips the LLM: ``None`` lets Phase 3 insert the static
|
|
fallback summary, or abort under ``abort_on_summary_failure`` exactly like a failed summary call."""
|
|
from agent.context_compressor import take_deterministic_summary_pin
|
|
if take_deterministic_summary_pin():
|
|
# Surfaces through the fallback summary's reason line and the host's one-shot user warning.
|
|
self._last_summary_error = (
|
|
"summary model stalled on every route; deterministic fallback summary inserted"
|
|
)
|
|
telemetry = getattr(self, "_active_compression_telemetry", None)
|
|
if isinstance(telemetry, dict):
|
|
telemetry["failure_class"] = "stall_deterministic_fallback"
|
|
return None
|
|
# Focus-topic derivation scans user turns; only pay when a summary is generated.
|
|
summary_kwargs: Dict[str, Any] = {
|
|
"focus_topic": focus_topic or self._derive_auto_focus_topic(messages),
|
|
"memory_context": memory_context,
|
|
}
|
|
if _accepts_keyword_argument(self._generate_summary, "bypass_cooldown"):
|
|
summary_kwargs["bypass_cooldown"] = bypass_cooldown
|
|
try:
|
|
return self._generate_summary(turns_to_summarize, **summary_kwargs)
|
|
except AuxiliaryExplicitCancellation:
|
|
# Cancellation is a true no-op: restore the scan's mutation before the exception escapes.
|
|
self._previous_summary = scan.previous_summary_before
|
|
self._summary_has_user_turn = scan.has_user_turn_before
|
|
raise
|