Files
hermes-agent/agent/context_compressor_summary.py
teknium1 303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00

61 lines
2.8 KiB
Python

"""Summary-hook dispatch and cancellation rollback for context compression."""
from __future__ import annotations
import inspect
from typing import TYPE_CHECKING, Any, Dict, List, Optional
from agent.auxiliary_client import AuxiliaryExplicitCancellation
if TYPE_CHECKING:
from agent.context_compressor import _HandoffScan
def _accepts_keyword_argument(callable_obj: Any, name: str) -> bool:
"""Return whether an inspectable callable accepts ``name`` as a keyword."""
try:
parameters = inspect.signature(callable_obj).parameters
except (TypeError, ValueError):
return False
if any(parameter.kind is inspect.Parameter.VAR_KEYWORD for parameter in parameters.values()):
return True
parameter = parameters.get(name)
return parameter is not None and parameter.kind in (
inspect.Parameter.POSITIONAL_OR_KEYWORD,
inspect.Parameter.KEYWORD_ONLY,
)
class SummaryDispatchMixin:
def _summarize_window(
self, messages: List[Dict[str, Any]], turns_to_summarize: List[Dict[str, Any]], scan: "_HandoffScan",
focus_topic: Optional[str], memory_context: str, bypass_cooldown: bool,
) -> Optional[str]:
"""Run the summary LLM; a cancellation rolls back the handoff scan's self-heal mutation first.
A deterministic pin (repeated stall, #112420) skips the LLM: ``None`` lets Phase 3 insert the static
fallback summary, or abort under ``abort_on_summary_failure`` exactly like a failed summary call."""
from agent.context_compressor import take_deterministic_summary_pin
if take_deterministic_summary_pin():
# Surfaces through the fallback summary's reason line and the host's one-shot user warning.
self._last_summary_error = (
"summary model stalled on every route; deterministic fallback summary inserted"
)
telemetry = getattr(self, "_active_compression_telemetry", None)
if isinstance(telemetry, dict):
telemetry["failure_class"] = "stall_deterministic_fallback"
return None
# Focus-topic derivation scans user turns; only pay when a summary is generated.
summary_kwargs: Dict[str, Any] = {
"focus_topic": focus_topic or self._derive_auto_focus_topic(messages),
"memory_context": memory_context,
}
if _accepts_keyword_argument(self._generate_summary, "bypass_cooldown"):
summary_kwargs["bypass_cooldown"] = bypass_cooldown
try:
return self._generate_summary(turns_to_summarize, **summary_kwargs)
except AuxiliaryExplicitCancellation:
# Cancellation is a true no-op: restore the scan's mutation before the exception escapes.
self._previous_summary = scan.previous_summary_before
self._summary_has_user_turn = scan.has_user_turn_before
raise