When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.
Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.
Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
67 lines
3.2 KiB
Python
67 lines
3.2 KiB
Python
"""Per-attempt recovery bookkeeping (``TurnRetryState``) for the conversation turn loop.
|
|
Dependency-free so it imports without a cycle."""
|
|
|
|
from __future__ import annotations
|
|
|
|
from dataclasses import dataclass
|
|
|
|
|
|
@dataclass
|
|
class TurnRetryState:
|
|
"""One-shot recovery guards + restart signals for a single API-call attempt.
|
|
|
|
A fresh instance is created per ``api_call_count`` iteration; each guard fires at
|
|
most once, and ``restart_with_*`` signals tell the loop to rebuild and retry.
|
|
Loop-control (``retry_count``, ``max_retries``) stays as plain loop locals."""
|
|
|
|
# Per-provider OAuth / credential refresh guards
|
|
codex_auth_retry_attempted: bool = False
|
|
anthropic_auth_retry_attempted: bool = False
|
|
nous_auth_retry_attempted: bool = False
|
|
nous_paid_entitlement_refresh_attempted: bool = False
|
|
# Nous free tier: one model move onto the tier's own model after a ``model_not_free``
|
|
# refusal, and one route re-read after a wrong-host refusal (``anon_on_paid_host``).
|
|
welcome_model_switch_attempted: bool = False
|
|
welcome_route_heal_attempted: bool = False
|
|
copilot_auth_retry_attempted: bool = False
|
|
# Copilot surfaces a stale credential as a 400 ``model_not_available_for_integrator``
|
|
# / ``model_not_supported``, not a 401 — separate guard from the 401 one.
|
|
copilot_stale_cred_retry_attempted: bool = False
|
|
vertex_auth_retry_attempted: bool = False
|
|
|
|
# Format / payload recovery guards
|
|
thinking_sig_retry_attempted: bool = False
|
|
invalid_encrypted_content_retry_attempted: bool = False
|
|
native_compaction_reject_retry_attempted: bool = False
|
|
image_shrink_retry_attempted: bool = False
|
|
multimodal_tool_content_retry_attempted: bool = False
|
|
reasoning_mandatory_retry_attempted: bool = False
|
|
oauth_1m_beta_retry_attempted: bool = False
|
|
llama_cpp_grammar_retry_attempted: bool = False
|
|
|
|
# Transport / rate-limit recovery
|
|
primary_recovery_attempted: bool = False
|
|
has_retried_429: bool = False
|
|
# Persistent 401/403 already escalated to the fallback chain once this attempt.
|
|
auth_failover_attempted: bool = False
|
|
# Post-exhaustion auto-recovery cycles spent on this API call (agent.auto_recovery_cycles caps it).
|
|
auto_recovery_cycles_used: int = 0
|
|
|
|
# Restart signals (read by the outer loop after the attempt)
|
|
restart_with_compressed_messages: bool = False
|
|
restart_with_length_continuation: bool = False
|
|
# A fallback activation (incl. content-filter stream stalls) rolled partial content
|
|
# off ``messages``; re-issue the call against the new provider.
|
|
restart_with_rebuilt_messages: bool = False
|
|
# A user correction cancelled the in-flight request: append a role-safe checkpoint +
|
|
# user message, rebuild the payload, and retry the same logical iteration.
|
|
restart_with_redirected_messages: bool = False
|
|
|
|
|
|
# ---- BEGIN PLUGIN-COMPAT (revert-scheduled; see COMPAT_MANIFEST.md) ----
|
|
# Names external plugins imported from this module before the Sep 2026 decomposition.
|
|
# Internal code MUST NOT use these (scripts/check_compat_pointers.py fails CI if it does).
|
|
# The whole block is removed by reverting the commit that added it.
|
|
from dataclasses import fields # noqa: F401,E402
|
|
# ---- END PLUGIN-COMPAT ----
|