Ports #119027::finalize_continuation_partial: when a stream drop entered the
continuation path and the next request exhausted retries before a token, the
_length_continuation_fragment/_nudge rows were persisted as-is, so resume
replayed a dangling synthetic user nudge. Collapse them into one assistant row
before persistence and feed the collapsed text to the #119081 partial-retention
path (the fragment rows are gone by then).
Co-authored-by: fangliquan <fangliquan@qq.com>
Extends PROVIDER_STREAM_PARSE_MARKERS (single home) with jiter's serde
vocabulary and excludes those ValueErrors from _is_local_validation_error so
the turn-level retry/fallback path runs instead of aborting as a local bug.
Grafted from #65154 (its conversation_loop regex is superseded by the markers).
Co-authored-by: Simplicio, Wesley (ext) <wesley.simplicio.ext@siemens-energy.com>
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
Trim the salvage of #115706 to the existing seams:
- The structured code joins _BILLING_ERROR_CODES and _status_404 consults that table
first, exactly like _status_429 (the status handler always returns, so _by_error_code
never saw the code). Drops the duplicate _CREDIT_EXHAUSTION_404_CODES set, the
message-substring scan, and the credit_exhaustion_code verdict marker.
- The log moves from two call sites in the turn loop into try_activate_fallback, the
one chokepoint every fallback switch passes through. A billing switch is a WARNING
naming the profile (resolved from the scoped home, so under multiplex it is the
failing profile, not the launch profile), both models, and the `hermes [-p X] model`
remedy; every other reason keeps the INFO line.
- Tests trimmed to two invariants (classifier row + control body; warning is
profile-scoped A->B and non-billing stays INFO), both red on origin/main.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
A 404 carrying insufficient_credits_for_paid_model (paid model ungated
by credits) classified as unknown: retryable with no fallback, burning
retries and never switching models. Treat it like 429-exhaustion --
billing with rotate+fallback -- and log an actionable ERROR naming the
credits and the fallback target on activation.
Fixes#115702
(cherry picked from commit a1f7f9996bb82230c945340dcb279ffa923e90a1)
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.
Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.
Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
The retryable kwarg only reached log_api_error_attempt from
turn_api_error.handle_api_error, yet both tests called the helper
directly, so dropping the forwarding line kept the suite green.
Replace the retryable=True direct-call test with one that drives the
production entry (classifier patched to a non-retryable 401) and
asserts 'attempt 1/3, not retryable' on the log and status buffer:
red with the kwarg removed, green with it restored.
Also note why _touch_activity keeps the plain counter: it is a
watchdog liveness label, not chat output.
A 401 on a static-key route has no credential to refresh and no pool
entry to rotate to, so it goes straight to the fallback chain — yet the
log said "API call failed (attempt 1/3)" and "attempt 2/3" never came.
Readers took the stuck counter for a retry bug (#73237). The
classifier's verdict now rides the same line on both surfaces (logger
warning and the buffered status trace): "attempt 1/3, not retryable".
Retryable failures keep the plain counter.
Part of #73237 — the policy question (retry an unchanged static
credential once before fallback) is left to the maintainer.
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:
- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
between auxiliary_client._is_unsupported_parameter_error and the new
classifier helper); is_reasoning_field_rejection() replaces
is_reasoning_disable_rejected() with the same token gate and a
symmetric "unsupported" window so both word orders match
("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
model "requires reasoning" (a chat-only relay does not); a second
reasoning-field rejection in the same turn is treated as spent and
takes the fallback chain instead of replaying the identical request
max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
docs mention the reversed wording and the main-loop recovery.
Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
NVIDIA NIM ("Please make sure your payload is below 26214400 bytes"), Alibaba
DashScope ("String value length (N) exceeds the maximum allowed (M, from
`StreamReadConstraints.getMaxStringLength()`)") and Nebius Token Factory (a
pydantic `string_type` detail on `body.messages.N.content` whose rejected input
carries the inline image) enforce their byte caps with a 400, so the shrink
recovery keyed on 413 / image vocabulary never ran: the turn died as a
non-retryable format_error, or (Nebius) the generic "Input should be a valid
string" was claimed by the multimodal tool-content rule and its one retry was
spent stripping tool images that were never there.
- add the two wordings to _IMAGE_TOO_LARGE_PATTERNS
- structural rule for the pydantic detail (message content loc, not tool-scoped,
rejected input holds a data:image part over the shrink target) checked ahead
of the keyword multimodal rule in _classify_400
- settle_unrecovered_error: once the single shrink attempt was spent without
recovering, treat image_too_large as a client error and fall back instead of
re-sending the byte-identical oversized body max_retries times (a
payload-scoped cap can be tripped by text alone)
Trimmed salvage of #111975 by @dacheah: exclusion-guard tables, the shrink-target
mirror constant and the dedicated fallback block are dropped in favour of the
existing client-error branch and an import of the real shrink target.
Fixes#112473
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
Two CI regressions from the previous commit, both mine:
- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
`should_fallback` to True for the MoA adapter-shape and missing-preset
verdicts. #55933 made those deliberately NOT fall back (a fallback would
silently replace the MoA route with a single model); restore
`retryable=False` only. The gate in `settle_unrecovered_error` now honours
that for real: on main these verdicts never reached the fallback branch
because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
`_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
block instead.
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).
- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
request-validation and overflow heuristics, returning format_error with
`retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
verdicts that legitimately reach that branch (policy block, TLS chain, MoA
shape/preset errors) now state `should_fallback=True` explicitly, so the gate
changes behaviour only for the new verdict. Local validation errors keep
their historical fallback.
Fixes#12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.
Co-authored-by: cuyua9 <2114364329@qq.com>
The two facades post-#102117 (chat_completion_helpers, agent_runtime_helpers) must not grow
behaviour; the marker/predicate belong with the sibling that already owns primary-cooldown
state shared by the fallback walk and restore_primary_runtime. Also drops the two
try/except-return-False wrappers around pool.entries() and normalize_model_for_provider
(neither raises for a live pool / the never-raising normalizer).
Tests trimmed to the two invariants (#106475): the marker fires only on the single-credential
Codex entitlement 400 and the walk skips the rejected slug in either form; restore is gated on
a rejected primary and still restores an unrelated one.
A Codex ChatGPT-account 400 ('The X model is not supported when using Codex
with a ChatGPT account.') names the model, so with a single credential the
slug is dead for that account. The fallback walk still re-selected it and
restore_primary_runtime switched back to the primary at the start of every
turn, announcing an unverified 'Primary model restored' — the two warnings
alternated forever with zero delivered answers (#106475).
Record the rejected (provider, model) pair on the non-retryable client-error
path (only when no multi-credential pool exists — rotation covers that case,
#71970), skip rejected entries during the fallback walk, and gate
restore_primary_runtime on the primary's slug so the session fails closed
with the terminal entitlement error instead of oscillating. Fixes#106475.
(cherry picked from commit 471435b10288f15387b2549a8b02551d82c0f670)