Commit Graph

27 Commits

Author SHA1 Message Date
kshitijk4poor
bfeafa2d68 fix(agent): collapse orphan continuation trail on retry exhaustion (#119001)
Ports #119027::finalize_continuation_partial: when a stream drop entered the
continuation path and the next request exhausted retries before a token, the
_length_continuation_fragment/_nudge rows were persisted as-is, so resume
replayed a dangling synthetic user nudge. Collapse them into one assistant row
before persistence and feed the collapsed text to the #119081 partial-retention
path (the fragment rows are gone by then).

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-24 22:33:58 +05:30
kshitijk4poor
144e9711b8 fix(agent): treat jiter SSE parse ValueErrors as transient provider errors (#65147)
Extends PROVIDER_STREAM_PARSE_MARKERS (single home) with jiter's serde
vocabulary and excludes those ValueErrors from _is_local_validation_error so
the turn-level retry/fallback path runs instead of aborting as a local bug.
Grafted from #65154 (its conversation_loop regex is superseded by the markers).

Co-authored-by: Simplicio, Wesley (ext) <wesley.simplicio.ext@siemens-energy.com>
2026-09-24 22:26:23 +05:30
fangliquanflq
125bdff601 fix(agent): honor provider reset for fallback cooldown 2026-09-20 17:00:43 -07:00
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
teknium1
1cd2c45fcb fix: credit-wall 404 is one billing table row; billing fallback warns with profile + remedy
Trim the salvage of #115706 to the existing seams:

- The structured code joins _BILLING_ERROR_CODES and _status_404 consults that table
  first, exactly like _status_429 (the status handler always returns, so _by_error_code
  never saw the code). Drops the duplicate _CREDIT_EXHAUSTION_404_CODES set, the
  message-substring scan, and the credit_exhaustion_code verdict marker.
- The log moves from two call sites in the turn loop into try_activate_fallback, the
  one chokepoint every fallback switch passes through. A billing switch is a WARNING
  naming the profile (resolved from the scoped home, so under multiplex it is the
  failing profile, not the launch profile), both models, and the `hermes [-p X] model`
  remedy; every other reason keeps the INFO line.
- Tests trimmed to two invariants (classifier row + control body; warning is
  profile-scoped A->B and non-billing stays INFO), both red on origin/main.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-19 22:40:16 -07:00
Yagna Vudathu
d861bca178 fix: route 404 insufficient_credits_for_paid_model to fallback_model chain
A 404 carrying insufficient_credits_for_paid_model (paid model ungated
by credits) classified as unknown: retryable with no fallback, burning
retries and never switching models. Treat it like 429-exhaustion --
billing with rotate+fallback -- and log an actionable ERROR naming the
credits and the fallback target on activation.

Fixes #115702

(cherry picked from commit a1f7f9996bb82230c945340dcb279ffa923e90a1)
2026-09-19 22:40:16 -07:00
teknium1
0752127c5e feat(agent): bounded auto-recovery ladder after retries and fallback are spent (#85426, #107307)
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.

Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.

Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
2026-09-19 12:38:55 -07:00
teknium1
b23f31c2bd fix: cover the attempt-line verdict through handle_api_error
The retryable kwarg only reached log_api_error_attempt from
turn_api_error.handle_api_error, yet both tests called the helper
directly, so dropping the forwarding line kept the suite green.
Replace the retryable=True direct-call test with one that drives the
production entry (classifier patched to a non-retryable 401) and
asserts 'attempt 1/3, not retryable' on the log and status buffer:
red with the kwarg removed, green with it restored.

Also note why _touch_activity keeps the plain counter: it is a
watchdog liveness label, not chat output.
2026-09-19 09:44:25 -07:00
teknium1
9c75ae13ee fix(agent): name a non-retryable API failure on the attempt line
A 401 on a static-key route has no credential to refresh and no pool
entry to rotate to, so it goes straight to the fallback chain — yet the
log said "API call failed (attempt 1/3)" and "attempt 2/3" never came.
Readers took the stuck counter for a retry bug (#73237). The
classifier's verdict now rides the same line on both surfaces (logger
warning and the buffered status trace): "attempt 1/3, not retryable".
Retryable failures keep the plain counter.

Part of #73237 — the policy question (retry an unchanged static
credential once before fallback) is left to the maintainer.
2026-09-19 09:44:25 -07:00
teknium1
a830061c9d fix(agent): recognise reversed "reasoning_effort 'none' unsupported" on every surface (#114460)
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:

- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
  between auxiliary_client._is_unsupported_parameter_error and the new
  classifier helper); is_reasoning_field_rejection() replaces
  is_reasoning_disable_rejected() with the same token gate and a
  symmetric "unsupported" window so both word orders match
  ("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
  model "requires reasoning" (a chat-only relay does not); a second
  reasoning-field rejection in the same turn is treated as spent and
  takes the fallback chain instead of replaying the identical request
  max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
  docs mention the reversed wording and the main-loop recovery.

Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
2026-09-18 10:21:05 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
Robin Fernandes
165271fceb fix(nous): keep anonymous errors exclusive to anonymous users 2026-09-17 12:28:25 +05:30
dacheah
270fe15c8b fix(agent): route oversize rejections that arrive as HTTP 400 to the image-shrink recovery
NVIDIA NIM ("Please make sure your payload is below 26214400 bytes"), Alibaba
DashScope ("String value length (N) exceeds the maximum allowed (M, from
`StreamReadConstraints.getMaxStringLength()`)") and Nebius Token Factory (a
pydantic `string_type` detail on `body.messages.N.content` whose rejected input
carries the inline image) enforce their byte caps with a 400, so the shrink
recovery keyed on 413 / image vocabulary never ran: the turn died as a
non-retryable format_error, or (Nebius) the generic "Input should be a valid
string" was claimed by the multimodal tool-content rule and its one retry was
spent stripping tool images that were never there.

- add the two wordings to _IMAGE_TOO_LARGE_PATTERNS
- structural rule for the pydantic detail (message content loc, not tool-scoped,
  rejected input holds a data:image part over the shrink target) checked ahead
  of the keyword multimodal rule in _classify_400
- settle_unrecovered_error: once the single shrink attempt was spent without
  recovering, treat image_too_large as a client error and fall back instead of
  re-sending the byte-identical oversized body max_retries times (a
  payload-scoped cap can be tripped by text alone)

Trimmed salvage of #111975 by @dacheah: exclusion-guard tables, the shrink-target
mirror constant and the dedicated fallback block are dropped in favour of the
existing client-error branch and an import of the real shrink target.

Fixes #112473
2026-09-16 17:09:30 -07:00
Robin Fernandes
51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1
79007efc48 fix(agent): a no-fallback verdict wins over the local-ValueError fallback allowance
`MoAPresetNotFoundError` subclasses ValueError, so `is_local_validation_error`
re-opened the fallback the classifier had just refused (#55933). Only an
UNCLASSIFIED (reason=unknown) local error keeps the historical fallback.
Test exercises the real exception through settlement, not the classifier bool.
2026-09-12 08:27:11 -07:00
teknium1
293cb54f28 fix(agent): keep MoA verdicts fallback-free; nest the gated fallback under one if
Two CI regressions from the previous commit, both mine:

- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
  `should_fallback` to True for the MoA adapter-shape and missing-preset
  verdicts. #55933 made those deliberately NOT fall back (a fallback would
  silently replace the MoA route with a single model); restore
  `retryable=False` only. The gate in `settle_unrecovered_error` now honours
  that for real: on main these verdicts never reached the fallback branch
  because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
  `_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
  site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
  broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
  block instead.
2026-09-12 08:27:11 -07:00
teknium1
dfd4aa4a94 fix(agent): malformed tool-call-argument 400s no longer walk the fallback chain
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).

- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
  request-validation and overflow heuristics, returning format_error with
  `retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
  verdicts that legitimately reach that branch (policy block, TLS chain, MoA
  shape/preset errors) now state `should_fallback=True` explicitly, so the gate
  changes behaviour only for the new verdict. Local validation errors keep
  their historical fallback.

Fixes #12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.

Co-authored-by: cuyua9 <2114364329@qq.com>
2026-09-12 08:27:11 -07:00
teknium1
6c89cb032e refactor(agent): move the entitlement rejection marker into agent/fallback_cooldown
The two facades post-#102117 (chat_completion_helpers, agent_runtime_helpers) must not grow
behaviour; the marker/predicate belong with the sibling that already owns primary-cooldown
state shared by the fallback walk and restore_primary_runtime. Also drops the two
try/except-return-False wrappers around pool.entries() and normalize_model_for_provider
(neither raises for a live pool / the never-raising normalizer).

Tests trimmed to the two invariants (#106475): the marker fires only on the single-credential
Codex entitlement 400 and the walk skips the rejected slug in either form; restore is gated on
a rejected primary and still restores an unrelated one.
2026-09-09 11:29:56 -07:00
liuhao1024
808c5b30aa fix(agent): fail closed when Codex account-model entitlement 400s exhaust the chain
A Codex ChatGPT-account 400 ('The X model is not supported when using Codex
with a ChatGPT account.') names the model, so with a single credential the
slug is dead for that account. The fallback walk still re-selected it and
restore_primary_runtime switched back to the primary at the start of every
turn, announcing an unverified 'Primary model restored' — the two warnings
alternated forever with zero delivered answers (#106475).

Record the rejected (provider, model) pair on the non-retryable client-error
path (only when no multi-credential pool exists — rotation covers that case,
#71970), skip rejected entries during the fallback walk, and gate
restore_primary_runtime on the primary's slug so the session fails closed
with the terminal entitlement error instead of oscillating. Fixes #106475.

(cherry picked from commit 471435b10288f15387b2549a8b02551d82c0f670)
2026-09-09 11:29:56 -07:00
Teknium
2f4534eaaa refactor(agent/turn): _NONRETRYABLE_LABELS table shared by terminal status + fallback announce; docstring trims 2026-09-02 19:39:27 -07:00
Teknium
82fc85a8d6 refactor(agent/turn_api_error): keep direct _verdict("break") at fallback sites (#84733 guard) 2026-09-02 19:23:48 -07:00
Teknium
b925acb5a9 refactor(agent/turn_recovery): abort_turn_on_interrupt shared by backoff sleep and handle_api_error 2026-09-02 18:55:11 -07:00
Teknium
e50afc20bb refactor(agent/turn_api_error): _is_local_validation_error predicate, _RETRYABLE_CLIENT_REASONS table, _fallback_break helper, verdict passthrough; trim turn_loop_errors/turn_retry_state 2026-09-02 18:42:15 -07:00
Teknium
c94ced6225 refactor(agent/turn): AST-neutral bracket/signature packing across r3-08 slice 2026-09-02 18:01:27 -07:00
Teknium
7dc43c94d1 refactor(turn): sub-split handle_api_error / check_api_response / run_tool_round so no new turn_* function exceeds 300 LOC 2026-09-02 17:28:22 -07:00
Teknium
2769937936 refactor(turn): lift iteration entry/announce, Nous rate guard, API interrupt, retry-restart consumer and preflight-timeout result out of run_conversation 2026-09-02 16:37:14 -07:00
Teknium
9fda4e5bac refactor(turn): extract retry-loop API error handler, request build, provider call and response check into agent/turn_api_*.py + agent/turn_response_check.py 2026-09-02 16:07:51 -07:00