Commit Graph

144 Commits

Author SHA1 Message Date
kshitijk4poor
763646c30d fix(agent): only 4xx/5xx body codes count as a stream error's HTTP status (#121270)
A relay wrapping an error as {"code": 200, "success": false} must not route to
_by_status as a 200. Also reuse _error_obj for the error.* candidate walk and
point at the string-parsing text-SSE sibling.
2026-09-24 22:36:08 +05:30
kshitijk4poor
c9a1c3d6f5 refactor(agent): route status-less render errors through _MESSAGE_HEAD_RULES (#62662)
Drop the bespoke _streaming_render_error stage and its hand-built verdict: the
status-less message path already has a head table and _V_FORMAT_ERROR is the
canonical format_error/no-retry/fallback verdict. _by_status returns None when
there is no status, so the stage position bought nothing.
2026-09-24 22:36:08 +05:30
AlexFucuson9
d0dfea836b fix(agent): classify status-less streaming render errors as format_error
LM Studio / llama.cpp raise a bare status-less APIError when chat-template
Jinja rendering fails mid-stream; classify as format_error with fallback
instead of burning same-provider retries. Stage fires only when status_code
is None. Hand-applied from #64834 (supersedes #62673) onto the staged
classifier pipeline.

Fixes #62662
2026-09-24 22:36:08 +05:30
liuhao1024
0a661b7c94 fix(agent): an in-stream SSE error object's numeric code is its HTTP status
An aggregator can deliver an upstream failure as {"error": {"code": 403}}
inside an HTTP-200 SSE stream; the SDK then raises a status-less APIError
whose body carries the status. _extract_status_code only read
exc.status_code/.status, so the ban climbed the transient retry ladder
(api_max_retries resends, then 'temporarily unavailable, wait a minute')
which can never succeed for e.g. a banned account.

Fall back to a body-carried numeric code (100-599, int only — symbolic
string codes stay with _code_from_payload), mirroring the text-SSE path's
_status_code_from_payload. A 403 now classifies as auth: exactly one
request, fallback-eligible, no wait-and-retry advice.

Fixes #121270

(cherry picked from commit 8a09fc53614ea3b954faebb4bfcccc176116338c)
2026-09-24 22:36:08 +05:30
teknium1
2c948e6aa2 fix(agent): an upstream account ban relayed in a 200 stream fails once instead of retrying as an outage
OpenRouter relays OpenAI's "this user has been blocked for a previous
policy violation" as HTTP 200 + an SSE error event. The SDK raises a
status-less APIError that no classifier rule matched, so it landed in
the retryable `unknown` bucket: every retry was re-sent (3 streamed
requests by default), a configured fallback only engaged after the
ladder, and the user was told the provider "looks temporarily
unavailable. Wait a minute and send /retry".

- error_classifier: match the ban phrase status-agnostically in
  _provider_special_cases -> provider_policy_blocked (non-retryable,
  fallback, no credential rotation: every key on a banned account is
  banned; a 403 variant is not a bad key).
- turn_failure_copy: provider_policy_blocked chat copy now covers an
  account block as well as data/privacy settings; add its cause gloss
  so cron and subagent notices explain it instead of printing raw text.
- cron: provider_policy_blocked action (retrying won't help; pin another
  model).
- docs: FAQ entry for the reply; auto-recovery ladder exclusion list.
2026-09-23 10:40:08 -07:00
teknium1
30d14a3f73 fix: name the CommandCode upstream-outage pattern
The cherry-pick conflict dropped the contributor comment hunk; keep the WHY next to the entry.
2026-09-20 15:28:21 -07:00
fangliquan
1f3c882d81 fix(agent): narrow upstream outage matching 2026-09-20 15:28:21 -07:00
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
liuhao1024
d4f88f39ee fix(agent): classify CommandCode 'Content Exists Risk' 400 as a content-policy block
The CommandCode gateway (OpenAI-compatible aggregator fronting DeepSeek
models) rejects filtered prompts with a hard HTTP 400 whose message is the
verbatim string "Content Exists Risk" and whose nested param envelope
marks the rejection isRetryable=false. The pattern table in
_CONTENT_POLICY_BLOCKED_PATTERNS did not know this wording, so the 400
fell through to the generic format_error verdict: no fallback model was
tried and the surfaced copy blamed a malformed request even though the
request shape was fine (the same profile re-fires per prompt).

Add the verbatim phrase to the per-prompt safety-filter pattern list so
the classifier returns content_policy_blocked with should_fallback=True —
retryable unchanged-prompt attempts stop burning and the configured
fallback chain gets its turn. See #115218.
2026-09-19 23:39:24 -07:00
teknium1
1cd2c45fcb fix: credit-wall 404 is one billing table row; billing fallback warns with profile + remedy
Trim the salvage of #115706 to the existing seams:

- The structured code joins _BILLING_ERROR_CODES and _status_404 consults that table
  first, exactly like _status_429 (the status handler always returns, so _by_error_code
  never saw the code). Drops the duplicate _CREDIT_EXHAUSTION_404_CODES set, the
  message-substring scan, and the credit_exhaustion_code verdict marker.
- The log moves from two call sites in the turn loop into try_activate_fallback, the
  one chokepoint every fallback switch passes through. A billing switch is a WARNING
  naming the profile (resolved from the scoped home, so under multiplex it is the
  failing profile, not the launch profile), both models, and the `hermes [-p X] model`
  remedy; every other reason keeps the INFO line.
- Tests trimmed to two invariants (classifier row + control body; warning is
  profile-scoped A->B and non-billing stays INFO), both red on origin/main.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-19 22:40:16 -07:00
Yagna Vudathu
d861bca178 fix: route 404 insufficient_credits_for_paid_model to fallback_model chain
A 404 carrying insufficient_credits_for_paid_model (paid model ungated
by credits) classified as unknown: retryable with no fallback, burning
retries and never switching models. Treat it like 429-exhaustion --
billing with rotate+fallback -- and log an actionable ERROR naming the
credits and the fallback target on activation.

Fixes #115702

(cherry picked from commit a1f7f9996bb82230c945340dcb279ffa923e90a1)
2026-09-19 22:40:16 -07:00
teknium1
bce20d0b1f feat: model-provider plugins classify their own API errors via ProviderProfile.classify_api_error
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).

Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.

Part of #116408

(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
2026-09-19 21:24:05 -07:00
Teknium
656eb28309 Merge pull request #115822 from NousResearch/fix/boa-streaming-stalls-timeouts-reasoning-streak
fix(agent): Codex reasoning-only stalls switch to the fallback provider instead of ending incomplete (#67321, salvage #67336)
2026-09-19 11:09:15 -07:00
teknium1
e85cb94da1 chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:40 -07:00
teknium1
8d1b2c5b6e chore: merge origin/main (resolve agent/chat_completion_helpers.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:36 -07:00
teknium1
05c95d216e fix: read a top-level detail error body so a descriptive 400 is not a bare 400
`agent/error_classifier.py::_body_message_candidates` never yielded the
FastAPI-style top-level `detail` key (string, or nested `{"message": ...}`),
so the Codex gateway's `{"detail": "The '<model>' model is not supported when
using Codex with a ChatGPT account."}` 400 read as a *bare* 400. On a large
session `_classify_400`'s generic-400 heuristic then classified it as
context_overflow: the loop burned compression attempts ("Context length
exceeded (109,962 tokens). Cannot compress further.") and, because
`is_client_error` excludes overflow, the entitlement marker from #106549 never
ran. Reading `detail` makes it a descriptive rejection (format_error: abort +
fall back, no compression) and surfaces the provider's text as the message
instead of `Error code: 400 - {...}`. The pydantic list shape of `detail` is
still handled by `_oversized_message_content_rejection` and is not yielded.

Slim re-port of #100783's detail-body half onto the rule-table classifier;
the session/weekly usage-limit half is a separate class and was dropped.

Refs #81558
Refs #106475
Co-authored-by: Oleg Nagornyy <nagornyy.o@gmail.com>
2026-09-19 10:23:44 -07:00
liuhao1024
3306e105b3 fix(agent): recover from strict OpenAI-compatible "regex lookaround is not supported" 400s
Strict OpenAI-compatible schema validators reject tool JSON Schemas whose
``pattern`` contains a lookahead/lookbehind with
"Invalid JSON schema: regex lookaround is not supported" (HTTP 400,
code invalid_json_schema). The classifier only knew the llama.cpp grammar
sentences, so the turn failed as a non-retryable client error even though
the existing strip-pattern/format retry path already fixes the request.

Match the lookaround sentence in the same grammar guard so it routes through
FailoverReason.llama_cpp_grammar_pattern and retries once with the sanitized
schemas. Ported from #42635 onto the current grammar_hit shape.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-19 10:16:27 -07:00
teknium1
6f6ed01355 fix: WAF 403s stop reading as key rejections; anthropic_messages routes send custom_providers extra_headers
Two gaps for custom providers behind a WAF/CDN:

- `build_anthropic_client` never consulted `custom_providers[].extra_headers`,
  so a relay in `anthropic_messages` mode that rejects the SDK User-Agent kept
  403ing even with `extra_headers: {User-Agent: ...}` configured, while the
  OpenAI-wire clients already applied it. The lookup now lives in
  `_new_sdk_client`, the one constructor every builder path goes through
  (init, /model switch, rebuild, auxiliary), keyed by the caller's raw route
  because entries are keyed by the `/v1` form the normalizer strips.
  Salvaged direction of #46002 (@wait4xx). Fixes #24293, #9721.

- `_status_403` classified every non-billing 403 as `auth`, so a WAF's plain
  "Your request was blocked." or a Cloudflare browser challenge printed "Your
  API key was rejected" and could rotate a healthy credential. A 403 carrying
  established block/challenge markers is now `upstream_blocked`: no rotation,
  no retry, fallback allowed, WAF/User-Agent guidance on every surface (CLI
  loop, chat copy, cli chat error copy, TUI gateway + Ink TUI copy). Generic
  403 and all 401 keep the auth verdict. Salvaged direction of #70567
  (@ooiuuii) and #53114 (@AgenticSpark). Fixes #53099, #70566.
2026-09-19 09:57:21 -07:00
teknium1
5d0ddb7dd8 fix(agent): exhausted plan-quota 429 names its reset window instead of 'wait a minute'
_summarize_api_error reduces the usage_limit_reached body to 'HTTP 429: The
usage limit has been reached', so no text-based parser could ever reach the
8.6h reset from the in-loop failure path. _status_429 now stamps
error_context['reset_at'] from the body's reset fields, Retry-After or the
message grammar (same table the credential pool uses), and
max_retries_exhausted_result hands the remaining seconds to exhausted_copy,
which says 'its usage limit resets in ~9h. Send /retry after that' once the
window is >= 2 minutes. A throttle with no window keeps the short-wait copy.
Proven with a real openai.RateLimitError carrying resets_in_seconds=30995
driven through the production result builder. (#89401 atom 1)
2026-09-19 09:49:39 -07:00
teknium1
1213109474 fix: retry a 403 stamped code=upstream_unavailable instead of failing as auth
Some OpenAI-compatible gateways answer an upstream outage with HTTP 403
and a structured error code ``upstream_unavailable`` ("Upstream service
temporarily unavailable. Please retry later."). The 403 handler only
distinguished billing from auth, so this transient outage was classified
``auth``/non-retryable: agent.api_max_retries was ignored, the turn
died on attempt 1 and the credential was treated as refused. Give the
structured transient code precedence over the generic 403 default and
classify it as overloaded (retry with backoff, no credential rotation),
matching how the same condition is handled when it arrives as 503.

Part of #75388 (case 1; the content-policy retry-budget case is a
deliberate design from 0554ef1aa3 and needs a maintainer decision).
2026-09-19 09:48:02 -07:00
teknium1
b580e2152e fix(error_classifier): read Gemini's error.status when error.code is numeric; pin anthropic RATE_LIMIT_ERROR rotation
Gemini's real wire body is {"error": {"code": 503, "status": "UNAVAILABLE",
"message": ...}}: the symbolic code lives in error.status while error.code
carries the HTTP status. _code_from_payload only read error.code/error.type,
so the numeric code was discarded and the gemini rows in
_PROVIDER_CODE_VERDICTS (keyed on unavailable/deadline_exceeded/internal) were
unreachable; a status-less 503/UNAVAILABLE classified as unknown. A non-string
code now falls back to error.status.

Also pins the anthropic rate_limit_error row: reason rate_limit with
should_rotate_credential and should_fallback both True, which the existing
parametrized case (rotate False) could not host.
2026-09-19 09:47:31 -07:00
luyifan
f2651cb0f6 fix: classify provider code-only errors into structured failover reasons
A provider error body that carries only a structured code and no HTTP
status (Gemini UNAVAILABLE / DEADLINE_EXCEEDED / INTERNAL, Anthropic
API_ERROR, OpenAI SERVER_ERROR) fell through every classifier stage to
FailoverReason.unknown, so the retry loop treated a provider overload
or timeout as a generic retryable failure and never reached the
overload/timeout recovery hints. Add a per-provider code table consulted
after the shared _ERROR_CODE_VERDICTS map; codes are scoped to their
provider family so a same-named code from another backend stays unknown.

Fixes #70414
Salvages #70425 (@ooiuuii) onto the rule-table classifier layout.
2026-09-19 09:47:31 -07:00
teknium1
7958cbf2de fix: rotate Codex pool credentials on a ChatGPT-account model entitlement 400
The exact `The '<model>' model is not supported when using Codex with a
ChatGPT account.` 400 classified as format_error, so a two-entry openai-codex
pool never tried its second account: the turn failed as a malformed request
even though the other account was entitled (#71970). #106475 covered the
single-credential case and explicitly left the multi-entry case to rotation,
but no rotation branch existed for it.

- classify the exact normalized text as FailoverReason.model_entitlement
  (rotate + fallback, never retry); arbitrary 400s stay format_error
- recover_with_credential_pool rotates once on that reason; the pool records
  it through the existing model_cooldowns path (Anthropic per-model 429), so
  only (credential, model) is benched: other models keep using the account,
  and `hermes auth` reset clears the marker with everything else
- _mark_entitlement_rejected_model gates on pool.has_available(model=...)
  instead of entry count, so once every account rejects the model it falls
  back to the #106475 session marker (fallback walk skips it, no oscillation)

Salvages the design of #71973 by @kilhyeonjun on the current classifier
tables and the model-cooldown substrate that landed since.

Co-authored-by: kilhyeonjun <kboxstar@gmail.com>
2026-09-19 09:37:23 -07:00
teknium1
c7cf3f8af1 fix(classifier): route two more rejected-reasoning-replay 400s to the replay strip
Two Responses-API 400s meant "the replayed encrypted reasoning was rejected"
but never reached the one-shot recovery in turn_recovery (disable replay,
strip codex_reasoning_items, retry):

- OpenAI's ``thinking_signature_invalid`` code contains "thinking" and
  "signature", so the Anthropic thinking-block heuristic claimed it first;
  that recovery strips Anthropic fields and resends the same stale encrypted
  item on every retry (#70595).
- The ChatGPT Codex backend returns a bare ``{"detail": "Unsupported content
  type"}`` for the same rejection, which matched nothing and aborted the turn
  as a non-retryable format_error (#51512). It joins the #92353 exact-envelope,
  provider-gated codex mapping, so a no-replay 400 keeps aborting as before
  (the recovery still requires cached codex_reasoning_items).

Co-authored-by: ooiuuii <al3060388206@gmail.com>
2026-09-19 09:23:54 -07:00
luyifan
0acd96a439 fix(agent): classify Codex account token failures 2026-09-19 09:22:50 -07:00
teknium1
2768f31628 fix: step reasoning up to the floor when a route refuses to disable it
Endpoints that understand the reasoning field but refuse the OFF ("Reasoning is
mandatory for this endpoint and cannot be disabled" -- the Nous Portal on
gpt-6-astra) got two different wrong answers:

* auxiliary lanes (title generation sends reasoning_config={"enabled": False})
  400'd outright -- the #112781 strip rung only matched "unsupported/unknown
  field" wordings, so every session on such a route stayed untitled;
* the main loop dropped the disable and let the route pick its default effort,
  the opposite of what a thinking-off user asked for.

Both now step the effort up to the lowest level every reasoning wire accepts
(low) instead. The aux ladder gets a rung ordered before the strip
(agent/auxiliary_reasoning_floor.py: lifts reasoning_effort / extra_body.reasoning
/ _reasoning_config, memoises the (route, model) so the next thinking-off aux call
starts at the floor without the guaranteed 400); the main loop's reasoning_mandatory
recovery records whether the route said mandatory (floor) or unknown-field (drop,
unchanged). error_classifier.is_reasoning_required_rejection separates the two.

Live on the portal (proxy wire capture): title none->400->low->200 titled;
main loop none->400->low->200; second aux call on the route sends low up front.
2026-09-19 02:37:36 -07:00
PRATHAMESH75
f7567a62af fix: hand Codex reasoning-only stalls to the fallback provider instead of the incomplete sentinel
Three consecutive Codex Responses answers that carry only (encrypted) reasoning —
no visible text, no tool call — used to exhaust the 3-continuation budget and end
the turn on "Codex response remained incomplete after 3 continuation attempts",
never touching configured fallback_providers (#67321). Encrypted reasoning items
replay byte-for-byte, so a bare retry deterministically repeats the stall.

- Track a per-turn `_codex_reasoning_only_streak` apart from the aggregate
  `_codex_incomplete_retries`: a visible partial resets the streak, so the mixed
  partial-then-stall variant still reaches its own recovery threshold while the
  turn-wide iteration budget stays the hard bound.
- At streak 3, `continue_codex_incomplete` activates the next fallback with the
  semantic `FailoverReason.incomplete_response`, grants exactly one grace call when
  the trigger consumed the last iteration, and returns `CODEX_FALLBACK_ACTIVATED`;
  the intake re-syncs the Model:/Provider: identity on the system prompt.
- Off the Codex wire the synthetic continuation nudge is stripped alongside the
  opaque replay state (`drop_nudge_marker`) so the Chat Completions payload keeps
  valid role ordering and no Codex-only control text.
- No fallback configured: unchanged terminal sentinel, still bounded at 3 calls.

Ported from PR #67336 by @PRATHAMESH75 onto the decomposed agent/turn_*.py siblings.
2026-09-19 00:04:42 -07:00
teknium1
e343fbd28b fix: recognize structured reasoning-field 400s (param / invalid_reasoning_effort) as reasoning rejections
A custom OpenAI-compatible Responses relay rejects an unsupported
`reasoning.effort` with a message-less structured 400 —
`{"error": {"param": "reasoning.effort", "error_code":
"invalid_reasoning_effort", "retryable": false}}` (#100536). No wording
rule in `is_reasoning_field_rejection` could match an empty message, so
the 400 fell through `_classify_400` to the generic large-session
overflow heuristic and the loop compressed a conversation that had
nothing to do with context size ("Context length exceeded (77 tokens).
Cannot compress further.").

Read the structured signal from the stringified body instead: a `param`
naming a reasoning wire field (`reasoning_effort`, `reasoning.effort`,
`thinking_config`, ...) or an `invalid_reasoning_effort` code is a
reasoning-field rejection whatever the message says. This is the same
predicate the auxiliary strip-and-retry rung uses, so title generation
recovers too. The main loop then takes the existing one-shot
reasoning-disable retry and, when spent, the fallback chain that
surfaces the provider's real error — never compression.

Control: a genuine context-window 400 still classifies context_overflow
with should_compress on; a `param: temperature` enum rejection does not
match.
2026-09-18 23:34:36 -07:00
liuhao1024
da8388c759 fix(agent): match enum-style reasoning_effort rejections with no unsupported marker
commandcode.ai rejects the top-level `reasoning_effort: "none"` (the
disabled-reasoning encoding CustomProfile projects) with an enum 400 —
"Invalid option: expected one of \"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\""
— whose message carries no "unsupported" marker at all: the field name
appears only in the structured 'param' tail, outside the ±32 near-window
of the token match, and no UNSUPPORTED_PARAM_MARKERS entry matches the
enum wording. is_reasoning_field_rejection never fires, so the
strip-and-retry rung (#112781) is skipped and title generation fails
outright (#115277).

Add the enum wording to UNSUPPORTED_PARAM_MARKERS so both surfaces (the
main-loop reasoning_mandatory rung and the auxiliary ladder) recognise
it; a non-reasoning 'param' (temperature) stays unmatched because no
reasoning field token appears anywhere.
2026-09-18 23:32:08 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
7530f40383 fix(errors): auth refusals from a non-stock route name the contacted host
The #113719 report asked that, when the request does fail, the error
names the endpoint that was actually contacted. classify_api_error now
appends "(endpoint: <host>)" to auth-class messages when base_url is set
and is not the provider's own route (provider_owns_route is not True),
so a credential posted to a stale model.base_url reads as a wrong
endpoint rather than a bad key. Stock routes and empty base_url keep
the plain message.
2026-09-18 12:49:40 -07:00
teknium1
95dee173e7 fix(agent): drop non-invariant disable-rung test; document thinking-state 400 trade-off (#114460)
test_disable_drop_rung_sets_session_flag_once stays green with agent/turn_recovery.py from
main (that hunk is comment/message text only), so it is not an invariant of this fix; the
spent-path test already carries the flag-false control. Docstring on
is_reasoning_field_rejection records the accepted trade-off for thinking-state 400s
("Function calling is not supported when thinking is enabled"): the marker sits 6 chars from
the token, so no proximity gate separates it from the forward wordings; cost is one
dropped-disable retry before the spent path falls back.
2026-09-18 10:21:05 -07:00
teknium1
a830061c9d fix(agent): recognise reversed "reasoning_effort 'none' unsupported" on every surface (#114460)
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:

- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
  between auxiliary_client._is_unsupported_parameter_error and the new
  classifier helper); is_reasoning_field_rejection() replaces
  is_reasoning_disable_rejected() with the same token gate and a
  symmetric "unsupported" window so both word orders match
  ("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
  model "requires reasoning" (a chat-only relay does not); a second
  reasoning-field rejection in the same turn is treated as spent and
  takes the fallback chain instead of replaying the identical request
  max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
  docs mention the reversed wording and the main-loop recovery.

Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
2026-09-18 10:21:05 -07:00
Yagna Vudathu
74398c540c fix(agent): map reversed reasoning-disable rejection to reasoning_mandatory (#114460) 2026-09-18 10:21:05 -07:00
teknium1
6368ec6883 fix(aux): aux billing keywords share the main classifier table; quarantine logs name the real reason; vision fallback skips text-only models
Three auxiliary error-classification defects with one shared cause: the aux
ladder kept its own copies of the classifier's phrase tables and its own
fixed log wording, and they drifted.

- `_PAYMENT_KEYWORDS` is now `error_classifier._BILLING_PATTERNS` plus the
  aux-only phrasings, and `_BILLING_PATTERNS` gains OpenRouter's org cap
  "budget limit exceeded". A 403 "Budget limit exceeded (monthly limit)" (or
  "Key limit exceeded") is billing for the main loop AND a payment error for
  the aux ladder, so compression/title/vision fall through to the configured
  fallback instead of raising while the main model is served elsewhere.
- `_mark_provider_unhealthy` takes the caller's reason and level; the skip
  line echoes it. Absent credentials (no OPENROUTER_API_KEY, no Nous auth)
  are DEBUG with the real reason; the confirmed-402 rungs keep the WARNING
  payment wording. Local-only users no longer read "payment / credit error"
  for providers they never configured.
- `_try_main_agent_model_fallback` applies the auto-route's vision capability
  gate, so a transient 429 on the pinned vision lane no longer hands an image
  to a text-only main model (guaranteed 400 "content.type is invalid").
  `_recover_provider_pool` no longer benches a pool entry on an
  upstream-capacity 429 ("at capacity", same `_OVERLOADED_PATTERNS` table).

Fixes #107166, #64144, #108349.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: awizemann <319078+awizemann@users.noreply.github.com>
Co-authored-by: kokhlo <47825603+kokhlo@users.noreply.github.com>
2026-09-17 08:53:54 -07:00
kshitijk4poor
564687b113 fix(nous): classify the desktop's escaped-exception card with the request credential
build_error_surface_from_exception classified without the credential, so an anonymous
session's gate refusal that escaped the turn loop read as a retryable rate limit. Thread
api_key from the TUI gateway's terminal-error path. Provider-gate the named welcome-host
400 like the 403 branch, and drop the provider conjunct the anonymous gate already implies.
2026-09-17 12:28:25 +05:30
kshitijk4poor
f7da523ae6 refactor(nous): one is_anonymous_agent seam for the five identity reads
The same two-getattr expression was spelled three ways across turn_recovery, turn_api_call
and turn_response_check. Pass _Ctx.anonymous by keyword, and rename the guard's flag now
that it means "anonymous", not "welcome route".
2026-09-17 12:28:25 +05:30
kshitijk4poor
c0e604cd45 fix(nous): keep the reconnect copy for a named account on the welcome host
The identity gate made the gateway's mirror 400 ("serves anonymous Hermes Agent accounts
only") fall through to the generic format_error copy ("malformed request... hermes doctor"),
which is wrong for it. It is the one welcome-tier refusal a signed-in credential receives,
so classify it before the gate and keep "This Nous account needs to reconnect. Run /model
and pick the Nous row again"; only the free_tier sign-in card is withheld, since its single
action is a sign-in the user has already done.
2026-09-17 12:28:25 +05:30
Robin Fernandes
165271fceb fix(nous): keep anonymous errors exclusive to anonymous users 2026-09-17 12:28:25 +05:30
teknium1
4655263f35 fix(agent): apply the oversize content-field rule on the 422 path too
pydantic relays (#104731) report the same content-field detail with HTTP 422; the 422 handler
ran only the keyword rules, so a Nebius-shaped rejection behind such a relay would still be
claimed by the multimodal tool-content rule. Sibling site of #112473.
2026-09-16 17:09:30 -07:00
dacheah
270fe15c8b fix(agent): route oversize rejections that arrive as HTTP 400 to the image-shrink recovery
NVIDIA NIM ("Please make sure your payload is below 26214400 bytes"), Alibaba
DashScope ("String value length (N) exceeds the maximum allowed (M, from
`StreamReadConstraints.getMaxStringLength()`)") and Nebius Token Factory (a
pydantic `string_type` detail on `body.messages.N.content` whose rejected input
carries the inline image) enforce their byte caps with a 400, so the shrink
recovery keyed on 413 / image vocabulary never ran: the turn died as a
non-retryable format_error, or (Nebius) the generic "Input should be a valid
string" was claimed by the multimodal tool-content rule and its one retry was
spent stripping tool images that were never there.

- add the two wordings to _IMAGE_TOO_LARGE_PATTERNS
- structural rule for the pydantic detail (message content loc, not tool-scoped,
  rejected input holds a data:image part over the shrink target) checked ahead
  of the keyword multimodal rule in _classify_400
- settle_unrecovered_error: once the single shrink attempt was spent without
  recovering, treat image_too_large as a client error and fall back instead of
  re-sending the byte-identical oversized body max_retries times (a
  payload-scoped cap can be tripped by text alone)

Trimmed salvage of #111975 by @dacheah: exclusion-guard tables, the shrink-target
mirror constant and the dedicated fallback block are dropped in favour of the
existing client-error branch and an import of the real shrink target.

Fixes #112473
2026-09-16 17:09:30 -07:00
teknium1
857133babd refactor: trim empty-frame salvage to the translated error only
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
2026-09-15 18:18:08 -07:00
moxian
71281fb266 fix(agent): contentless SSE keepalive frames retry without streaming
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.

An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.

Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
2026-09-15 18:18:08 -07:00
kshitijk4poor
683ca44046 fix(classifier): a welcome-host 403 naming the free tier itself is the tier refusing, not a billing wall
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
2026-09-15 20:44:42 +05:30
Robin Fernandes
2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
KeyArgo
0968fa2631 fix(agent): classify NVIDIA NIM serde rejection of list-type tool content
NVIDIA NIM's Rust gateway 400s on list-type tool message content with a serde
error that names the enum rather than the field: "data did not match any
variant of untagged enum ChatCompletionRequestToolMessageContent". None of the
existing _MULTIMODAL_TOOL_CONTENT_PATTERNS match that wording, so the error
classifies as format_error (non-retryable) and the existing recovery —
downgrade the image-bearing tool message to text, remember (provider, model),
retry once — never fires: every retry re-sends the same shape and fails
identically in a session bound to that model.

Add the serialized enum name to the pattern tuple so the NVIDIA wording routes
to multimodal_tool_content_unsupported like the MiMo/Alibaba/Console Go
wordings already do (#111231).
2026-09-15 06:18:31 -07:00
KoNit-K
7d03ea3adf fix(agent): recover opencode zen encrypted replay 2026-09-15 06:18:31 -07:00
Fangliquan
51ebdff570 fix(codex): scope encrypted-reasoning replay to the issuing model
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.

Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.

Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.

Closes #95834
2026-09-15 10:49:19 +05:30
teknium
3ed62fc5e8 fix(agent): classify OpenAI spend/usage-limit error codes as billing
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
2026-09-13 20:49:13 -07:00