Commit Graph

116 Commits

Author SHA1 Message Date
teknium1
2768f31628 fix: step reasoning up to the floor when a route refuses to disable it
Endpoints that understand the reasoning field but refuse the OFF ("Reasoning is
mandatory for this endpoint and cannot be disabled" -- the Nous Portal on
gpt-6-astra) got two different wrong answers:

* auxiliary lanes (title generation sends reasoning_config={"enabled": False})
  400'd outright -- the #112781 strip rung only matched "unsupported/unknown
  field" wordings, so every session on such a route stayed untitled;
* the main loop dropped the disable and let the route pick its default effort,
  the opposite of what a thinking-off user asked for.

Both now step the effort up to the lowest level every reasoning wire accepts
(low) instead. The aux ladder gets a rung ordered before the strip
(agent/auxiliary_reasoning_floor.py: lifts reasoning_effort / extra_body.reasoning
/ _reasoning_config, memoises the (route, model) so the next thinking-off aux call
starts at the floor without the guaranteed 400); the main loop's reasoning_mandatory
recovery records whether the route said mandatory (floor) or unknown-field (drop,
unchanged). error_classifier.is_reasoning_required_rejection separates the two.

Live on the portal (proxy wire capture): title none->400->low->200 titled;
main loop none->400->low->200; second aux call on the route sends low up front.
2026-09-19 02:37:36 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
7530f40383 fix(errors): auth refusals from a non-stock route name the contacted host
The #113719 report asked that, when the request does fail, the error
names the endpoint that was actually contacted. classify_api_error now
appends "(endpoint: <host>)" to auth-class messages when base_url is set
and is not the provider's own route (provider_owns_route is not True),
so a credential posted to a stale model.base_url reads as a wrong
endpoint rather than a bad key. Stock routes and empty base_url keep
the plain message.
2026-09-18 12:49:40 -07:00
teknium1
95dee173e7 fix(agent): drop non-invariant disable-rung test; document thinking-state 400 trade-off (#114460)
test_disable_drop_rung_sets_session_flag_once stays green with agent/turn_recovery.py from
main (that hunk is comment/message text only), so it is not an invariant of this fix; the
spent-path test already carries the flag-false control. Docstring on
is_reasoning_field_rejection records the accepted trade-off for thinking-state 400s
("Function calling is not supported when thinking is enabled"): the marker sits 6 chars from
the token, so no proximity gate separates it from the forward wordings; cost is one
dropped-disable retry before the spent path falls back.
2026-09-18 10:21:05 -07:00
teknium1
a830061c9d fix(agent): recognise reversed "reasoning_effort 'none' unsupported" on every surface (#114460)
Reshape the salvage of #114461 (@whyyagswhy) so the reasoning-field
rejection matcher lives once, in agent.error_classifier, and both the
auxiliary retry ladder and the main conversation loop consume it:

- UNSUPPORTED_PARAM_MARKERS is the single marker tuple (was duplicated
  between auxiliary_client._is_unsupported_parameter_error and the new
  classifier helper); is_reasoning_field_rejection() replaces
  is_reasoning_disable_rejected() with the same token gate and a
  symmetric "unsupported" window so both word orders match
  ("unsupported reasoning_effort", "reasoning_effort 'none' unsupported").
- Main loop: the reasoning_mandatory rung message no longer claims the
  model "requires reasoning" (a chat-only relay does not); a second
  reasoning-field rejection in the same turn is treated as spent and
  takes the fallback chain instead of replaying the identical request
  max_retries times (mirrors image_too_large's shrink_spent).
- Tests trimmed to invariants (one per surface) plus the spent path;
  docs mention the reversed wording and the main-loop recovery.

Live against a stand-in replaying the Otari gateway's documented 400:
before, title generation failed and the thinking-only continuation died
with "Non-retryable client error"; after, both retry once without
reasoning_effort and complete.
2026-09-18 10:21:05 -07:00
Yagna Vudathu
74398c540c fix(agent): map reversed reasoning-disable rejection to reasoning_mandatory (#114460) 2026-09-18 10:21:05 -07:00
teknium1
6368ec6883 fix(aux): aux billing keywords share the main classifier table; quarantine logs name the real reason; vision fallback skips text-only models
Three auxiliary error-classification defects with one shared cause: the aux
ladder kept its own copies of the classifier's phrase tables and its own
fixed log wording, and they drifted.

- `_PAYMENT_KEYWORDS` is now `error_classifier._BILLING_PATTERNS` plus the
  aux-only phrasings, and `_BILLING_PATTERNS` gains OpenRouter's org cap
  "budget limit exceeded". A 403 "Budget limit exceeded (monthly limit)" (or
  "Key limit exceeded") is billing for the main loop AND a payment error for
  the aux ladder, so compression/title/vision fall through to the configured
  fallback instead of raising while the main model is served elsewhere.
- `_mark_provider_unhealthy` takes the caller's reason and level; the skip
  line echoes it. Absent credentials (no OPENROUTER_API_KEY, no Nous auth)
  are DEBUG with the real reason; the confirmed-402 rungs keep the WARNING
  payment wording. Local-only users no longer read "payment / credit error"
  for providers they never configured.
- `_try_main_agent_model_fallback` applies the auto-route's vision capability
  gate, so a transient 429 on the pinned vision lane no longer hands an image
  to a text-only main model (guaranteed 400 "content.type is invalid").
  `_recover_provider_pool` no longer benches a pool entry on an
  upstream-capacity 429 ("at capacity", same `_OVERLOADED_PATTERNS` table).

Fixes #107166, #64144, #108349.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: awizemann <319078+awizemann@users.noreply.github.com>
Co-authored-by: kokhlo <47825603+kokhlo@users.noreply.github.com>
2026-09-17 08:53:54 -07:00
kshitijk4poor
564687b113 fix(nous): classify the desktop's escaped-exception card with the request credential
build_error_surface_from_exception classified without the credential, so an anonymous
session's gate refusal that escaped the turn loop read as a retryable rate limit. Thread
api_key from the TUI gateway's terminal-error path. Provider-gate the named welcome-host
400 like the 403 branch, and drop the provider conjunct the anonymous gate already implies.
2026-09-17 12:28:25 +05:30
kshitijk4poor
f7da523ae6 refactor(nous): one is_anonymous_agent seam for the five identity reads
The same two-getattr expression was spelled three ways across turn_recovery, turn_api_call
and turn_response_check. Pass _Ctx.anonymous by keyword, and rename the guard's flag now
that it means "anonymous", not "welcome route".
2026-09-17 12:28:25 +05:30
kshitijk4poor
c0e604cd45 fix(nous): keep the reconnect copy for a named account on the welcome host
The identity gate made the gateway's mirror 400 ("serves anonymous Hermes Agent accounts
only") fall through to the generic format_error copy ("malformed request... hermes doctor"),
which is wrong for it. It is the one welcome-tier refusal a signed-in credential receives,
so classify it before the gate and keep "This Nous account needs to reconnect. Run /model
and pick the Nous row again"; only the free_tier sign-in card is withheld, since its single
action is a sign-in the user has already done.
2026-09-17 12:28:25 +05:30
Robin Fernandes
165271fceb fix(nous): keep anonymous errors exclusive to anonymous users 2026-09-17 12:28:25 +05:30
teknium1
4655263f35 fix(agent): apply the oversize content-field rule on the 422 path too
pydantic relays (#104731) report the same content-field detail with HTTP 422; the 422 handler
ran only the keyword rules, so a Nebius-shaped rejection behind such a relay would still be
claimed by the multimodal tool-content rule. Sibling site of #112473.
2026-09-16 17:09:30 -07:00
dacheah
270fe15c8b fix(agent): route oversize rejections that arrive as HTTP 400 to the image-shrink recovery
NVIDIA NIM ("Please make sure your payload is below 26214400 bytes"), Alibaba
DashScope ("String value length (N) exceeds the maximum allowed (M, from
`StreamReadConstraints.getMaxStringLength()`)") and Nebius Token Factory (a
pydantic `string_type` detail on `body.messages.N.content` whose rejected input
carries the inline image) enforce their byte caps with a 400, so the shrink
recovery keyed on 413 / image vocabulary never ran: the turn died as a
non-retryable format_error, or (Nebius) the generic "Input should be a valid
string" was claimed by the multimodal tool-content rule and its one retry was
spent stripping tool images that were never there.

- add the two wordings to _IMAGE_TOO_LARGE_PATTERNS
- structural rule for the pydantic detail (message content loc, not tool-scoped,
  rejected input holds a data:image part over the shrink target) checked ahead
  of the keyword multimodal rule in _classify_400
- settle_unrecovered_error: once the single shrink attempt was spent without
  recovering, treat image_too_large as a client error and fall back instead of
  re-sending the byte-identical oversized body max_retries times (a
  payload-scoped cap can be tripped by text alone)

Trimmed salvage of #111975 by @dacheah: exclusion-guard tables, the shrink-target
mirror constant and the dedicated fallback block are dropped in favour of the
existing client-error branch and an import of the real shrink target.

Fixes #112473
2026-09-16 17:09:30 -07:00
teknium1
857133babd refactor: trim empty-frame salvage to the translated error only
Drop the raw JSONDecodeError belt from _is_provider_stream_empty_frame_error:
every chat stream iteration goes through _iter_provider_stream_chunks, which
already translates the decode failure, so the belt guarded a path that does
not exist. Drop the explicit classifier entry for the new code: unknown codes
already resolve to the same retryable unknown verdict (verified live: identical
ClassifiedError with and without the entry).
2026-09-15 18:18:08 -07:00
moxian
71281fb266 fix(agent): contentless SSE keepalive frames retry without streaming
An empty `data:` frame (or a frame carrying only `event:` / `id:`) is a legal SSE
no-op, but the OpenAI SDK still hands it to `json.loads`, which raises
`JSONDecodeError` with an empty document. The streaming helper translated that
into `ProviderStreamError(provider_stream_non_json_data)`, nothing recognised it
as recoverable, and the main loop retried streaming -- identically -- until the
retry budget ran out: "API call failed after 3 retries: Provider stream returned
non-JSON SSE data". A degraded gateway answers EVERY streaming request that way,
so the retries only repeated the failure while the same request sent
non-streaming succeeded seconds later.

An empty document means no payload, which is a different fact from a malformed
payload: give it its own code, and when it appears before any delta switch the
session to non-streaming (the existing `_disable_streaming` mechanism, already
used for "stream not supported", Bedrock IAM denials and adapter-returned final
responses) so the retry goes out on a channel the degraded gateway can answer.
Non-empty payloads keep today's fatal semantics; failures after deltas keep the
existing stream-drop handling.

Tests: the turn-level recovery test drives the real SDK decoder over a real
httpx response and is red on base (three identical streaming attempts, turn
lost); the boundary test pins the malformed-payload path so a future "ignore bad
frames" change cannot swallow genuine provider errors.
2026-09-15 18:18:08 -07:00
kshitijk4poor
683ca44046 fix(classifier): a welcome-host 403 naming the free tier itself is the tier refusing, not a billing wall
The "says something else" guard reuses the billing table, which carries the
Nous gateway's own free-tier phrases ("not available on the free tier",
"model_not_supported_on_free_tier"). On the welcome route those words mean
exactly "the tier refused"; routing them to billing prints a credits check to
an anonymous session that has none. Leave those two phrases on the
tier_disabled path.
2026-09-15 20:44:42 +05:30
Robin Fernandes
2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
KeyArgo
0968fa2631 fix(agent): classify NVIDIA NIM serde rejection of list-type tool content
NVIDIA NIM's Rust gateway 400s on list-type tool message content with a serde
error that names the enum rather than the field: "data did not match any
variant of untagged enum ChatCompletionRequestToolMessageContent". None of the
existing _MULTIMODAL_TOOL_CONTENT_PATTERNS match that wording, so the error
classifies as format_error (non-retryable) and the existing recovery —
downgrade the image-bearing tool message to text, remember (provider, model),
retry once — never fires: every retry re-sends the same shape and fails
identically in a session bound to that model.

Add the serialized enum name to the pattern tuple so the NVIDIA wording routes
to multimodal_tool_content_unsupported like the MiMo/Alibaba/Console Go
wordings already do (#111231).
2026-09-15 06:18:31 -07:00
KoNit-K
7d03ea3adf fix(agent): recover opencode zen encrypted replay 2026-09-15 06:18:31 -07:00
Fangliquan
51ebdff570 fix(codex): scope encrypted-reasoning replay to the issuing model
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.

Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.

Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.

Closes #95834
2026-09-15 10:49:19 +05:30
teknium
3ed62fc5e8 fix(agent): classify OpenAI spend/usage-limit error codes as billing
Clean-room port of the billing-code coverage from zed-industries/zed#63208: credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded, organization_usage_limit_exceeded now classify as billing (rotate + fallback) instead of falling through to generic buckets.
2026-09-13 20:49:13 -07:00
Teknium
92df11f81d fix(classifier): terminal_quota_exhausted 429s classify as billing, not rate_limit
Port from code-yeongyu/oh-my-openagent#6677 (credit: @niStee).

LiteLLM proxies stamp a structured `terminal_quota_exhausted` code on
hard-cap 429s. Hermes' `_status_429` handler always returns a verdict, so
`_by_error_code` (which maps _BILLING_ERROR_CODES to billing) never saw
the code: the exhausted key classified as rate_limit, earned the 429
cooldown, and got retried against a wall that cannot clear until someone
pays. Upstream this respawned duplicate subagent sessions.

- `_status_429` now honors a structured billing code first (decisive
  signal outranks message heuristics).
- `terminal_quota_exhausted` joins _BILLING_ERROR_CODES so every path
  (429, 402, status-less) agrees.
- "hard billing limit" free text joins _BILLING_PATTERNS ("billing hard
  limit" was already there; providers use both orders). "terminal billing
  limit" text is deliberately NOT matched: substring rules cannot negate
  the "non-terminal billing limit" wording — the structured code covers it.
2026-09-12 21:52:48 -07:00
teknium1
293cb54f28 fix(agent): keep MoA verdicts fallback-free; nest the gated fallback under one if
Two CI regressions from the previous commit, both mine:

- `_moa_special_cases` was switched to `_ABORT_FALLBACK`, which flips
  `should_fallback` to True for the MoA adapter-shape and missing-preset
  verdicts. #55933 made those deliberately NOT fall back (a fallback would
  silently replace the MoA route with a single model); restore
  `retryable=False` only. The gate in `settle_unrecovered_error` now honours
  that for real: on main these verdicts never reached the fallback branch
  because the flag was never consulted.
- `tests/agent/test_prompt_cache_ttl_propagation.py` pins every
  `_try_activate_fallback` reference to a direct `if agent._try_activate_fallback():`
  site (#84733 restart discipline); `fallback_allowed and agent._try_activate_fallback()`
  broke that shape. Nest the two calls under one `if should_fallback or local_validation:`
  block instead.
2026-09-12 08:27:11 -07:00
teknium1
dfd4aa4a94 fix(agent): malformed tool-call-argument 400s no longer walk the fallback chain
When a proxy (Ollama, OpenRouter) rejects the MODEL's own unparseable
tool-call JSON with `400 invalid tool call arguments`, the classifier returned
the generic format_error verdict (`should_fallback=True`) and the non-retryable
client-error path cascaded through every fallback provider: 4-5 sequential
calls, 20-60s per occurrence, ending on a model that produced the same broken
JSON (#12770).

- error_classifier: explicit `_MALFORMED_TOOL_ARGS_PATTERNS` checked before the
  request-validation and overflow heuristics, returning format_error with
  `retryable=False, should_fallback=False`.
- turn_api_error: the client-error settlement honours `should_fallback`; the
  verdicts that legitimately reach that branch (policy block, TLS chain, MoA
  shape/preset errors) now state `should_fallback=True` explicitly, so the gate
  changes behaviour only for the new verdict. Local validation errors keep
  their historical fallback.

Fixes #12770. Pattern list and gating approach from #16022 by @cuyua9 (stale
base); tests trimmed to two invariants.

Co-authored-by: cuyua9 <2114364329@qq.com>
2026-09-12 08:27:11 -07:00
Siddharth Balyan
4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Erosika
5db0df28de fix(classifier): treat Azure Foundry's continuation-identity 400 as a rejected reasoning replay
Azure Foundry (gpt-6-astra) answers a request that replays encrypted
reasoning from more than one prior response with HTTP 400 "Conflicting
authenticated continuation identities", code invalid_value. The classifier
mapped that to a non-retryable client error, so the turn aborted and every
later turn in the session failed the same way.

Map the message to invalid_encrypted_content. The existing recovery in
turn_recovery then strips the replay state and retries once.

Refs #105369
2026-09-09 13:01:48 -07:00
liuhao1024
d3515a9d18 fix(agent): classify Codex patch-budget image 400s as image_too_large
OpenAI Codex Responses rejects an image whose tile-patch budget exceeds
its 30000-patch ceiling with wording ("requires N patches after
processing, exceeding the limit") that contains none of the existing
image-size vocabulary, so the 400 fell through to format_error
(non-retryable). The reactive image-shrink recovery in turn_recovery was
therefore bypassed and the session kept failing — or failover re-sent
the identical oversized image to another model.

Route the patch-budget wording to image_too_large (retryable) so the
existing shrink pass re-encodes the image under the ceiling and retries.

Fixes #106337
2026-09-09 12:06:53 -07:00
KeyArgo
43ecad8fe6 fix(agent): classify Novita 'server overload' 429 as overloaded
Novita returns HTTP 429 with message 'server overload, please try again
later' and error type 'server_overload' when its server is genuinely busy.
Neither phrase was in _OVERLOADED_PATTERNS, so the 429 fell through to
_V_RATE_LIMIT and set should_fallback=True + should_rotate_credential=True —
rotating the credential / falling back early instead of retrying the same
key. Add 'server overload' and 'server_overload' to the overload tuple so
this reaches the existing _V_OVERLOADED verdict (retryable, no rotation).
Closes #106205
2026-09-09 10:47:49 -07:00
Teknium
e1838c5b5a fix: trim opencode-go 422 salvage to the invariant set
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
2026-09-09 03:52:47 -07:00
ericmaddox
bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Tim Kaufmann
b7e4712029 fix(agent): classify local-inference memory-ceiling rejections as overloaded
oMLX/MLX prefill memory-guard rejections name an allocation peak in BYTES but
close with "Reduce context length", so _CONTEXT_OVERFLOW_PATTERNS claims them
and the turn enters the compress-and-shrink loop. Compression cannot lower a
prefill peak — the prompt is usually far below the window — so it burns the
compression budget, re-hits the wedged server on every attempt and ends in
"Cannot compress further" plus a destructive session reset.

Classify them as `overloaded` instead: retry with backoff, no compression, no
session reset (mirrors 503/529 recovery).

The guard runs before the overflow check AND before the usage-limit
disambiguation. The second ordering matters more than it looks: "memory limit
exceeded" contains "limit exceeded", so a status-less rejection — a proxy that
flattened the body — is currently classified as `billing` and rotates a
healthy credential.

Sites covered:
  - _OVERFLOW_AS_5XX_RULES, which _400_TAIL_RULES extends → 400, 500, 502,
    503, 529
  - _MESSAGE_HEAD_RULES for the status-less path (ahead of usage-limit)
  - _ERROR_CODE_VERDICTS for the structured oMLX codes
  - _classify_400, because _by_status runs before _by_error_code, so a body
    whose wording a proxy stripped would otherwise fall through to
    format_error

Every pattern names memory/allocation in bytes, never a token or window count,
so the list stays disjoint from _CONTEXT_OVERFLOW_PATTERNS. Both oMLX wordings
are kept: 0.5.6 says "Prefill would require ~13.87 GB peak", 0.5.7 reworded it
to "predicted peak would require/exceed" and both are in the field. A test
pins that a genuine window overflow still compresses.

Refs #52261. Supersedes #52289, which predates the classifier rewrite and can
no longer be merged.
2026-09-09 12:17:29 +05:30
Xipong
190c73328c chore(aux): carry local layer evolution onto current main 2026-09-06 05:48:39 -07:00
Teknium
e0592b6d43 fix(codex): masked "invalid_prompt: Request blocked." replay rejection reaches the replay-strip recovery
The ChatGPT Codex backend answers a rejected encrypted-reasoning replay with the
same bare 400 {code: invalid_prompt, message: "Request blocked."} it uses for
genuine blocks. classify_api_error bucketed it as format_error, so the one-shot
repair that already exists for invalid_encrypted_content (disable replay, strip
cached codex_reasoning_items, retry once) never ran and the turn aborted.

Classify exactly that envelope from provider openai-codex — SDK 400 body, SSE
``error`` frame, or the ``response.failed`` text — as invalid_encrypted_content
while keeping format_error's abort-and-fallback hints, so the only behavioural
delta is turn_recovery's replay strip, which still requires cached reasoning
items. A block with nothing to strip aborts exactly as before; other providers,
other messages/codes and the #18028 safety refusals are untouched.

evals/codex_masked_replay_review.py drives a real AIAgent (codex_responses)
against a synthetic local Responses server for the before/after check.
Refs #92353, #92357, #94077.
2026-09-06 05:40:02 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
022785a541 Merge origin/main (63279301bc): reasoning-mandatory 400 recovery folded into turn_recovery/error_classifier/models_reasoning_caps 2026-09-03 05:36:21 -07:00
Teknium
f6bd1633f7 fix(reasoning): GLM-5.3 on Nous/OpenRouter no longer 400s when thinking is disabled
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.

- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
  no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
  queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
  (configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
  background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
  mandatory (parity with the Nous profile).

Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
2026-09-03 05:20:11 -07:00
Teknium
ea3a1cec8c refactor(agent/error_classifier,display,i18n): flatten remaining guard ladders and log-call formatting 2026-09-02 19:28:25 -07:00
Teknium
b4b3baeda1 refactor(agent/display,error_classifier): simplify shell tokenizers, share abort-fallback verdict hints 2026-09-02 19:19:34 -07:00
Teknium
b9cd14d755 refactor(agent/error_classifier): derive _Ctx fields in __post_init__, collapse extractor guards 2026-09-02 19:06:49 -07:00
Teknium
eac3b11dde refactor(agent/error_classifier,display): compact WHY comments to 1-3 lines, collapse guard ladders and small wrappers 2026-09-02 18:49:20 -07:00
Teknium
a5129759e3 refactor(agent/error_classifier): pack pattern/rule tables, unify cause-chain + OpenRouter body readers 2026-09-02 18:27:55 -07:00
Teknium
ebd4b1002b refactor(agent/error_classifier): stage pipeline + status dispatch table, verdict dicts, 0 fuzz mismatches 2026-09-02 18:14:19 -07:00
Teknium
695f7da3cf refactor(agent): error_classifier — rule-table verdicts, shared match helpers, compacted pattern rationale (2244 -> 1357 LOC) 2026-09-02 13:29:36 -07:00
Koduri Mahesh Bhushan Chowdary
b3f4f50771 fix(agent): classify "media exceeds size limit" as image_too_large
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.

That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.

Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.

Fixes #76039
2026-08-28 04:58:06 -07:00
Sora-bluesky
564113572e fix(agent): classify xAI's downloaded-response wording as a corrupt image
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.

Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.

Reported by ryuhaneul in #69078.
2026-08-28 03:46:24 -07:00
Sora-bluesky
d61411d131 fix(agent): narrow #69078 image-corrupt recovery to the classifier route
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.

Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.

Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.

Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky
8aeb3f6ee3 fix(agent): un-brick sessions on non-retryable 400s that carry image parts
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.

Two recovery layers, deliberately separate:

- Semantic split: new FailoverReason.image_corrupt with
  _IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
  checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
  _classify_by_message. Corrupt bytes route to strip-and-retry, never
  to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
  still contain image parts gets one strip-and-retry via the existing
  _strip_images_from_messages helper, guarded by a new
  stripped_images_this_turn one-shot flag on TurnRetryState. This
  un-bricks the session for any current or future provider wording
  without adding another pattern list to maintain.

Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Teknium
0c3a507535 fix(classifier): 429 quota walls route to billing across providers; reset signals stay rate-limited
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:

- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
  the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
  exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
  credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
  exclusion so an explicit 'Rate limit exceeded' never promotes to
  non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
  reached. Resets in 6hr 29min.' wrongly read as terminal billing because
  main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
  signals. (credit @jtstothard, #74785)

Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.

Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.

Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
2026-08-23 20:02:07 -07:00
fangliquanflq
654d537088 fix(agent): honor structured quota reset signals 2026-08-23 18:43:12 -07:00