Two gaps for custom providers behind a WAF/CDN:
- `build_anthropic_client` never consulted `custom_providers[].extra_headers`,
so a relay in `anthropic_messages` mode that rejects the SDK User-Agent kept
403ing even with `extra_headers: {User-Agent: ...}` configured, while the
OpenAI-wire clients already applied it. The lookup now lives in
`_new_sdk_client`, the one constructor every builder path goes through
(init, /model switch, rebuild, auxiliary), keyed by the caller's raw route
because entries are keyed by the `/v1` form the normalizer strips.
Salvaged direction of #46002 (@wait4xx). Fixes#24293, #9721.
- `_status_403` classified every non-billing 403 as `auth`, so a WAF's plain
"Your request was blocked." or a Cloudflare browser challenge printed "Your
API key was rejected" and could rotate a healthy credential. A 403 carrying
established block/challenge markers is now `upstream_blocked`: no rotation,
no retry, fallback allowed, WAF/User-Agent guidance on every surface (CLI
loop, chat copy, cli chat error copy, TUI gateway + Ink TUI copy). Generic
403 and all 401 keep the auth verdict. Salvaged direction of #70567
(@ooiuuii) and #53114 (@AgenticSpark). Fixes#53099, #70566.
Review follow-ups on the inline-image guard:
- `image/jpg` is the JPEG alias every other image site accepts
(vision_message_prep, conversation_compression, image_gen_provider,
mcp_tool_content); the Responses guard downgraded it to a text
placeholder, losing a valid image. It now counts as JPEG.
- The Anthropic converter forwarded data:image/svg+xml (bmp, tiff) verbatim
as media_type, which 400s every turn once the part is in history. It now
applies the same rule: SVG is rasterized to PNG when a rasterizer exists,
any other unsupported inline subtype becomes a text placeholder, and
image/jpg is normalized to image/jpeg.
Both wire paths share one helper next to the existing supported-set
constant in tools/vision_tools_image_prep (import-time deps: hermes_constants
only, no cycle).
A/B: new Anthropic test and the jpg assertion red on the previous head,
green now.
The send-layer guard from c9f8cb72a6a downgrades every inline
data:image/svg+xml part to a text placeholder, so the model never sees a
drawing the agent (or a user attachment) wanted it to look at. The
reporter's ask was to rasterize instead of drop when that is possible.
Reuse the vision_analyze rasterizer chain (cairosvg / svglib+reportlab /
rsvg-convert / inkscape, all soft deps) through a small
rasterize_svg_data_url helper in tools/vision_tools_image_prep: the SVG
payload is decoded to a temp file, rasterized, re-encoded as a PNG data URL
and forwarded as input_image; the temp files are removed. Without a
rasterizer the existing placeholder still applies, so the request can never
carry SVG source. Import is lazy inside _input_image_part (no cycle; the prep
module only depends on hermes_constants at import time).
Live (real converter on a temp HERMES_HOME): base with cairosvg installed ->
input_text placeholder; fixed with cairosvg -> input_image
data:image/png (PNG magic verified); fixed without any rasterizer ->
placeholder.
A `data:image/svg+xml` (or BMP/TIFF/...) part reaching the Codex Responses
converter was forwarded verbatim as `input_image`; the backend rejects the
WHOLE request with 400 "The image data you provided does not represent a
valid image", and because the part is baked into history (a persisted
vision_analyze tool result, a user turn) every later continuation re-trips
the same 400 until the user drops the turn.
The inputs that create SVG parts are already guarded (image_routing skips
SVG for native attach; vision_analyze rasterizes it), so this is the
converging send-layer seam: `_input_image_part` now downgrades any inline
image whose subtype is outside jpeg/png/gif/webp to an `input_text`
placeholder. It is the single seam for user/tool message content,
`function_call_output.output`, and the preflight validator, so all three
carriers are covered; remote http(s) URLs pass through untouched (the
provider owns their validation). Valid images in the same message still go
as `input_image`.
The 400-recovery matcher already recognises the Codex wording since
0241619068, so only the send side was missing.
Docs: vision.md now documents `agent.image_input_mode` (auto|native|text)
and the explicit `auxiliary.vision` override — the way to keep an
openai-codex main model while routing image analysis to another vision
provider when the backend answers image requests with server_error.
Salvages #47299 (@hanzckernel) — same fix direction, redone against the
current converter (the original diffs predate the `_input_image_part` seam).
With a home channel configured, a deliver=origin job created over api_server rerouted
there silently and the create response said nothing, so the agent could still promise
"I'll report back here" to the HTTP client. _local_delivery_notice now returns a one-line
note naming the origin_fallback target when the session cannot receive async delivery.
hermes_cli/suggestions_cmd._resolve_origin used its own env mirror without the
async_delivery guard, so /suggestions accept still stamped platform=api_server; it now
delegates to tools.cronjob_job_args._origin_from_env.
A cron job created from an api_server turn captured origin.platform="api_server".
That adapter's send() is a stub (request/response only, supports_async_delivery=False),
so every fire ran fine (last_status=ok) and then recorded
"API server uses HTTP request/response, not send()" — nothing reached the creator and
nothing warned them at creation time (#69304).
Producer: _origin_from_env() stamps no origin when the session cannot receive async
delivery, so the job takes the existing no-origin path — home-channel fallback at fire
time and the creation-time "local-only" notice (its wording now names stateless HTTP API
sessions alongside CLI/TUI). Fire time: _resolve_origin() treats an already-stamped
api_server origin as missing so existing jobs.json entries fall back too.
Co-authored-by: Chen Jin <243284244@qq.com>
wait_notice_text and WaitNoticeState.should_emit now both derive near-deadline
from the watchdog tuple, so the two call sites stop recomputing it and passing
it as a kwarg. Two docstrings still described the removed copy as the 60s notice.
providerWaitText only matched frames starting with 'waiting on', so the
near-deadline update minted by agent/chat_completion_wait_notice.wait_notice_text
('⏳ still waiting on <model> — …') was rejected and Desktop kept the stale first
notice until the reconnect. Widen the matcher; pin it with the exact string the
helper produces.
Both wait-status builders — the Codex Responses request poller
(chat_completion_nonstream) and the chat-completions stream monitor — rewrote
the status line on every 30s heartbeat once a request had been silent for 60s,
always with "provider may be slow or overloaded" and an unlabeled
"auto-reconnect at Ns". Successful long calls (p95 ~53s on large contexts)
therefore looked unhealthy, and a genuine zero-event stall produced the same
text, so the two states were indistinguishable.
The 30s heartbeat stays (it is gateway liveness) but the visible notice now
comes from one shared, presentation-only module (chat_completion_wait_notice):
- emitted once when the silence crosses the threshold, then only when the wait
phase changes or the applicable watchdog is within 15s of firing;
- neutral phase wording: "waiting for the first provider event" (also
"...after reconnect") vs "provider stream active; Ns without stream events";
stream path: "waiting for the first stream chunk" vs "stream open; Ns
without stream output";
- the reconnect hint names the watchdog (TTFB / stream idle / wall-clock stale
/ stream stale) and the seconds left before it fires, which is unambiguous
across the elapsed-vs-silence timelines the old "at Ns total elapsed" mixed.
No watchdog thresholds or retry behaviour change. The old
_codex_wait_notice_recovery helper (deadline as a bare string) is replaced by
codex_watchdog_deadline (label + seconds remaining). Design ported from the
candidate fix in #92657 (phase-keyed dedupe, labeled deadlines, imminent
window), which targeted a pre-refactor call site that no longer exists.
Fixes#92550
Co-authored-by: chelsealong <chelsealong@126.com>
The dedupe test now drives conversation_loop.run_conversation with a native-parts
user message and a stubbed API whose tool loop calls vision_analyze on the attached
path, asserting the already_in_context text result (replaces the direct-call scope
test; replacing the call site by nullcontext() goes red). vision.md notes that a
same-turn native attachment is not re-embedded and that region still zooms.
When a surface (Telegram gateway, CLI, TUI, delegated child) attaches an
image natively to the user turn and the model then calls vision_analyze on
the same path, the native fast path embedded the identical pixels a second
time as a multimodal tool result in the same request (#76411).
conversation_loop.run_conversation now scopes the turn's native image
handles (the [Image attached at: ...] hints build_native_content_parts
writes) in a ContextVar for the duration of the turn; the tool answers with
a short text result instead. Region crops and other images still embed;
the scope ends with the turn so later turns can re-load after compression.
compression.context_timeout_seconds (default 120s, floored at the effective
auxiliary.compression.timeout, itself >= 300s) and hygiene_timeout_seconds
(default 30s) abort a silent summary before a larger no_progress_timeout can
fire (#108104).
test_codex_client_receives_configured_no_progress_timeout monkeypatched
_get_task_no_progress_timeout and is subsumed by
test_real_config_value_reaches_the_stream_guard_per_task, which drives the real
config.yaml -> call_llm -> _CodexStreamGuard path.
Follow-up to the salvaged #108119 hunk that adds
auxiliary.<task>.no_progress_timeout: a value that is not a positive number
used to fall back to the 60s default silently, which is exactly the
"my 600s request aborted after 60s and nothing told me why" confusion the
key exists to remove. The resolver now logs a warning naming the task and
the rejected value before applying the default.
Documents the key on the configuration page (independent of
auxiliary.<task>.timeout; keepalive frames do not re-arm; capped at the
request timeout; host deadline/cancel still win; per-task scope) and adds a
real-config invariant test that drives call_llm through the genuine
CodexAuxiliaryClient path on a temp HERMES_HOME.
Fixes#108104
auxiliary.compression.timeout only bounds the overall request; the
substantive-progress window inside _CodexStreamGuard was hardcoded to 60s,
so raising the overall timeout couldn't widen the actual stall window that
aborts slow-but-live reasoning/summary streams.
Add auxiliary.<task>.no_progress_timeout (default: unset, keeps the
existing 60s behavior). It's threaded only to Codex/Responses-shim clients
(CodexAuxiliaryClient / AsyncCodexAuxiliaryClient) — real OpenAI-SDK-shaped
clients don't accept the extra kwarg.
Fixes#108104
OpenAICompatibleVideoGenProvider has a configurable <NAME>_BASE_URL, so a
local/custom endpoint hit the same macOS system-proxy symptom as the image
provider (#64888 sibling). Route it through build_keepalive_http_client(base_url);
one generate()-level test with a plain httpx.Client() control.
httpx 0.28 binds getproxies at import (httpx._utils), so patching
urllib.request.getproxies never reached it and the test stayed green with
_build_client neutered to a plain httpx.Client(). Patch httpx._utils.getproxies,
assert a plain httpx.Client() control DOES mount the fake system proxy, and
assert the http_client generate() hands openai.OpenAI carries no HTTPProxy
mount. Red with the neutered client, green on the branch.
Custom model ids (#97928): a value of image_gen.openai.model / OPENAI_IMAGE_MODEL outside the
gpt-image-2 / gpt-image-2.5 catalog used to fall through silently to the default tier, so a
gateway serving its own image model names always received `model: gpt-image-2` +
`quality: medium`. resolve_static_model(passthrough=True) now returns the unknown id as the
API model with quality=None, and the request omits `quality` for it — OpenAI-compatible
gateways reject enum values they do not know, and a foreign model has no OpenAI quality
tiers. Only the provider-scoped key and env var pass through; the shared top-level
image_gen.model can hold another provider's id (a FAL path) and is never forwarded.
Named custom endpoint (#83080, config-reuse half): image_gen.openai.provider names a
providers:/custom_providers: entry; its base_url and api_key/key_env fill in whatever
image_gen.openai.base_url / key_env leave unset, so a gateway already declared for chat is
not re-declared with a duplicated key. Explicit image_gen.openai values keep precedence; an
unknown name logs a warning and falls back to the OpenAI variables. The lookup goes through
hermes_cli.runtime_provider._get_named_custom_provider, the same resolver the auxiliary
clients use, so aliases and legacy list entries behave identically.
image_gen is an open config root (deliberately absent from DEFAULT_CONFIG), so
`image_gen.openai.provider` already validates as a known key; docs cover both behaviours.
Live probe (fake /v1/images/generations gateway recording the body, real provider on a temp
HERMES_HOME): before — `custom-image-model` sent as model=gpt-image-2 quality=medium; named
provider → is_available False / auth_required. After — model=custom-image-model, no quality;
named provider → request to the entry's URL with `Bearer <its key_env>`; the catalog model
still maps to gpt-image-2 + quality=medium.
What: plugins/image_gen/openai resolves its endpoint and credential through one
resolver — image_gen.openai.base_url → OPENAI_BASE_URL → SDK default, and the env
var named by image_gen.openai.key_env → OPENAI_API_KEY — shared by is_available()
and generate() so the two cannot disagree. The client is built on
build_keepalive_http_client (env-only proxy policy) and sends a blank
OpenAI-Project header.
Why: the image endpoint could only be routed via the process-wide OPENAI_BASE_URL /
OPENAI_API_KEY, so a local or third-party image gateway could not be configured
independently of the chat provider (#65309, #97928, #13798). openai.OpenAI() with
trust_env routed localhost endpoints through a macOS system proxy whose
ExceptionsList httpx never sees (#64888). An OPENAI_PROJECT_ID set for chat made
/v1/images/generations 403 model_not_found on projects with a model allow-list
even though the key already carries the project (#60748).
Slim redo of the contributor direction in #18796 (@y0shua1ee), #37208/#37209
(@charzhou), #65312/#65323/#64893 (@asdlem), #60749 (@perelin); all predate the
StaticImageGenProvider refactor and no longer apply.
_anthropic_stream_created now records the attempt's httpx response and calls
_reabort_if_cancelled like the chat_completions wire, so an interrupt that raced
create()'s connect/TLS window shuts down the socket once headers arrive on
anthropic-protocol endpoints (incl. anthropic-compatible custom ones); the
re-abort picks the Anthropic slot sweep for that api_mode.
Tests: the direct _chat_stream_created test is replaced by one driving
interruptible_streaming_api_call (monitor _abort_for_interrupt during the connect
window -> on_stream_created wiring), parametrized over both wires.
The /stop and /reset abort path shut down only the sockets the request
client's pool sweep could see. Two shapes escaped it (#98974): the
connection checked out for the in-flight body read (the pool sweep skips
it — the stale-kill path already covered this via
_shutdown_stale_attempt_socket, the interrupt path did not), and an abort
that fires during create()'s connect/TLS window, when no socket exists
yet, so the request came up afterwards with nothing to stop it. The log
said "no sockets found; tcp_force_closed=0" and the self-hosted serve kept
generating into a dropped consumer.
Now _abort_for_interrupt also shuts down the attempt's own socket, and
_chat_stream_created re-aborts when the attempt was cancelled before the
response arrived. Both remain shutdown-only (never a cross-thread close,
#30858): the worker unwinds and releases descriptors on its own thread.
Direction shared with #98989 (@liuhao1024); the re-abort-on-create hunk
is reimplemented against the current _StreamingCall shape.
The 'completed response instead of an iterator' replay branch read reasoning only
through attributes; apply the same model_extra fallback so reasoning from
non-SDK message objects is still emitted (#56516).
_ChatStreamAccumulator.feed read reasoning only via getattr(delta, 'reasoning') /
('reasoning_content'); a non-SDK delta carrying reasoning only in model_extra
registered no progress and dropped the text. Same two-key fallback as the main
streaming path (#56516); one test drives _create_with_progress.
The chat-completions stream intake read reasoning from `delta.reasoning_content`
/ `delta.reasoning` attributes only. OpenAI-compatible gateways that expose the
field as an unknown extra land it in the SDK's `delta.model_extra`, so a
reasoning-only stream accumulated nothing and either tripped
"Provider returned an empty stream with no finish_reason" or dropped the
reasoning entirely on finish_reason=length. Mirror the `model_extra` fallback
the non-streaming path (and the streaming reasoning_details/refusal readers)
already use, for both carrier keys.
Salvages #56789 by @Xingkai98 (ported onto the phase-split intake; the original
test asserted a synthetic "stop" on a finish-less stream, which current main
deliberately reports as a drop, so the invariant test uses finish_reason=length
from the issue's SSE trace).
Fixes#56516
A fallback whose every entry sits in an exhaustion cooldown (the weekly quota
case, resets_in_seconds=30995) was still switched to and failed the turn the
same way the primary did. _should_skip_fallback_candidate now consults the
candidate's pool before any client is built and skips it when the earliest
recovery is further out than the retry loop's 600s Retry-After cap; a short
throttle still gets its chance. (#89401 atom 2)
_summarize_api_error reduces the usage_limit_reached body to 'HTTP 429: The
usage limit has been reached', so no text-based parser could ever reach the
8.6h reset from the in-loop failure path. _status_429 now stamps
error_context['reset_at'] from the body's reset fields, Retry-After or the
message grammar (same table the credential pool uses), and
max_retries_exhausted_result hands the remaining seconds to exhausted_copy,
which says 'its usage limit resets in ~9h. Send /retry after that' once the
window is >= 2 minutes. A throttle with no window keeps the short-wait copy.
Proven with a real openai.RateLimitError carrying resets_in_seconds=30995
driven through the production result builder. (#89401 atom 1)
The tightened 401 token dropped 'HTTP 401: Unauthorized' (the exact shape
_summarize_api_error emits), 'returned 401.' and 'HTTP 401, ...' — every real
envelope fell to the generic 'kept failing' reply with no /login hint. Only the
lookbehind is needed for the '05:14:15,401' timestamp case: (?<![\d:,.])401(?!\d).
Control assertions cover both directions.
The gateway's private reset-seconds regex is gone: 'resets_in_seconds' joins
agent.retry_utils.RETRY_DELAY_PATTERNS and _rate_limit_reply calls
reset_delay_from_message, so 'resets in 4hr 5min' weekly-limit text now names
its window too, and format_reset_window is the one renderer.
The direct-call user_text() test is replaced by one driving TurnRunner.run_sync
with _resolve_session_agent_runtime raising the codex_rate_limited RuntimeError;
it goes red when the is_rate_limited branch is removed. (#89401)
Rate-limit now wins over the auth pattern in the gateway's provider-error
reply table, `401` only matches as a standalone status token (a bare
`\b401\b` also hit timestamp fragments like `05:14:15,401`), and the
rate-limit reply names the reset window when the envelope carries
`resets_in_seconds` or `retry after Ns` instead of "wait a moment" for a
weekly quota. The pre-turn credential-resolution failure on chat surfaces
and the api_server `_ProviderAuthResolutionError` label follow the same
rule via `is_rate_limited_auth_error` on the cause chain.
WHY: a quota/429 envelope often also carries an auth-shaped preamble and
"Credentials are still valid"; classifying it as auth sent operators to
re-login working credentials across profiles (#89401).
'Next, I'll create the script.' / 'First, let me check the directory.' /
'Okay — running the tests.' followed by a closing {"cmd": ...} object were not
classified as leaked tool calls because the lead-in pattern anchored the action
verb at line start. Allow an optional Next/First/Then/Okay/OK/Alright marker
(with comma/dash) before the existing prefixes; the 'closes the message' and
'cmd' key constraints are unchanged, so bare/explained JSON answers still pass
through (#56920).
gpt-5.x on the Codex Responses backend sometimes serializes the tool call it
meant to make as Codex-CLI shell JSON closing the assistant text
('Creating the script now.\n{"cmd": "mkdir -p ..."}') instead of a structured
function_call item. _normalize_codex_response only recognised the Harmony
`to=functions.<name>` leak, so this shape was normalised as plain content with
finish_reason=stop: the JSON was printed, nothing ran, and the turn ended.
Classify it through the same leak seam: a trailing `{"cmd": ...}` object (scalar
siblings only) whose previous line is an action lead-in is a failed tool call, so
the turn is incomplete and the existing Codex continuation re-elicits a real
function_call. A bare or explained `{"cmd": ...}` payload, or one that does not
close the message, stays a legitimate answer (false-positive guard).
Leaked text (both leak shapes) also no longer populates codex_message_items, so
the continuation cannot replay it as a completed assistant message.
Salvaged from #56958 (@YanzhongSu, commit by @minhngoc25a): compact lead-in
regex, one detector for both leak shapes, tests trimmed to two invariants.
Fixes#56920
Replace the direct-queue chat-completions test with one that drives
_spawn_stream_agent with the real _run_agent/_create_agent path and a fake
AIAgent invoking reasoning_callback, asserting the delta.reasoning_content
frame reaches the SSE writer. Unwiring either the _spawn_stream_agent kwarg
or the AIAgent kwarg now fails the test.
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
#116138 taught the bare-`custom` switched arm to keep the session endpoint when the credential
ladder answered with the `OPENROUTER_BASE_URL` mirror. That test was host/URL-only, so it also
fired when the ladder's answer came from an endpoint configured *ahead* of the OpenRouter rung —
`CUSTOM_BASE_URL`, or a trusted `model.base_url` holding the same URL as the mirror (two env vars
aimed at one proxy). In that configuration the switch paired the new model with the previous
provider's host and key again, i.e. the #73680 shape this arm exists to prevent.
The guard now asks the ladder's own question: the mirror counts as a fallback only when
`CUSTOM_BASE_URL` and a trusted `model.base_url` are both absent, so a URL that equals the mirror
because the user configured it there still wins.
Second act of `test_switch_to_bare_custom_ignores_an_openrouter_mirror` pins it: with both env
vars pointed at the same proxy the switch adopts that endpoint; without the precision check the
test fails (`api.anthropic.com` kept instead of the configured proxy).
Some OpenAI-compatible gateways answer an upstream outage with HTTP 403
and a structured error code ``upstream_unavailable`` ("Upstream service
temporarily unavailable. Please retry later."). The 403 handler only
distinguished billing from auth, so this transient outage was classified
``auth``/non-retryable: agent.api_max_retries was ignored, the turn
died on attempt 1 and the credential was treated as refused. Give the
structured transient code precedence over the generic 403 default and
classify it as overloaded (retry with backoff, no credential rotation),
matching how the same condition is handled when it arrives as 503.
Part of #75388 (case 1; the content-policy retry-budget case is a
deliberate design from 0554ef1aa3 and needs a maintainer decision).
Gemini's real wire body is {"error": {"code": 503, "status": "UNAVAILABLE",
"message": ...}}: the symbolic code lives in error.status while error.code
carries the HTTP status. _code_from_payload only read error.code/error.type,
so the numeric code was discarded and the gemini rows in
_PROVIDER_CODE_VERDICTS (keyed on unavailable/deadline_exceeded/internal) were
unreachable; a status-less 503/UNAVAILABLE classified as unknown. A non-string
code now falls back to error.status.
Also pins the anthropic rate_limit_error row: reason rate_limit with
should_rotate_credential and should_fallback both True, which the existing
parametrized case (rotate False) could not host.
A provider error body that carries only a structured code and no HTTP
status (Gemini UNAVAILABLE / DEADLINE_EXCEEDED / INTERNAL, Anthropic
API_ERROR, OpenAI SERVER_ERROR) fell through every classifier stage to
FailoverReason.unknown, so the retry loop treated a provider overload
or timeout as a generic retryable failure and never reached the
overload/timeout recovery hints. Add a per-provider code table consulted
after the shared _ERROR_CODE_VERDICTS map; codes are scoped to their
provider family so a same-named code from another backend stays unknown.
Fixes#70414
Salvages #70425 (@ooiuuii) onto the rule-table classifier layout.
_persist_multimodal_text_parts persisted the text part and then text_summary
under the same tool_use_id, so the second write overwrote the spill file with
the (shorter) summary while the part's <persisted-output> preview advertised
the longer text. Reuse the part's bounded replacement for an oversized summary
instead of persisting twice; the test fixture now uses distinct text vs
text_summary (real browser_exec shape) and asserts the file holds the part text.
_record_persisted_path_for_stub only read string results; a spilled multimodal
envelope now has its persisted path extracted from text_summary / text parts
so a duplicate-result reference stub can point at the spill file.
#95429 acceptance criterion 3: retrying a zero-event oversized request must not
resend the same pathological payload unchanged. run_codex_stream's in-place
transport reconnect re-sent dict(api_kwargs) verbatim with no diagnostic.
When an attempt dies before any stream event was received and the serialized
input exceeds the per-turn budget, inline function_call_output strings over the
per-result threshold (results that escaped commit-time persistence) are spilled
through maybe_persist_tool_result -- bounded preview + recoverable
<persisted-output> reference -- and one warning records
serialized_input_bytes=before -> after with the attempt number. When nothing is
prunable the warning says the payload is being resent unchanged. Only the
retried wire payload changes; the caller's kwargs and history are untouched.
No new config or env vars.
Before: zero-event reconnect resent 760K-char tool output byte-identical, no log.
After: retried input < 1/10 the size, spill file holds the full text, one
"Codex zero-event retry (attempt 1/2): ... serialized_input_bytes=N -> M" line.
A browser_exec call that captured a screenshot returns a multimodal envelope
whose text part carries the full stdout. _finalize_tool_result exempted every
multimodal envelope from maybe_persist_tool_result, and on vision-capable
routes the part list also slips past enforce_turn_budget (len(list) counts
parts, not chars). A 760K-char browser result therefore stayed inline in hot
context and was re-sent on every later request (#95429: 2.47 MB requests,
repeated no-first-byte stalls on retry).
Route each TEXT part (and text_summary) of a multimodal envelope through the
same per-tool persistence threshold; image parts stay untouched (their size is
already governed by the vision embed budget). Normal-sized envelopes are
returned unchanged. Same defect class as PR #95458 (@fangliquanflq), ported
minimally.
Drop the length-independence case: the chat-vs-Responses parity test already
fails on base by the same base64/4 inflation, and the string-output control
covers the negative side.
OpenAI's original overflow wording — "(36865 in the messages, 65536 in the
completion)" and the legacy "(771 in your prompt; 4000 for the completion)" —
is copied by vLLM and llama-cpp-python. It names neither max_tokens nor
"output tokens" and ends with "reduce the length", so both
parse_available_output_tokens_from_error and is_output_cap_error rejected it
and the 400 went to the compressor, which re-sent the same oversized
max_tokens until "cannot compress further".
The first figure is the prompt the server measured, so window - prompt is a
real output budget: return it when >= 1 and route the turn to the max_tokens
clamp. When the measured prompt alone fills the window the parser keeps
returning None and a genuine input overflow still compresses.
Slim port of #90612 (regex and semantics by @Dhruv7201); the three-helper
shape was folded into one budget helper and the tests trimmed to two
invariants. Fixes#90607.
The retryable kwarg only reached log_api_error_attempt from
turn_api_error.handle_api_error, yet both tests called the helper
directly, so dropping the forwarding line kept the suite green.
Replace the retryable=True direct-call test with one that drives the
production entry (classifier patched to a non-retryable 401) and
asserts 'attempt 1/3, not retryable' on the log and status buffer:
red with the kwarg removed, green with it restored.
Also note why _touch_activity keeps the plain counter: it is a
watchdog liveness label, not chat output.
A 401 on a static-key route has no credential to refresh and no pool
entry to rotate to, so it goes straight to the fallback chain — yet the
log said "API call failed (attempt 1/3)" and "attempt 2/3" never came.
Readers took the stuck counter for a retry bug (#73237). The
classifier's verdict now rides the same line on both surfaces (logger
warning and the buffered status trace): "attempt 1/3, not retryable".
Retryable failures keep the plain counter.
Part of #73237 — the policy question (retry an unchanged static
credential once before fallback) is left to the maintainer.
Groq's OpenAI-compatible wire accepts top-level reasoning_effort only as
"none" or "default" (#75089, Defect 1); the custom profile forwarded the
configured graded level ("medium"/"high") and the request still 400'd on
both the main transport and the auxiliary path. The clamp lives in
CustomProfile.build_api_kwargs_extras, keyed on the resolved base_url host,
so both paths share it. The aux bare-key case now drives call_llm with a
fake client instead of the private _build_call_kwargs.
An unconfigured provider name carrying just an explicit base_url keeps the
generic nested reasoning fallback; the fireworks control contract
(test_auxiliary_reasoning_wire_shape[unregistered-gateway]) pins that, and
#75089 only needs keyed providers:/custom_providers entries to route through
CustomProfile.