Ports #119027::finalize_continuation_partial: when a stream drop entered the
continuation path and the next request exhausted retries before a token, the
_length_continuation_fragment/_nudge rows were persisted as-is, so resume
replayed a dangling synthetic user nudge. Collapse them into one assistant row
before persistence and feed the collapsed text to the #119081 partial-retention
path (the fragment rows are gone by then).
Co-authored-by: fangliquan <fangliquan@qq.com>
abort_turn_on_interrupt closes an open tool sequence with the caller's
specific interrupt text and then persists. When a Stop lands during
empty-response recovery, the synthetic assistant+nudge pair still sits
after the executed tool result, so the close sees no exposed tool tail
and the generic close in _persist_session wins with "Operation
interrupted.". Strip only the request-local scaffold first, so the exit
owner closes the tail with its own reason.
(taken from 413ad9647b72b27e39c49d8e0daa05bac5411ec1 in #120883,
agent/turn_recovery.py hunk only)
An organization with no fast allocation for a model gets a 429 whose
anthropic-fast-*-tokens-limit header is 0, with no retry-after. Hermes
treated it as a rate limit: backoff, credential rotation that benched a
key that works at standard speed, then provider fallback, so /fast
failed every turn.
The pre-classification recovery now stops sending `speed` to that model
for the rest of the session and retries once. A 429 with a real fast
limit still takes the retry-after path.
The image-rejection recovery no longer switches the session to
text-only or strips history; it records the (provider, model) and
build_api_request strips images from that model's requests only. Three
places still described the old behaviour:
- recover_before_classification docstring and the adjacent comment said
"switch session to text-only" / "mark session vision-unsupported".
- TestStripImagesDropsStaleApiContent's rationale claimed the strip runs
on persistent history and that leaving api_content would replay
rejected images every turn because the recovery gates on
_image_rejecting_models — false on both counts: current callers pass
per-call clones. Reworded as the generic helper contract (a rewritten
persisted row must drop its sidecar) with the no-op note.
- _strip_images_from_messages docstring, same fix.
Comments and docstrings only; no code change.
Two branches in turn_recovery.py carried the same isinstance +
_strip_images_from_messages guard and the same "Provider rejected a
corrupted image" notice: the new pre-classification corrupt branch and
the FailoverReason.image_corrupt branch. Both now call one module
helper, _strip_request_images_and_retry(agent, api_messages) -> bool,
so the strip-this-attempt-only policy lives in one place.
_IMAGE_REJECTION_PHRASES was rebound as unsupported + corrupt, which
made _looks_like_image_content_rejection silently cover corrupt payloads
and forced the recovery to test the corrupt list a second time to undo
that. The name stays (external references) but is now the
unsupported-only tuple; _IMAGE_CORRUPT_PHRASES is disjoint, and the
recovery asks the two questions explicitly:
`_corrupt or (model not yet rejected and unsupported)`.
Tests: the phrase-isolation matcher asks the same disjoint question the
recovery does; the image_corrupt source-contract check looks for the
helper call instead of the inlined strip.
recover_before_classification's capability branch stripped api_messages
itself right after recording the (provider, model) in
_image_rejecting_models. That strip is dead work: the verdict is
`continue`, which re-enters build_api_request with the SAME api_messages
object, and strip_images_for_rejecting_model runs there — before
_build_api_kwargs on every attempt — and strips them because the key is
now in the set. Keeping both meant two places had to agree on the send
policy. The corrupt-image branch keeps its local strip on purpose: it
deliberately does not record the model, so the send path would not.
test_canonical_history_keeps_its_images now asserts the text-only wire
copy through strip_images_for_rejecting_model (the real send path) and
checks build_api_request still calls it ahead of _build_api_kwargs;
commenting that call out turns the test red.
init_agent seeds `_image_rejecting_models = set()` at the same site as
`_force_ascii_payload`, which the adjacent sanitize_outbound_kwargs already
reads as a plain attribute. The getattr/isinstance guards and the lazy
re-create in recover_before_classification implied the attribute could be
missing or mistyped on a real agent; it cannot, and defensive fallbacks for
impossible states hide wiring bugs instead of surfacing them. Every test
fixture that reaches these paths seeds the attribute already.
message_sanitization.image_model_key duplicated vision_message_prep's
_provider_model_key — the (provider, model) key the same mixin already uses
for its per-model vision bookkeeping. Two keying functions for the same
concept drift: one normalised the provider (.strip().lower()), the other did
not, so a provider spelled "OpenAI" in one place and "openai" in another
would have been tracked as two models. Key `_image_rejecting_models` on the
existing helper and delete the duplicate (and its __all__ entry).
No import cycle: vision_message_prep imports only lazy_forward,
tool_dispatch_helpers and utils, none of which import message_sanitization
or turn_recovery.
recover_before_classification matches _IMAGE_REJECTION_PHRASES, which
mixes two kinds of body: capability rejections ("does not support
images", "only text content type is supported", ...) and bad-payload
rejections. After the per-model tracking landed, BOTH added the
(provider, model) to agent._image_rejecting_models, so one truncated
screenshot rejected with "failed to decode image" made every later
request to that model text-only for the rest of the session, even
though the model can see fine.
Split the bad-payload phrases into _IMAGE_CORRUPT_PHRASES:
- "image data you provided does not represent a valid image"
(ChatGPT-account Codex backend)
- "failed to decode image" (Kimi / Moonshot and other
OpenAI-compatible providers)
_IMAGE_REJECTION_PHRASES stays the union so the turn still recovers on
them. For a corrupt match the recovery now strips the current attempt
only and retries iff something was stripped (mirroring the existing
image_corrupt branch), and leaves the model unmarked; only a capability
rejection records the model.
One test: a "failed to decode image" body strips the wire copy, keeps
history, and leaves _image_rejecting_models empty so the next request
to the same model carries its images.
Image rejections are now tracked per (provider, model) in
agent._image_rejecting_models, and recover_before_classification gates
on that set. Nothing reads agent._vision_supported any more, so the
write in turn_recovery and the per-turn reset entry in turn_context are
dead state. Remove both so the next reader doesn't assume a turn-global
vision gate still exists; update the two tests that asserted or seeded
the attribute.
The first head stored a single rejecting (provider, model) and kept the
turn-global `_vision_supported` as the recovery guard. In a fallback
chain that fails: model A rejects images and retries text-only, a later
error activates model B, the restart rebuilds api_messages from history
so B receives the images, and when B rejects them too the branch is
skipped because `_vision_supported` is already False — the request
falls through to generic error handling. Recording B also overwrote A,
so A was no longer treated as text-only on later turns.
`_image_rejecting_models` is now a set of every rejecting model, and it
is also the guard: each model's first rejection runs the recovery and a
repeat rejection from the same model still falls through, so the retry
cannot loop. image_model_key() names the key in one place.
Adds a test for the two-model sequence (fails on the previous head) and
one pinning that a repeat rejection from the same model does not retry.
Thanks to @ehz0ah for the review.
(cherry picked from commit 225fd76ccf90cd3ec509a9327f725ac02d996f24)
When a provider 4xx's on image content, recover_before_classification
ran _strip_images_from_messages on the canonical `messages` list and
reset _db_flush_scan_prefix. Since #117569 that function pops
_db_persisted on every rewritten dict, so the next flush rewrote those
rows: every image in the session — and every image-only message, which
the stripper deletes outright — was removed from state.db for good.
The rejection describes what the CURRENT model accepts, not what the
conversation holds. An automatic fallback to a text-only provider, or a
single /model switch, was enough to erase images the user had sent to a
vision model, and switching back found them gone. It is the same failure
as the ASCII strip in #117802, on the image path; neither open fix for
that issue touches this branch.
Keep the repair on the send path, where the per-call copy already lives
(_clone_message_for_send exists so send-path rewrites never reach the
persisted transcript, #80498):
- record the rejecting (provider, model) on the agent, a session-scoped
flag initialised beside _force_ascii_payload;
- strip the in-flight api_messages copy for the immediate retry;
- build_api_request calls strip_images_for_rejecting_model() on each
attempt's api_messages BEFORE provider conversion. The stripper knows
Hermes's own part types; a converted payload would slip past it
(Bedrock Converse image blocks carry no `type`). Keyed on the model,
so one that accepts images gets them again.
_strip_images_from_messages itself is unchanged, so its role-alternation
and sidecar guarantees still hold on the wire copy. The notice no longer
claims "text-only mode for this session" (_vision_supported resets every
turn) or that images were stripped from history.
(cherry picked from commit cbdb184c27f915ab138b2087f878aed7fcc7c76b)
Both callers immediately reduced the (headers_sanitized, credential_sanitized)
tuple with `or`; nobody distinguished the two. Returning one
`transport_repaired` bool removes the unpacking at both call sites and the
redundant flag bookkeeping inside the helper without changing behaviour.
The ASCII branch of _recover_unicode_encode_error popped/stripped/restored
`tools` on the failed attempt's api_kwargs and tracked `_tools_sanitized`,
but that dict is discarded: build_api_request rebuilds api_kwargs from
agent.tools on every retry iteration and sanitize_outbound_kwargs already
strips the whole payload once `_force_ascii_payload` is set. Only
api_messages (reused across retries) and the local active_system_prompt
still need stripping here.
Also removes the `_force_ascii_payload = False` assignment in the UTF-8
branch (the flag is initialised False in agent_init and only ever set True
in this function), the now-unused `_sanitize_tools_non_ascii` import and
the dead `import copy`. The test drops its assertion on the recovery
dict's tools identity; the chokepoint assertion (agent.tools byte-stable
after sanitize_outbound_kwargs) is kept.
_recover_unicode_encode_error carried two copies of the same block that
strips non-ASCII from _client_kwargs["default_headers"] and the API key
(agent.api_key, _client_kwargs["api_key"], client.api_key). The UTF-8
runtime branch was a silent copy; the ASCII runtime branch also printed
the "bad copy-paste?" hint. Two copies drift, and a user whose key was
repaired under a UTF-8 runtime got no hint about why auth might now fail.
Extract one module-level _repair_transport_credentials(agent) returning
(headers_sanitized, credential_sanitized) and call it from both branches.
The hint is emitted whenever the key changed, in both runtimes; the retry
decisions in each branch are unchanged. Net -7 LOC.
api_kwargs["tools"] is rebuilt from agent.tools on every attempt
(_build_api_kwargs_for_mode: tools_for_api = agent.tools; transports set
api_kwargs["tools"] = tools without copying), so the list usually aliases
the canonical tool schemas. The ASCII retry path then runs
sanitize_outbound_kwargs with _force_ascii_payload set, and its in-place
strip rewrote agent.tools for the rest of the session.
Move the guard to the chokepoint: when the flag is set and tools IS
agent.tools, deepcopy before stripping. The deepcopy the contributor pick
added inside _recover_unicode_encode_error only protected the failed
request's kwargs, which are discarded before the retry; drop it and have
recovery skip an aliased tools list entirely (request-local lists are
still stripped for the diagnostic message).
The kept test now exercises the chokepoint directly: with the flag set on
kwargs whose tools aliases agent.tools, agent.tools must be byte-stable
afterwards. Verified red against the previous chokepoint.
Under a UTF-8 runtime, _recover_unicode_encode_error only repairs the
api_key and _client_kwargs['default_headers']. When neither changed, the
retry is byte-identical to the failed request and cannot succeed; it just
consumed both _unicode_sanitization_passes before the error finally
surfaced. Return False in that case so recover_before_classification
falls through to the normal classification/error path immediately.
The existing api_key + default_headers sanitization is unchanged; the
"retrying unchanged request content" message is removed because that
branch no longer retries. The PR's UTF-8 invariant test keeps asserting
the request copy is not stripped, and now asserts the False return.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The _db_persisted marker asserts that a message dict's persisted row is
durable as written; any in-place mutation must pop it or the flush scan
identity-skips the dict and state.db keeps the stale row forever. Six
mutation sites violated the contract:
- micro_compaction._merge_adjacent_user_turns rewrote content on a
carried-forward dict after superseding a stale micro marker. On the
archive-failure path nothing re-stamps, so the merged text never
reached state.db.
- repair_message_sequence passes mutated stamped survivors in place:
_merge_assistant_into (tool_calls union, content join,
reasoning_content carry), _prune_unanswered_tool_calls (tool_calls
rewrite), _merge_consecutive_users (content join).
- sanitize_tool_call_arguments rewrote corrupted/blank
function.arguments and prepended the corruption marker onto stamped
resumed rows, leaving the corrupt bytes durable and self-perpetuating
across resumes.
- _sanitize_messages (surrogate and non-ASCII recovery) and
_strip_images_from_messages (image-rejection recovery) rewrote live
dicts on the recovery path.
Each site now pops the marker when a persisted field actually changes,
and the agent-aware callers (repair_message_sequence_with_cursor,
turn_iteration_prep, turn_recovery) invalidate the bounded flush-scan
prefix so repaired rows are rewritten on the next flush.
Trim the salvage of #115706 to the existing seams:
- The structured code joins _BILLING_ERROR_CODES and _status_404 consults that table
first, exactly like _status_429 (the status handler always returns, so _by_error_code
never saw the code). Drops the duplicate _CREDIT_EXHAUSTION_404_CODES set, the
message-substring scan, and the credit_exhaustion_code verdict marker.
- The log moves from two call sites in the turn loop into try_activate_fallback, the
one chokepoint every fallback switch passes through. A billing switch is a WARNING
naming the profile (resolved from the scoped home, so under multiplex it is the
failing profile, not the launch profile), both models, and the `hermes [-p X] model`
remedy; every other reason keeps the INFO line.
- Tests trimmed to two invariants (classifier row + control body; warning is
profile-scoped A->B and non-billing stays INFO), both red on origin/main.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
A 404 carrying insufficient_credits_for_paid_model (paid model ungated
by credits) classified as unknown: retryable with no fallback, burning
retries and never switching models. Treat it like 429-exhaustion --
billing with rotate+fallback -- and log an actionable ERROR naming the
credits and the fallback target on activation.
Fixes#115702
(cherry picked from commit a1f7f9996bb82230c945340dcb279ffa923e90a1)
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
- turn_recovery: the upstream_blocked hint no longer advertises model.default_headers,
which _apply_user_default_headers skips on anthropic_messages/bedrock_converse.
- cron: an upstream_blocked run says to set a User-Agent via extra_headers or pin
another provider, instead of 'run it again' — a retry never heals a firewall block.
- agent_runtime_helpers: in anthropic_messages the request dump read agent.client
(None) and printed 'Authorization: Bearer None' with a /chat/completions URL; it now
reads the live _anthropic_client key and uses /messages.
Two gaps for custom providers behind a WAF/CDN:
- `build_anthropic_client` never consulted `custom_providers[].extra_headers`,
so a relay in `anthropic_messages` mode that rejects the SDK User-Agent kept
403ing even with `extra_headers: {User-Agent: ...}` configured, while the
OpenAI-wire clients already applied it. The lookup now lives in
`_new_sdk_client`, the one constructor every builder path goes through
(init, /model switch, rebuild, auxiliary), keyed by the caller's raw route
because entries are keyed by the `/v1` form the normalizer strips.
Salvaged direction of #46002 (@wait4xx). Fixes#24293, #9721.
- `_status_403` classified every non-billing 403 as `auth`, so a WAF's plain
"Your request was blocked." or a Cloudflare browser challenge printed "Your
API key was rejected" and could rotate a healthy credential. A 403 carrying
established block/challenge markers is now `upstream_blocked`: no rotation,
no retry, fallback allowed, WAF/User-Agent guidance on every surface (CLI
loop, chat copy, cli chat error copy, TUI gateway + Ink TUI copy). Generic
403 and all 401 keep the auth verdict. Salvaged direction of #70567
(@ooiuuii) and #53114 (@AgenticSpark). Fixes#53099, #70566.
_summarize_api_error reduces the usage_limit_reached body to 'HTTP 429: The
usage limit has been reached', so no text-based parser could ever reach the
8.6h reset from the in-loop failure path. _status_429 now stamps
error_context['reset_at'] from the body's reset fields, Retry-After or the
message grammar (same table the credential pool uses), and
max_retries_exhausted_result hands the remaining seconds to exhausted_copy,
which says 'its usage limit resets in ~9h. Send /retry after that' once the
window is >= 2 minutes. A throttle with no window keeps the short-wait copy.
Proven with a real openai.RateLimitError carrying resets_in_seconds=30995
driven through the production result builder. (#89401 atom 1)
A 401 on a static-key route has no credential to refresh and no pool
entry to rotate to, so it goes straight to the fallback chain — yet the
log said "API call failed (attempt 1/3)" and "attempt 2/3" never came.
Readers took the stuck counter for a retry bug (#73237). The
classifier's verdict now rides the same line on both surfaces (logger
warning and the buffered status trace): "attempt 1/3, not retryable".
Retryable failures keep the plain counter.
Part of #73237 — the policy question (retry an unchanged static
credential once before fallback) is left to the maintainer.
With a multi-entry openai-codex pool, a 401 `token_expired` caused by a stale
replayed `encrypted_content` blob went pool-first: each healthy entry was
force-refreshed (single-use refresh token) or benched STATUS_EXHAUSTED before
the strip in `_recover_format_errors` finally ran, so pooled users lost every
account for the bench window over a session-state problem. The single-credential
path likewise burned a forced OAuth refresh on a bearer that was fine.
`recover_after_classification` now takes the Codex stale-reasoning strip first
(after the Nous welcome-tier repair) when the 401 carries `token_expired` and
the transcript still holds `codex_reasoning_items`; the same one-shot latch
bounds it, so a real expiry pays one extra round-trip and then takes the pool /
refresh path exactly as before. The strip body moved into
`_recover_stale_codex_reasoning`, shared with the 400 `invalid_encrypted_content`
branch. Docs: one line in the OpenAI Codex path section describing the
self-heal.
Tests: pooled control with a real two-entry CredentialPool (red on the previous
ordering: pool rotated, entry benched; green now: strip runs, nobody benched);
the single-credential test now asserts no refresh is burned before the strip.
A resumed openai-codex session failed every prompt with HTTP 401
"Provided authentication token is expired" while a fresh session on the
same bearer worked: the Codex backend rejects a stale replayed
`encrypted_content` reasoning blob with the auth signature instead of
the 400 `invalid_encrypted_content` Hermes already recovers from, so the
error went to OAuth recovery (which had nothing new to adopt) and the
turn died on "sign in again".
`_recover_format_errors` now also takes a 401 `token_expired` once the
one-shot Codex/xAI OAuth refresh has run and changed nothing, and only
while the transcript still carries `codex_reasoning_items`: it reuses
the existing `AIAgent._disable_codex_reasoning_replay` strip + retry and
the same one-shot latch. A 401 without cached reasoning (a real expiry)
or without the token_expired code stays on the credential path; the
classifier is untouched (401 remains FailoverReason.auth).
Slim redo of #88541 by @StanleyStetson (same predicate, at the recovery
chain's current home in agent/turn_recovery.py instead of the loop).
Fixes#88510
Co-authored-by: StanleyStetson <StanleyStetson@users.noreply.github.com>
Drops _CODEX_APP_SERVER_FALLBACK_REASONS (same three members as _RATE_LIMIT_REASONS
defined above it) and hoists classify_api_error to the existing top-level
error_classifier import. Documents under Known limitations that codex_app_server
honours fallback_providers on quota / rate-limit failures only; auth, timeout and
model-not-found failures do not fail over on this runtime.
Re-applies kingrubic's #71642 logic on the _LoopState loop shape: classify the
codex app-server result["error"] text after the turn, guard on
_has_pending_fallback(), fail over on billing / rate_limit / upstream_rate_limit,
carry the failed API call in api_call_count and resync the failover system
message before retrying the same user turn on the generic loop.
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
Temporarily removes the hunks that 08ad795b9db carried over line-for-line from
kingrubic's PR #71642 (codex_result error + _has_pending_fallback guard,
classify_api_error(RuntimeError(str(err)), ...), the eligible-reason frozenset,
api_call_count accounting, _sync_failover_system_message) so the next commit can
re-apply them under the contributor's authorship. Tree is intentionally
non-functional between this commit and the next.
The codex_app_server dispatch returned the turn result unconditionally, so a
billing/usage-limit/rate-limit error carried as result["error"] text never reached
the classify_api_error -> _try_activate_fallback chain every other runtime uses.
Classify that text after the codex turn; on a fallback-eligible verdict activate
the configured fallback and continue the same user turn on the generic loop,
keeping the projected codex rows and the failed API call in the accounting.
Re-port of PR #71642 (@kingrubic) onto the _LoopState loop shape; the invariant
test drives the real activate -> retry path against a local fake OpenAI server.
Fixes#71633
Endpoints that understand the reasoning field but refuse the OFF ("Reasoning is
mandatory for this endpoint and cannot be disabled" -- the Nous Portal on
gpt-6-astra) got two different wrong answers:
* auxiliary lanes (title generation sends reasoning_config={"enabled": False})
400'd outright -- the #112781 strip rung only matched "unsupported/unknown
field" wordings, so every session on such a route stayed untitled;
* the main loop dropped the disable and let the route pick its default effort,
the opposite of what a thinking-off user asked for.
Both now step the effort up to the lowest level every reasoning wire accepts
(low) instead. The aux ladder gets a rung ordered before the strip
(agent/auxiliary_reasoning_floor.py: lifts reasoning_effort / extra_body.reasoning
/ _reasoning_config, memoises the (route, model) so the next thinking-off aux call
starts at the floor without the guaranteed 400); the main loop's reasoning_mandatory
recovery records whether the route said mandatory (floor) or unknown-field (drop,
unchanged). error_classifier.is_reasoning_required_rejection separates the two.
Live on the portal (proxy wire capture): title none->400->low->200 titled;
main loop none->400->low->200; second aux call on the route sends low up front.
A "HTTP 429: The usage limit has been reached" turn offered only Retry and
never said when a retry would work, so users guessed or babysat the app
(#98852). The provider already tells us: Retry-After / resets_at /
retry_after are parsed into the turn's error context (extract_api_error_context)
and honoured by the backoff, but the datum died there.
- agent/turn_recovery.py::_stamp_limit_reset: both terminal paths
(max_retries_exhausted_result, nonretryable_client_error_result) stamp
failure_resets_at (epoch s) on the failed result and append one plain line
("Limit resets at 14:05 (in 1h 00m).") to final_response, which every text
surface (CLI, Ink TUI, messaging gateway) renders.
- agent/error_surface.py: result path forwards failure_resets_at as
surface.resets_at; the exception path derives it from the same context.
- tui_gateway/contracts/events.py::ErrorSurface.resets_at + regenerated
apps/shared gateway-contract outputs.
- apps/desktop lib/error-surface.ts: parse resets_at -> resetsAt,
formatLimitReset("HH:mm (in 1h 05m)", null once passed), diagnostics line;
the error card renders "Limit resets at …" next to Retry (i18n copy in every
full locale).
- Docs: website/docs/user-guide/desktop.md error-card section.
Informational only: no scheduled or automatic retry is added — firing a turn
unattended on a subscription is the maintainer's call (#98872, #103048).
The enriched "Rate limited. Resets in ~13m. Waiting ..." text only lived in
the buffered status, which replays solely when every retry fails. The line a
user actually watches during the wait — _emit_diagnostic_wait on CLI, the
messaging gateway, the Ink TUI and Desktop — was still the anonymous
"waiting on provider — retrying in 60s". It now reads "rate limited — resets
in ~13m, retrying in 60s (attempt 2/3)" when the reset is known and keeps the
old wording otherwise.
extract_api_error_context also learns the per-bucket reset headers plain
OpenAI 429s (x-ratelimit-reset-requests/-tokens, "6m0s"/"1.5s"/"20ms") and
Anthropic 429s (anthropic-ratelimit-*-reset, ISO-8601) carry when no
Retry-After is present, at the lowest priority after Retry-After and
x-ratelimit-reset; past timestamps are ignored.