The Codex Responses API reports quota exhaustion as HTTP 200 with
`response.status == "failed"` and `error.code == "usage_limit_reached"`.
The SDK never raises, so the exception path's credential-pool rotation
(`recover_after_classification` -> `_recover_with_credential_pool`) never
saw it: `retry_invalid_response` retried the same exhausted account
`max_retries` times and then went straight to cross-provider fallback,
leaving a healthy sibling pool entry unused and the dead entry unmarked.
Fix: `turn_recovery.classify_codex_soft_failure` reshapes `response.error`
as an SDK-style error body and runs it through the existing
`classify_api_error` / `extract_api_error_context` (no second pattern
list); `retry_invalid_response` then tries same-provider pool recovery
FIRST for rate_limit / billing / auth verdicts and falls through to the
existing eager fallback + retry path otherwise. Content-policy and other
non-quota failures never rotate. The `ResponseError` repr that leaked into
the retry trace now shows the provider's message.
Live probe (fake Responses SSE server, 2-entry pool, real AIAgent loop):
before -> 3 requests on tok-A, pool untouched, turn failed;
after -> 1 soft failure on tok-A, cred-0 exhausted (usage_limit_reached),
rotated to tok-B, turn completed; content_policy control -> no
rotation on either side.
Fixes#24159
Salvages #24173 (@jmmaloney4) — recovery order and insertion point; the
hand-rolled pattern classifier is replaced by the shared error classifier.
Co-authored-by: Jack Maloney <jmmaloney4@gmail.com>