The rotation test drove pool.mark_exhausted_and_rotate directly, so reverting the
recover_with_credential_pool model_entitlement branch stayed green. It now drives the
production recovery entry point with the classifier verdict and asserts (True, False),
the swap to cred-1 and the model-only bench on cred-0.
A ChatGPT-account entitlement is a plan property, not a quota window: the model bench
now lasts until the explicit reset path (hermes auth reset) clears model_cooldowns,
instead of re-probing the unentitled account hourly on the 429 TTL (#71970).
The exact `The '<model>' model is not supported when using Codex with a
ChatGPT account.` 400 classified as format_error, so a two-entry openai-codex
pool never tried its second account: the turn failed as a malformed request
even though the other account was entitled (#71970). #106475 covered the
single-credential case and explicitly left the multi-entry case to rotation,
but no rotation branch existed for it.
- classify the exact normalized text as FailoverReason.model_entitlement
(rotate + fallback, never retry); arbitrary 400s stay format_error
- recover_with_credential_pool rotates once on that reason; the pool records
it through the existing model_cooldowns path (Anthropic per-model 429), so
only (credential, model) is benched: other models keep using the account,
and `hermes auth` reset clears the marker with everything else
- _mark_entitlement_rejected_model gates on pool.has_available(model=...)
instead of entry count, so once every account rejects the model it falls
back to the #106475 session marker (fallback walk skips it, no oscillation)
Salvages the design of #71973 by @kilhyeonjun on the current classifier
tables and the model-cooldown substrate that landed since.
Co-authored-by: kilhyeonjun <kboxstar@gmail.com>
Follow-up to the salvaged #111787 commit (@KoNit-K), same mechanism as #75578 (@adikpb).
What:
- Move the per-model cooldown logic out of the 2.7k-line credential_pool.py facade into
agent/credential_pool_model_cooldowns.py (mixin + module helpers), keeping only the
select/has_available/next_available_at hooks and the mark_exhausted_and_rotate branch in the facade.
- The model cooldown uses the same TTL policy as a credential-wide 429 (_exhausted_ttl: provider
reset_at wins, a sole credential keeps its 60s bench) instead of a flat 1h, so a single-credential
user is never benched LONGER for the failed model than before.
- Drop the `quota_scope == "account"` check: nothing in the tree produces that key.
- Also keep billing_unverified 429s credential-wide, matching the credential-wide branch.
- _rebind_primary_credential_pool read `rt` that was not in its scope (NameError on every
post-fallback restore); pass primary_model from the caller instead.
- resolve_anthropic_token(model=...) gates only model-aware callers; model-less diagnostics
(usage display, model discovery) keep the key as before.
- _anthropic_token_or_raise names the cooled model instead of claiming no credentials exist.
- Simplify the salvaged call sites (unconditional select(model=)/resolve_anthropic_token(model=)),
widen test stubs that lacked the new kwargs, trim the tests to two invariants on a real temp store.
- Docs: credential-pools.md documents per-model Anthropic 429 cooldowns.
Why: a generic Anthropic 429 is a per-model rate limit; benching the whole credential took every
other Claude model offline while the env/borrowed token path handed the same benched key straight
back (#111769, #61451).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: adikpb <67222969+adikpb@users.noreply.github.com>