Two holes in the forced retry. With an explicit live-agent api_key the
retry dropped it and re-resolved from the singleton/pool, so a 401 on pool
entry B rendered account A's usage — the cross-account leak the surrounding
comments say the tiering prevents. The retry now force-refreshes the pool
entry that issued the key (or the singleton when it IS the singleton token)
and fails open when neither matches.
In a pool-only setup (empty singleton) resolve_codex_runtime_credentials
returned the pool token before consulting force_refresh, so the retry resent
the identical revoked bearer. The pool-only branch now rotates the pool
entry through try_refresh_matching when force_refresh is set.
Also: when the forced refresh itself raises inside redeem_codex_reset_credit,
surface the original 401 (re-login hint) instead of the refresh error.
Review finding: explicit-key force refresh re-resolved another account; pool-only tier ignored force_refresh.
Review follow-ups on the refresh transaction:
- `_provider_state_transaction` takes a `timeout_seconds` applied to BOTH the
active and the root lock. The refresh passes max(default, refresh timeout
+ 5 s); before, only the profile lock used that budget and root's lock kept
the 15 s default, so the waiting profile raised TimeoutError instead of
adopting whenever the peer's POST ran long. The regression test now holds
the endpoint past the lock floor and fails without the passthrough.
- Peer adoption requires a stored access token as well as a rotated refresh
token; an incomplete stored pair falls through to the refresh.
- `_save_codex_tokens` keeps its body: the per-path lock is reentrant, so the
refresh calls it inside the open transaction (as the CLI-recovery path
already did) instead of a split-out helper.
Follow-up to #110024 (ehz0ah's review thread). `_refresh_codex_auth_tokens` POSTed the
single-use refresh token to OpenAI first and only then entered
`_provider_state_transaction("openai-codex")` for the write-back, so root's lock covered the
save alone. `resolve_codex_runtime_credentials` holds only the caller's own profile lock, so
two profiles borrowing the same ROOT grant could both submit `old-rt`; last root save won and
OpenAI answered `refresh_token_reused` / revoked the family — the failure #87503 exists to
prevent.
The transaction now spans re-read -> endpoint refresh -> write-back:
- enter `_provider_state_transaction` first; the yielded state is root's, re-read under root's
lock. If its refresh token already differs from the one we were about to submit, a peer
rotated it: adopt the stored pair and return without touching the endpoint.
- otherwise POST and write back through `_store_codex_tokens_in`, the body of
`_save_codex_tokens` split out so it can run inside an already-open transaction.
`_save_codex_tokens` keeps its signature for the login/import/CLI-recovery callers.
Holding the advisory flock across the network call is safe here and already the established
shape: `resolve_codex_runtime_credentials` holds the active-store lock across the same POST,
and every waiter's timeout is `max(AUTH_LOCK_TIMEOUT_SECONDS, refresh_timeout + 5)`, i.e. it
outlives one full endpoint timeout. `_load_auth_store` readers never take the lock, so
readers are not blocked; `_file_lock` is reentrant per thread per path, so the nested
transaction inside the caller's lock and the CLI-recovery save inside the transaction both
re-enter cleanly. A release-POST-retake variant would reopen the window it is meant to close.
Test: two refreshers with the same stale pre-read pair against a rotate-once endpoint that
rejects any replay — the endpoint sees `old-rt` exactly once, both callers end with the
rotated pair, root holds it, the profile store stays unshadowed. Red on origin/main
(`refresh_token_reused` surfaces for the second caller).
Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.
Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).
Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes#87503
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.
Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
`_read_codex_pool_entries` had become a pass-through whose only residual
behaviour was taking the active-store lock around a read every other
`read_credential_pool` caller does unlocked (and which never covered the
root file it fell back to). Both consumers now call the helper directly.
`clear_codex_pool_quota_cooldowns` decided "profile owns rows?" with an
unlocked pre-read, then re-read the same file under the lock — a TOCTOU
against a concurrent `hermes auth add` and a wasted parse. It now tries the
active store under its lock and falls back to the borrowed root only when
that store has no codex rows.
With reads now inheriting the global-root pool, a profile hitting a stale
root cooldown probes quota, sees it restored, and calls
clear_codex_pool_quota_cooldowns() — which only ever edited the (empty)
profile store, so the next resolve raised quota_exhausted again forever.
Pick the store the same way agent/credential_pool.py persists borrowed
rows (_profile_owns_pool_provider / _borrowed_single_use_pool_root) and
lock/save against that path. The fallback test now also binds
profile-wins precedence; one new test pins the root write.
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
__cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
SSLEOFError whose text httpx did not repeat still gets the hint. The
hint now names the TLS 1.2 diagnostic and links the providers docs
note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
classic-groups snippet (EN + existing zh-Hans copy).
Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.
- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
(device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
workaround hint; KeyboardInterrupt handling is unchanged
(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.