A chat-completions stream that ends with no finish_reason while a tool call's
arguments are cut at a value boundary leaves the assembler holding a prefix
that _repair_tool_call_arguments closes into valid JSON: '"timeout": 600' cut
after its first digit becomes timeout=6, a todo list cut after its second item
replaces the whole list with two. The repaired call was no longer flagged as
truncated, so _finish_chat_stream stamped the turn "stop" and dispatched it
with every key and digit that never arrived missing. Only a cut inside a
string stayed on the partial-stream stub path.
The assembler now repairs only calls the provider finished (any
finish_reason). A dropped stream's malformed arguments are flagged and take
the existing stub retry, like the empty-argument (#80498) and repetition
guards beside it. Repair of completed responses (#115061, GLM via Ollama) is
unchanged.
Ported from the second commit of #105017 onto
_StreamingCall._assemble_tool_calls. Its finish_reason="length" arm is left
out: that path already retries the call without executing it.
(cherry picked from commit 0895ffe009613dc6b6087c99d2cf599211ec87dc)
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
An auth/quota error whose text also says "overloaded" went through the
overload branch: it bumped the durable sustained-overload streak even though
it always aborts as summary_auth_failure (#29559). After the credential was
fixed, the next real 503 then inherited that streak and committed the lossy
fallback immediately, skipping the one-abort grace (#115906).
Compute the access/quota classification once and keep such errors out of the
overload branch. With that, the overload-escalation path never runs for an
auth error, so the keep_auth escape hatch (and its hard-coded flag-name
compare) in _clear_terminal_summary_failures is dead and goes away.
The existing [403_overloaded] case now also asserts the streak stays 0.
Co-authored-by: Yuan Li <dskwelmcy@163.com>
The two-line clear-all loop over _TERMINAL_SUMMARY_FAILURES had three
copies (init, summary success, overload escalation), and the escalation
copy now needs to spare the auth/quota flag. Extract
_clear_terminal_summary_failures(keep_auth=...) so the three sites share
one definition and the #29559 exception lives in one place.
_reset_consecutive_overload_aborts skipped the DB write unless the
in-memory count was armed (or settle=True). The in-memory count is not
authoritative when two agents share a session: agent A can bump the row
to 2 while agent B still holds 0, so B's summary success or runtime
switch left A's streak alive and the next overload escalated early.
Drop the settle parameter and always persist; one helper, one contract.
The escalation clear-all of terminal summary flags also cleared the
auth/quota flag the SAME error had just set. A 403/402 or quota-worded
error that also says "overloaded"/"at capacity" classifies as both, so on
strike 3 it committed the lossy fallback instead of aborting - breaking
the #29559 invariant that access/quota failures always preserve the
session. Keep the auth flag when the current error is access/quota; a
stale auth flag from an earlier failure is still cleared.
The kept escalation test is parametrized with a 403 "provider overloaded"
case that must abort all three times with summary_auth_failure.
The why (retry can win vs. deferred compression_exhausted wipe, and why the
streak is durable) was spelled out in full at the constant, in
_on_summary_failure, in record_completed_compaction and in the durable
getter docstring. Keep it once on _CONSECUTIVE_OVERLOAD_ABORT_ESCALATION
and cut the other copies to one line, matching the sibling fallback-streak
getter. Comment-only.
The zero-and-persist pair was pasted at the summary-success, runtime-switch
and completed-boundary sites, and every healthy compaction wrote
compression_overload_streak=0 twice even when it was already 0.
_reset_consecutive_overload_aborts() skips the UPDATE when the in-memory
streak is 0; record_completed_compaction passes settle=True and still
writes unconditionally because another agent on the same session may have
bumped the durable row since this object last read it.
On a compressor with a distinct summary_model, the first overload takes the
aux->main one-shot retry, which pre-sets telemetry failure_class to
aux_model_fallback; the `telemetry.get(...) or` chain then hid the
escalation and operators never saw the documented
summary_overload_degraded class. The degraded flag now wins. It is always
initialised (__init__ and _begin_compress_attempt, the only path into
_fallback_summary_for_window), so the defensive getattr goes too; no test
builds a compressor via __new__ and reaches this method without compress().
The kept fresh-bind test now asserts the label instead of excusing it.
Only a successful summary clears the network/empty/truncated/auth failure
flags, and _abort_on_summary_failure aborts on ANY of them. One earlier
streaming_closed failure on a long-lived (gateway-cached) compressor
therefore kept every later sustained-overload attempt aborting past the
3-strike escalation, so the #123167 compression_exhausted wipe persisted.
When the escalation fires, the latest failure class decides: clear the
other terminal flags so compress() commits the degraded fallback.
The durable overload budget was a read-modify-write (memory += 1, then
set the column), so two agents compressing the same session concurrently
could each write N+1 and lose a strike, delaying the fallback escalation.
Bump the column with one UPDATE ... RETURNING inside _execute_write and
take the returned row value as authoritative; memory-only counting is
kept for unbound compressors. No try/except or legacy-setter shim: the
existing _durable_read plumbing already degrades on unsupported DBs.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Review P1 on #123186 (ehz0ah): the N=3 sustained-overload budget was
object-local. The gateway binds a fresh compressor to the same session on
every turn / cache eviction, and restart / API-server requests construct
one too, so each fresh instance restarted the budget at zero and a
sustained summary-provider outage walked the session back into
compression_exhausted and auto-reset — the exact wipe the PR exists to
prevent. Reproduced by the reviewer with three fresh compressors bound to
one session: every attempt ended at counter=1.
Persist the streak as sessions.compression_overload_streak through the
same durable channel the fallback streak and recovery deadline already
use (#100185): SessionCompressionMixin get/set pair over the sessions
row, loaded in the compressor's durable-load block, written on every
change (overload abort, successful summary, committed boundary, runtime
model switch).
- Compression rotation carries the streak to the child row at the
boundary, same as the fallback streak: the parent's value is read
before the bind, re-applied, and persisted onto the fresh child row.
- A completed compaction boundary settles the budget to 0 — including
the committed degraded fallback, so a recovered provider regains a
full budget instead of the session staying degraded forever.
- Schema event appended to SCHEMA_HISTORY["sessions"] (seq 28, after
compression_recovery_deadline) so salvage replay maps the new column.
- Fresh-instance regression: two aborts on one compressor, then a fresh
compressor bound to the same session inherits streak=2 and the third
session-wide attempt commits the fallback; plus a rotation carry-over
test. Both fail red against the memory-only implementation.
(cherry picked from commit 3742fee8d03d7077d23b0dac3a3c79fe48e102ac)
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.
After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.
Fixes#123167
(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
When the heal recreated the session row but its single retry write then
failed (database locked, turn lease, disk), the replay was carried only by
the per-call _replay_history argument. The next flush found a live row, hit
no FK error, ran no heal, and stamped the history prefix durable without
writing it: the durable transcript silently lost everything but the tail
(scratch probe: rows ['a2'] instead of ['one','a1','two','a2']).
Set _session_row_replay_pending in the heal branch right after the markers
are stripped, so any exit after a heal (recreate failed, retry failed)
leaves the replay pending until a write succeeds. That makes the separate
recreate-failed assignment and the _replay_history parameter redundant;
the retry is a plain _adoption_budget=0 call. The fails-closed test now
covers FK -> heal -> locked retry -> full replay on the next flush.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
When the session row is gone and the heal cannot recreate it, the flush
fails closed with the flush markers already stripped. The next flush then
recreates the row up front via _ensure_db_session, so no FK error fires,
no heal runs, and _db_flush_collect stamps the history prefix as durable
without writing it. The recreated row ends up holding only the new tail
(probe: rows ['a2'] instead of ['one', 'a1', 'two', 'a2']).
Remember the failed heal per session id and have the next flush for that
session replay the history prefix; clear it once a flush succeeds.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
The FK classifier behind session_row_missing also matches the sessions
table's own parent_session_id / system_prompt_hash FKs, and session
create is an upsert. Healing on the classification alone could replay
the whole transcript into a still-live row and duplicate it. Heal only
when get_session(session_id) confirms the row is gone.
The delegate-child branch's get_session(parent_id) was unguarded: a
store error there escaped the flush's error handler. Route both lookups
through one guarded helper (as the compression-tip helper does); a failed
lookup is not proof of deletion, so it falls through to the existing
"will retry next flush" return.
While here: _db_flush_failed now returns which retry to take
("adopted" / "healed" / None) instead of the caller string-comparing
_last_persistence_error_cause; its docstring covers the heal branch;
messages is required (the only caller always passes it); the marker
reset reuses _strip_persistence_markers. Tests drop the change-detector
assertion on the lingering cause after a successful heal and the Scenario
B docstring now says fail-closed, which is what it asserts.
The session-row heal retried with conversation_history=None so the
history prefix would be replayed onto the recreated row. On a muted
notification-reply turn that also made _db_flush_collect treat every
history row as new: each replayed row was written display_kind=hidden
and the flag was stamped onto the live in-memory history dicts, hiding
the user's whole past transcript from pollers from then on.
Keep the history set on the heal retry and pass an explicit
replay_history flag instead: history rows are written again (not stamped
durable) but keep their original visibility, and only this turn's new
rows get the mute treatment.
Deleting a parent session cascades to its delegate children. A live child
then can never be recreated: create_session(parent_session_id=...) hits
the parent FK on every flush for the agent's lifetime. When the recreate
fails and the parent row is gone, create once without the parent for that
call only; agent._parent_session_id is restored because the relay and
hooks still key on it.
Without entries in _STORAGE_FAILURES and _PERSISTENCE_CAUSE_EXPLANATIONS a
failed heal fell back to the generic disk/lock advice, which sends users
chasing the wrong problem.
The heal predicate (errno 787 or a 'foreign key constraint' substring)
and the classifier phrase were two separate definitions of the same
failure. Move the SQLITE_CONSTRAINT_FOREIGNKEY code check into
classify_persistence_error, document the bucket, and branch the heal on
agent._last_persistence_error_cause == 'session_row_missing' so the two
can't drift.
The heal retry passed the caller's conversation_history through, so
_db_flush_collect's id()-based history shortcut treated the prefix as
already durable and skipped it. The row was recreated with only the new
tail and the flush returned True: a silent partial restore. On the
session_row_missing retry, pass conversation_history=None so the full
in-memory transcript lands on the recreated row.
Co-authored-by: 赵桂雄 <daniel21436@hotmail.com>
`_ensure_db_session` trusts the cached `_session_db_created` flag as proof
the row exists, and the flush path only retries row creation while that
flag is False. Any store-side removal of the row under a live agent —
`hermes sessions delete`, the Desktop/web delete, bulk prune, a
profile-repair move, an in-place store rebuild — leaves the flag stale, so
every later turn's append fails the FK and is dropped with one WARNING per
turn. The agent keeps answering; the durable transcript silently stops
growing, and because the failed transaction leaves no rows behind there is
no post-hoc trace in the store.
Classification (`hermes_state_errors.py`): add the `session_row_missing`
cause — matched by `SQLITE_CONSTRAINT_FOREIGNKEY` (787) when the code
survives, else the RPC-wrapped phrase — so the turn-end explanation names
the real failure instead of "unknown".
Heal (`agent/session_persistence.py::_db_flush_failed`): on that cause,
drop the stale flag, reset the flush markers, call `_ensure_db_session()`,
and replay once within the flush's existing `_adoption_budget`. Because the
deletion erased the session's message rows too, the replay clears the
per-message persisted markers so the FULL in-memory transcript lands on the
recreated row, not just the current tail (mirrors
`_db_flush_adopt_compression_tip`). If row creation also fails, return
False without appending into a guaranteed rollback — fail-open, batch stays
unmarked for the next flush. No new fail-closed path.
(cherry picked from commit ad97047da0f9801a4fdca4478bfd68085f02c3a8)
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
Follow-up to the #123857 salvage. request_overrides is static config, so the
drop warning fired on every turn, retry and subagent call. It now fires once
per process. The Astra sanitizer docstring no longer suggests it handles the
key. The test is parametrized per route, and it pins the extra_body escape
hatch that the warning points to.
request_overrides are merged into the top-level Responses.create() kwargs,
where prompt_cache_options has no SDK parameter: the call fails with
TypeError before any request is sent, on every route. Drop the key after
the merge with a warning, matching the wire-boundary sanitizer precedent.
Wire-only fields still reach capable proxies via extra_body.
(cherry picked from commit c8e8227fdde75236246233565f9cb8ed74c6bd9c)
The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.
Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
DEFAULT_CONTEXT_LENGTHS declares claude-opus-5 as 1M, but the Bedrock
static table never got the entry, so the offline path resolved the 128K
default and the agent compressed context ~8x early. Add the key and pin
the table pairing with a test so the next 1M generation cannot drift.
Co-authored-by: JiaDe-Wu <JiaDe-Wu@users.noreply.github.com>
(cherry picked from commit c1be16ecbdd267b14124dd35c0b92ccfcd201818)
The streaming stale monitor re-fires every stale window while its worker
has not started/dispatched the attempt yet or is still unwinding a kill,
and each re-fire bumped agent._consecutive_stale_streams. One provider
attempt could therefore count twice (or more), so a turn with four hung
attempts reached HERMES_STREAM_STALE_GIVEUP=5 and the breaker refused the
NEXT turn although the provider only ever saw four stale requests. The
non-streaming and inline watchdogs already count once per call.
Count each started stream attempt once (attempt 0, i.e. nothing started,
never counts), and let the interrupted-wait bump honour the same ledger.
This is the root cause of the intermittent
test_agent_turn_liveness[provider_hang] red on CI (fault calls 4, five
"Stream stale for 3s" kills, probe requests 0, breaker text "5
consecutive stale attempts"): on a busy runner the first attempt's
worker takes >3 s to reach the wire, the monitor kills it before
dispatch (tcp_force_closed=0) and again 3 s after dispatch.
Repro: a 3.6 s sleep before the worker opens its first stream
reproduces the CI signature 4/4 on origin/main and 0/6 with the fix;
under taskset CPU starvation the unmodified module goes red 2/3 on base.
- _mark_finish_seen helper on the local _diag; also set on the Anthropic
path when message_delta carries stop_reason
- emit_stream_drop status: 'attempt N/M dropped, reconnecting'
- clean-EOF log wording: server or proxy closed the stream cleanly
- truncated_unreported copy drops the raw finish_reason placeholder
- turn_tool_validation compares against FINISH_REASON_LENGTH
When the 4 truncated-tool-call retries are exhausted on a clean-EOF stub
(no transport error, no finish_reason), use a dedicated stream_closed_tool_call
copy and failure_reason=truncated instead of stream_dropped_tool_call
('check your network') stamped as timeout. Also compute the truncation
banner in a plain if/elif chain instead of a 4-arm conditional expression.
Hand-applied from PR #91738 (1da9001929): the truncation detector moved from
conversation_loop.py to the _truncated branch of agent/turn_tool_validation.py.
Only claim the output cap when finish_reason == 'length'; otherwise use new
hedged site copy 'truncated_unreported' (agent/turn_failure_copy.py).
MessageStream._retry_after_drop passed attempt + 2 to _emit_stream_drop, so
the first drop was announced as attempt 2/N. Hand-applied from PR #90254
onto main's MessageStream split (original hunks targeted the pre-split
interruptible_streaming_api_call).
Response truncated - stream ended before completion" and the log line
"...treating as a mid-stream drop" fired identically for two distinct
failure classes: a genuine transport exception mid-stream, and the
provider closing the connection cleanly (no exception) without ever
sending a terminal finish_reason chunk. The second case is not a
network drop, but the shared wording sent users chasing a network
problem that never existed (#102766).
_build_partial_stream_stub() now tags its stub _clean_eof=True (every
caller is a clean-EOF site: the streaming loop ended without an
exception). The conversation loop uses that tag to pick distinct
wording; the exception-driven stub is unchanged. The two clean-EOF
log lines in chat_completion_helpers.py no longer say "drop".
stream_diag's per-attempt dict also gains finish_reason_seen, set the
moment a terminal chunk is observed, so log_stream_retry's retry/
exhaustion WARNING can say whether the dying attempt had already seen
a finish_reason before the exception hit.
(cherry picked from commit 2c1855ed6dd0037eebe0b1068628999832617109)
Desktop chat goes through tui_gateway, which never applied the gateway's
first-message profile-build sidecar note. Add shared first_contact_turn_note
helper and wire it into the prompt turn so fresh installs get the same
opt-in profile setup offer as messaging surfaces.
Adapted to current main (salvage of open PR #82765, fixes#82750):
_run_prompt_submit now lives in tui_gateway/prompt_turn.py; the note is
staged on agent._gateway_turn_context_notes (the existing sidecar channel,
consumed by agent.turn_context on the user message) so the system prompt
stays byte-stable; the prior-session probe rides _session_db(session).
The low-value onboarding test classes purged from main were not re-added —
only new coverage for first_contact_turn_note and the staging path.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
(cherry picked from commit b843900ca818912f0480f6926a059d292b63ffe6)
A session asked to clean up older Pythons removed the uv-managed base
interpreter its own venv depended on; the next boot died with 'uv
trampoline failed to spawn Python child process' and no agent tool could
repair it, because the agent itself no longer started (#58748). Prior
uninstall detection (85ce25687e) only flagged package-manager commands.
Add agent/runtime_self_protection.py and wire it into both layers:
- The approval floor (_floor_block) now blocks shell commands that
delete the running interpreter, its own venv, the pyvenv.cfg base, or
the uv-managed install directory — rm/rmdir/rd/del/erase/Remove-Item
with any flags, find <root> -delete, and uv python uninstall of the
running version (including --all). The floor runs before yolo /
approvals.mode=off / cron approve mode, so no session setting can
bypass it.
- The file-safety write classifier denies write/patch/move/delete to the
same paths, so the file tools cannot overwrite the interpreter either.
Only the runtime the process itself boots from is protected; every other
venv and interpreter on the machine stays manageable.
Fixes#58748
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.
Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
(_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
explicit operator stops (process_manage kill, /stop slash + RPC mirror,
CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
the host is going away and survivors would become PPID=1 orphans
Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
A turn that falls out of the loop after a tool result with no follow-up
assistant text left the durable transcript ending at a raw tool row and
returned a silent result: Desktop/TUI showed a ready composer (or kept
spinning) with no final message. The pending_tool_result explainer copy
already existed but nothing minted the exit reason.
finalize_turn now detects the non-interrupted tool tail, fails the turn
with turn_exit_reason=pending_tool_result, and synthesizes the visible
assistant close before persistence so the durable tail is alternation-
safe. Stream-recovered turns (#95514) and interrupted tails keep their
existing paths.
Fixes#55316Fixes#54756
Co-authored-by: blakehermes9 <blakehermes9@users.noreply.github.com>
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
Addresses two P1 review findings on #121614.
1. `provider_model_ids`: a failed or empty relay probe fell through to the
canonical per-provider fetcher, sending the provider credential to exactly
the vendor host the user routed away from — recreating #121387 on the
failure path. A configured `model.base_url` relay is now TERMINAL for live
catalog egress and degrades to the local curated list instead. The curated
tail is extracted as `_static_catalog` and shared by both paths.
Fetchers that already resolve `model.base_url` themselves and degrade
locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
simple api-key fetchers) are excluded from interception via
`_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
produce a better-merged catalog.
2. `_try_anthropic`: an `explicit_base_url` that failed
`_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
the ambient/canonical host and continuing with the explicit credential —
a silent retarget of authority, not the refusal the PR body claimed. It now
returns unavailable before client construction.
Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
Cron Codex now runs inline via direct_api_call, whose stale budget skipped the
openai-codex large-context floor (600/900/1200s) and HERMES_CODEX_HARD_TIMEOUT_SECONDS
cap applied on the worker path, so healthy >10k-token cron turns were killed at 90s.
Extract the floor+cap into _bound_openai_codex_stale_timeout, used by both paths.
Correct the should_use_direct_api_call docstring (Codex streaming goes through
_stream_codex_passthrough -> _interruptible_api_call, not _StreamingCall) and document
that the worker-only TTFB/idle watchdogs don't run inline. Replace the non-guarding
inline Codex watchdog test with one asserting the >=600s budget.
Cron Codex Responses calls still went through the spawned interrupt worker
(the #62151 nested-pool deadlock path). Route them through direct_api_call;
the Codex dispatch builds its client via make_client, so the inline stale
watchdog aborts a silent Codex stream even though the worker-only TTFB/idle
watchdogs are bypassed. Delegated children stay chat_completions-only.
Ported onto main's _InlineRequest refactor from PR #70087 (cherry picked
from e0028a5d73). Fixes#69734.
MiniMax M2.x reasoning models emit reasoning_content blocks before
their first content token (#17924). During extended thinking phases,
they routinely exceed the default 180s chat-model stale-stream
timeout, causing the stale-stream detector to kill the connection
mid-think.
This adds minimax-m2 to _REASONING_STALE_TIMEOUT_FLOORS with a
300s floor — generous enough to cover the documented 240s stall
(test_streaming.py:1270-1278) with margin.
Fixes#62353
(cherry picked from commit d66b2a75dee78f583edb75b699ba055ffcb93013)
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.