Drop the duplicate chunk-reading helper, run_agent facade forward and
_last_serving_provider agent state; the chat-completions loop already captures
chunk.provider, so stamp it on the per-attempt diag there and read the hook's
upstream_provider from the assembled response.provider.
Relay routing is re-rolled per request and the winning downstream is reported only
inside the delta chunk bodies, so the header snapshot in agent/stream_diag.py could
not attribute a mid-stream drop to a provider. The per-attempt diag now carries
serving_provider (first non-empty chunk-body provider), log_stream_retry prints it,
and the post_api_request payload exposes it as upstream_provider for plugins auditing
route compliance. Fixes#90216.
(cherry picked from commit f0cf4fe4f4b47e384532008b1bf56ad259c87dc2)
Three consecutive Codex Responses answers that carry only (encrypted) reasoning —
no visible text, no tool call — used to exhaust the 3-continuation budget and end
the turn on "Codex response remained incomplete after 3 continuation attempts",
never touching configured fallback_providers (#67321). Encrypted reasoning items
replay byte-for-byte, so a bare retry deterministically repeats the stall.
- Track a per-turn `_codex_reasoning_only_streak` apart from the aggregate
`_codex_incomplete_retries`: a visible partial resets the streak, so the mixed
partial-then-stall variant still reaches its own recovery threshold while the
turn-wide iteration budget stays the hard bound.
- At streak 3, `continue_codex_incomplete` activates the next fallback with the
semantic `FailoverReason.incomplete_response`, grants exactly one grace call when
the trigger consumed the last iteration, and returns `CODEX_FALLBACK_ACTIVATED`;
the intake re-syncs the Model:/Provider: identity on the system prompt.
- Off the Codex wire the synthetic continuation nudge is stripped alongside the
opaque replay state (`drop_nudge_marker`) so the Chat Completions payload keeps
valid role ordering and no Codex-only control text.
- No fallback configured: unchanged terminal sentinel, still bounded at 3 calls.
Ported from PR #67336 by @PRATHAMESH75 onto the decomposed agent/turn_*.py siblings.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
When a Responses-API call (muse-spark, gpt-*, grok-* routes) ends with
status=incomplete / incomplete_details.reason=max_output_tokens and no visible
text, reasoning consumed the whole output budget. continue_codex_incomplete
re-sent the identical request three times: same max_output_tokens, same
reasoning effort, plus a nudge. The model re-burned the same budget each time
and the turn ended as "Codex response remained incomplete after 3 continuation
attempts" with nothing for the user (#90393; measured in #103483 at xhigh:
out_tokens == max_output_tokens, reasoning_tokens == out - 3).
The chat-completions length path already handles this with two one-shot
overrides (_ephemeral_reasoning_off, _ephemeral_max_output_tokens); reuse them:
on a budget-exhausted empty fragment set reasoning off and double the cap
(2x, 4x, capped at 32768 / the configured cap) for the next attempt, and make
_build_codex_kwargs consume the ephemeral cap it never read before. A
reasoning-only status=completed response (Codex "still thinking") is untouched.
Live repro (fake Responses provider, real loop): before 2000/high, 2000/high,
2000/high -> partial; after 2000/high, 4000/no reasoning, 8000/no reasoning.
Chat-completions control unchanged (2000->4000->8000->16000, effort none).