Some OpenAI-compatible routers answer an upstream connect timeout with a 200
ChatCompletion whose only assistant text is "Connect timeout, please try again
later." and zero completion tokens. The cherry-picked predicate only guarded
ChatCompletionsTransport.validate_response(); the streaming assembler emitted
the text to live callbacks before that check, the iteration-limit summary
normalized the response directly, and the auxiliary _validate_llm_response()
only checked the message shape — so the router failure surfaced as the answer
(#68396, maintainer keep_open review on #68433).
Share the predicate (is_router_timeout_shim) and apply it where each path reads
the response:
- streaming: hold shim-prefix text in the existing pending-text seam used for
echoed SSE, and release it at assembly only when the finished response is not
a shim; the assembled object then fails validate_response and the loop retries
- iteration summary: a shim reads as an empty summary, taking the retry slot
- auxiliary: a shim raises like a malformed response so the fallback chain moves on
Live probe (fake OpenAI-compatible server, first reply shim, real AIAgent, temp
HERMES_HOME): before, non-stream / stream / summary all returned the shim text
(stream also emitted it live); after, all three retry and return the real answer;
control (no shim) still makes one call.