16 Commits

Author SHA1 Message Date
kshitijk4poor
1f041552d7 fix(stream): tidy clean-EOF diagnostics and truncation copy
- _mark_finish_seen helper on the local _diag; also set on the Anthropic
  path when message_delta carries stop_reason
- emit_stream_drop status: 'attempt N/M dropped, reconnecting'
- clean-EOF log wording: server or proxy closed the stream cleanly
- truncated_unreported copy drops the raw finish_reason placeholder
- turn_tool_validation compares against FINISH_REASON_LENGTH
2026-09-26 06:10:53 +05:30
chelsealong
a2a19bcc77 fix(stream): distinguish clean-EOF no-finish_reason from a transport drop
Response truncated - stream ended before completion" and the log line
"...treating as a mid-stream drop" fired identically for two distinct
failure classes: a genuine transport exception mid-stream, and the
provider closing the connection cleanly (no exception) without ever
sending a terminal finish_reason chunk. The second case is not a
network drop, but the shared wording sent users chasing a network
problem that never existed (#102766).

_build_partial_stream_stub() now tags its stub _clean_eof=True (every
caller is a clean-EOF site: the streaming loop ended without an
exception). The conversation loop uses that tag to pick distinct
wording; the exception-driven stub is unchanged. The two clean-EOF
log lines in chat_completion_helpers.py no longer say "drop".

stream_diag's per-attempt dict also gains finish_reason_seen, set the
moment a terminal chunk is observed, so log_stream_retry's retry/
exhaustion WARNING can say whether the dying attempt had already seen
a finish_reason before the exception hit.

(cherry picked from commit 2c1855ed6dd0037eebe0b1068628999832617109)
2026-09-26 06:10:53 +05:30
kshitijk4poor
1e04c64ef3 fix(stream_diag): feed the existing chunk-body upstream_provider into diag + hook payload (#90216)
Drop the duplicate chunk-reading helper, run_agent facade forward and
_last_serving_provider agent state; the chat-completions loop already captures
chunk.provider, so stamp it on the per-attempt diag there and read the hook's
upstream_provider from the assembled response.provider.
2026-09-24 22:25:06 +05:30
kouyichi
e058815c83 chore(stream_diag): log the swallowed provider-note failure at DEBUG
Review follow-up on #118759: the bare `except Exception: pass` in
stream_diag_note_serving_provider is intentional (best-effort annotation that
must never break streaming) but indistinguishable from a missed error-handling
gap. Keep the swallow, add a DEBUG log with exc_info so the intent is legible.

(cherry picked from commit 6674471642c4bdb81ed212cd1cee0242647b60ed)
2026-09-24 22:25:06 +05:30
kouyichi
0370beae2e fix(agent): record the serving downstream provider in stream-drop diagnostics
Relay routing is re-rolled per request and the winning downstream is reported only
inside the delta chunk bodies, so the header snapshot in agent/stream_diag.py could
not attribute a mid-stream drop to a provider. The per-attempt diag now carries
serving_provider (first non-empty chunk-body provider), log_stream_retry prints it,
and the post_api_request payload exposes it as upstream_provider for plugins auditing
route compliance. Fixes #90216.

(cherry picked from commit f0cf4fe4f4b47e384532008b1bf56ad259c87dc2)
2026-09-24 22:25:06 +05:30
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
Teknium
fcbe4acbef simplify(compat): tools/mcp_tool — repoint 20 non-test callers to the defining mcp_tool_* siblings 2026-09-03 13:29:35 -07:00
Teknium
46020c32c5 refactor(agent/stream_diag): compact retry log call and drop-suffix computation 2026-09-02 19:43:36 -07:00
Teknium
9a0fa659af refactor(agent/thread_scoped_output,stream_diag,stream_single_writer): shared fence-call helper, compact installs 2026-09-02 19:42:28 -07:00
Teknium
731e03ac77 refactor(agent/stream_diag,status_output): compact signatures, replay dispatch, chain rendering 2026-09-02 19:21:14 -07:00
Teknium
9d7b2e5ff1 refactor(agent/stream_diag): extract retry-log diag field formatting 2026-09-02 19:08:30 -07:00
Teknium
6fe8f9e76c refactor(agent/stream_diag,repetition_guard,thread_scoped_output,stream_single_writer): compact loops and docstrings 2026-09-02 18:35:30 -07:00
Teknium
a53d886584 refactor(agent/transports): compact ACP/SSL/stream helper modules (copilot_acp_client, acp_openai_bridge, ssl_*, stream_*, jiter_preload) 2026-09-02 13:29:48 -07:00
Teknium
67011cc0d7 feat(agent): buffer retry/fallback status, surface only on terminal failure (#33816)
Users report that the CLI/gateway floods them with confusing retry chatter
during transient failures: a single 429 can produce 10+ "Provider/Endpoint/
Retrying in 5s..." lines before the request eventually succeeds. The same
firehose hits Telegram, Discord, Slack, etc. via _emit_status.

This patch defers all retry/fallback/compression status messages until we
know the outcome:
  - if the turn ultimately succeeds (any path: primary recovers, fallback
    activates, compression unsticks the request), the buffer is silently
    dropped — the user sees nothing.
  - if every retry and fallback exhausts and the turn fails, the buffer
    is flushed at the terminal-failure return so the user sees the full
    retry trace alongside the final error.

Backend logging (agent.log) is unchanged — every emission site still
writes to logger.warning/info, so post-mortem diagnosis is intact.

## What changed

run_agent.py: four new methods on AIAgent:
  _buffer_status(msg)   — defer an _emit_status call
  _buffer_vprint(msg)   — defer a _vprint(force=True) line
  _clear_status_buffer() — drop pending messages on success
  _flush_status_buffer() — replay pending messages on terminal failure

agent/conversation_loop.py:
  - converted ~30 mid-process emit/vprint sites in the retry, fallback,
    compression, empty-response, and stream-watchdog paths to the buffered
    helpers
  - added _flush_status_buffer() at every terminal-failure return so users
    still see the trace when it actually matters
  - added _clear_status_buffer() at the "non-empty assistant content"
    point (NOT at "API call returned bytes" — empty responses still loop
    through the empty-retry path and would otherwise lose their trace
    between iterations)
  - silenced the two "(´;ω;`) oops, retrying..." / "(╥_╥) error,
    retrying..." spinner final-frame messages — the spinner now stops
    cleanly so retries leave no visible residue

agent/chat_completion_helpers.py: same conversion for codex TTFB / stale-
stream / fallback-activation status messages.

agent/stream_diag.py: _emit_stream_drop now buffers instead of emitting
directly.

## Tests

tests/run_agent/test_retry_status_buffer.py: 7 unit tests covering
accumulate→flush, clear-on-success, mixed kinds, empty-buffer no-op,
re-buffer after flush, exception swallowing.

Updated 3 existing tests that mocked _emit_status to also mock (or use)
_buffer_status:
  - tests/run_agent/test_run_agent.py::test_empty_response_emits_status_for_gateway
  - tests/run_agent/test_stream_drop_logging.py (2 tests)
  - tests/agent/test_codex_ttfb_watchdog.py (TTFB hint test)

## Validation

Live test: hermes chat -q against an unreachable endpoint with no fallback
exhausts retries and prints the full trace at the end. Same flow against
a working endpoint prints zero retry chatter.
2026-05-28 04:53:27 -07:00
teknium1
57f6762ca0 refactor(run_agent): extract stream diagnostics to agent/stream_diag.py
Move the five stream-drop diagnostic helpers + the headers tuple:

* STREAM_DIAG_HEADERS — cf-ray, x-openrouter-provider, x-request-id, etc.
* stream_diag_init — fresh per-attempt diagnostic dict
* stream_diag_capture_response — snapshot upstream headers + HTTP status
* flatten_exception_chain — compact Outer(msg) <- Inner(msg) rendering
* log_stream_retry — structured WARNING with provider/bytes/elapsed/ttfb
* emit_stream_drop — user-facing status line + activity touch

AIAgent keeps thin forwarder methods (and exposes the headers tuple as
_STREAM_DIAG_HEADERS for back-compat).  All test patches and call sites
unchanged.

tests/run_agent/ + tests/agent/: 4313 passed (same pre-existing
test_auxiliary_client failure).

run_agent.py: 13470 -> 13227 lines (-243).
2026-05-16 18:28:17 -07:00