Commit Graph

15 Commits

Author SHA1 Message Date
teknium1
24f03c41c1 fix(agent): hosted providers no longer get the local-server "wait and /retry" context rejection
The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.

The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
2026-09-23 04:38:27 -07:00
teknium1
f1737e1f0a fix: output-cap failure copy stops pointing at the removed model.max_tokens key
fd3565deec removed every user-facing output-cap control, but the unparseable
output-cap failure in agent/turn_overflow.py::_recover_context_length still told users
to "Lower model.max_tokens in config.yaml" — a key Hermes no longer reads. The notice
now explains what happened (the provider's cap was exceeded and the error did not
state the limit) and where the cap actually lives (the endpoint's default max output
tokens on the server or proxy) without naming a config key that does not exist.

Tests: one parametrized invariant for the Azure and SGLang ceiling wordings plus the
SGLang input-alone-over-window control (red on base).
2026-09-19 00:59:27 -07:00
teknium1
5e95050608 fix(agent): a server context rejection the transcript cannot explain is no longer "conversation too long"
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).

_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.

Fixes #114644
2026-09-18 19:18:57 -07:00
Victor Kyriazakos
541b20290c fix(bedrock): restore Grok context with provider-confirmed cache provenance 2026-09-17 18:36:10 -07:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
Teknium
89fbd5d4d3 simplify(compat): conversation_loop — drop 9 re-exports, repoint 7 callers, 35 test sites 2026-09-03 13:12:23 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
1a025a4cc6 refactor(agent/turn): AST-neutral packing of multi-line f-string literals 2026-09-02 18:51:40 -07:00
Teknium
761cff3221 refactor(agent/turn): working-state carriers subclass their Verdict (drop duplicate field copies) 2026-09-02 18:49:52 -07:00
Teknium
498c7a1c08 refactor(agent/turn): AST-neutral packing of multi-line string literals 2026-09-02 18:45:24 -07:00
Teknium
4cf65d7d89 refactor(agent/turn_overflow): unify attempt counting and token-scored compression across the overflow handlers 2026-09-02 18:30:28 -07:00
Teknium
c94ced6225 refactor(agent/turn): AST-neutral bracket/signature packing across r3-08 slice 2026-09-02 18:01:27 -07:00
Teknium
9c99769ad8 refactor(turn): split recover_from_overflow into per-error handlers on a _Recovery state (501-LOC function -> max 76) 2026-09-02 15:29:36 -07:00
Teknium
645db06053 refactor(agent): extract 413/context-overflow compression recovery into agent/turn_overflow.py 2026-09-02 13:30:17 -07:00