fix(agent): hosted providers no longer get the local-server "wait and /retry" context rejection

The unexplained-rejection gate from #114644 ended the turn with "another request on the same
server was probably holding its capacity ... wait and /retry" whenever a server said "context
exceeded" without a count while the local estimate sat under half the known window. That cause
only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any
public endpoint) the same rejection means the route's real window is smaller than the one Hermes
assumes, so /retry failed identically every turn and the conversation was never compressed:
Discord bots on claude-opus-5-5 were stuck repeating the message.

The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container
DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ
entry says which endpoints get the message.
This commit is contained in:
teknium1
2026-09-23 03:41:21 -07:00
committed by Teknium
parent c80d12b9b9
commit 24f03c41c1
3 changed files with 18 additions and 6 deletions

View File

@@ -23,7 +23,7 @@ from agent.conversation_compression import (
from agent.error_classifier import FailoverReason
from agent.message_sanitization import serialized_messages_bytes
from agent.model_metadata import (
get_context_length_from_provider_error, is_output_cap_error,
get_context_length_from_provider_error, is_local_endpoint, is_output_cap_error,
parse_available_output_tokens_from_error,
)
from agent.turn_failure_copy import site_copy, stamp_failure
@@ -402,11 +402,14 @@ def _recover_context_length(st: _Recovery, _retry: TurnRetryState, error_msg: st
# request — a background review from an earlier session — holds their context. Name that,
# keep the turn retryable and transient: no "conversation too long", no gateway auto-reset.
# Only when the server quoted NO measurement of its own: "prompt is too long: 233153 tokens
# > 200000" is the server's count and beats the local estimate.
# > 200000" is the server's count and beats the local estimate. Local endpoints only: a hosted
# route has no shared slot, so the same rejection means its real window is smaller than the one
# Hermes assumes, and "wait and /retry" would fail identically forever; compress instead.
window = agent.context_compressor.context_length
request_tokens = st.request_tokens() + max(0, int(getattr(agent, "max_tokens", 0) or 0))
if (
not re.search(r"\d{4,}", error_msg)
is_local_endpoint(agent.base_url)
and not re.search(r"\d{4,}", error_msg)
and isinstance(window, int) and window > 0
and request_tokens < window * _UNEXPLAINED_REJECTION_FRACTION
):

View File

@@ -38,7 +38,8 @@ def test_overflow_exhaustion_is_non_retryable_context_overflow_with_slash_comman
assert build_error_surface_from_result(result)["code"] == "context_overflow"
def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 500: Context size has been exceeded."):
def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 500: Context size has been exceeded.",
base_url: str = "http://127.0.0.1:1234/v1"):
"""Drive ``_recover_context_length`` with a provider "context exceeded" and a request the rough
estimator prices at ``request_tokens`` against a ``window``-token model (no output cap)."""
from unittest.mock import patch
@@ -50,7 +51,7 @@ def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 50
st.compression_attempts = 0
st.agent.max_tokens = None
st.agent.context_compressor = SimpleNamespace(context_length=window)
st.agent.provider, st.agent.base_url, st.agent.tools = "lmstudio", "http://127.0.0.1:1234/v1", None
st.agent.provider, st.agent.base_url, st.agent.tools = "lmstudio", base_url, None
st.agent._buffer_vprint = st.agent._buffer_diagnostic_status = lambda *a, **k: None
compressed = []
st.agent._compress_context = lambda msgs, *a, **k: (compressed.append(1) or [{"role": "user", "content": "x"}], None)
@@ -87,6 +88,14 @@ def test_context_rejection_near_the_window_still_compresses():
assert compressed and verdict.action == "break"
def test_hosted_context_rejection_far_below_the_known_window_compresses():
"""A hosted route has no shared slot to wait out: a small request rejected there means the
route's real window is below the one Hermes assumes, so /retry would fail forever. Compress."""
verdict, compressed = _context_rejection(
47_000, window=1_000_000, base_url="https://api.anthropic.com",
error="This model's maximum context length was exceeded. Please reduce the length of the messages.",
)
assert compressed and verdict.action == "break"
def test_empty_response_exhaustion_has_one_text_everywhere():

View File

@@ -339,7 +339,7 @@ Look at the CLI startup line — it shows the detected context length (e.g., `
**Local servers (llama.cpp, Ollama) that go silent instead of erroring:** when a provider rejects a request as too large, Hermes compacts the conversation and rebuilds the request. Hermes re-measures the *complete* rebuilt request (system prompt + tool schemas + messages) before retrying, and runs further bounded compaction passes if it is still over the threshold. If the request still cannot fit, the turn ends with `Context length exceeded: compression could not reduce the rebuilt request below the safe threshold` rather than sending an oversized request that llama.cpp would silently truncate (`stop processing: n_tokens = 65535, truncated = 1` in the server log). If you hit that message, the fix is almost always the configured `context_length` above: make it match the server's actual `-c` / `--ctx-size`.
**"The model server rejected this request as too large, but this conversation is only about N tokens…":** the server said "context exceeded" without quoting any measurement, while Hermes's own estimate of the request is far below the window it knows for the model — so it does **not** compress or blame the conversation, and the turn stays retryable. On single-slot local servers (LM Studio, Ollama) this is almost always another request holding the server's context at that moment — typically a background memory review from an earlier session (`thread=bg-review` in `logs/agent.log`). Wait a moment and `/retry`. If it recurs with no other Hermes process running, the server is loading the model with a smaller window than Hermes assumes: raise the server's context setting or lower `model.context_length` to match it.
**"The model server rejected this request as too large, but this conversation is only about N tokens…":** a local server (localhost, LAN, Tailscale) said "context exceeded" without quoting any measurement, while Hermes's own estimate of the request is far below the window it knows for the model — so it does **not** compress or blame the conversation, and the turn stays retryable. On single-slot local servers (LM Studio, Ollama) this is almost always another request holding the server's context at that moment — typically a background memory review from an earlier session (`thread=bg-review` in `logs/agent.log`). Wait a moment and `/retry`. If it recurs with no other Hermes process running, the server is loading the model with a smaller window than Hermes assumes: raise the server's context setting or lower `model.context_length` to match it. Hosted providers never get this message: they have no shared slot to wait out, so the same rejection there means the route's real window is smaller than Hermes assumes, and Hermes compresses and retries instead.
To fix context detection, set it explicitly: