fix(agent): hosted providers no longer get the local-server "wait and /retry" context rejection
The unexplained-rejection gate from #114644 ended the turn with "another request on the same server was probably holding its capacity ... wait and /retry" whenever a server said "context exceeded" without a count while the local estimate sat under half the known window. That cause only exists on single-slot local servers. On a hosted route (Anthropic, Nous, OpenRouter, any public endpoint) the same rejection means the route's real window is smaller than the one Hermes assumes, so /retry failed identically every turn and the conversation was never compressed: Discord bots on claude-opus-5-5 were stuck repeating the message. The gate now also requires is_local_endpoint(base_url) (loopback, LAN, Tailscale, container DNS). Hosted endpoints return to the compress-and-retry path they had before #114644. The FAQ entry says which endpoints get the message.
This commit is contained in:
@@ -23,7 +23,7 @@ from agent.conversation_compression import (
|
||||
from agent.error_classifier import FailoverReason
|
||||
from agent.message_sanitization import serialized_messages_bytes
|
||||
from agent.model_metadata import (
|
||||
get_context_length_from_provider_error, is_output_cap_error,
|
||||
get_context_length_from_provider_error, is_local_endpoint, is_output_cap_error,
|
||||
parse_available_output_tokens_from_error,
|
||||
)
|
||||
from agent.turn_failure_copy import site_copy, stamp_failure
|
||||
@@ -402,11 +402,14 @@ def _recover_context_length(st: _Recovery, _retry: TurnRetryState, error_msg: st
|
||||
# request — a background review from an earlier session — holds their context. Name that,
|
||||
# keep the turn retryable and transient: no "conversation too long", no gateway auto-reset.
|
||||
# Only when the server quoted NO measurement of its own: "prompt is too long: 233153 tokens
|
||||
# > 200000" is the server's count and beats the local estimate.
|
||||
# > 200000" is the server's count and beats the local estimate. Local endpoints only: a hosted
|
||||
# route has no shared slot, so the same rejection means its real window is smaller than the one
|
||||
# Hermes assumes, and "wait and /retry" would fail identically forever; compress instead.
|
||||
window = agent.context_compressor.context_length
|
||||
request_tokens = st.request_tokens() + max(0, int(getattr(agent, "max_tokens", 0) or 0))
|
||||
if (
|
||||
not re.search(r"\d{4,}", error_msg)
|
||||
is_local_endpoint(agent.base_url)
|
||||
and not re.search(r"\d{4,}", error_msg)
|
||||
and isinstance(window, int) and window > 0
|
||||
and request_tokens < window * _UNEXPLAINED_REJECTION_FRACTION
|
||||
):
|
||||
|
||||
@@ -38,7 +38,8 @@ def test_overflow_exhaustion_is_non_retryable_context_overflow_with_slash_comman
|
||||
assert build_error_surface_from_result(result)["code"] == "context_overflow"
|
||||
|
||||
|
||||
def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 500: Context size has been exceeded."):
|
||||
def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 500: Context size has been exceeded.",
|
||||
base_url: str = "http://127.0.0.1:1234/v1"):
|
||||
"""Drive ``_recover_context_length`` with a provider "context exceeded" and a request the rough
|
||||
estimator prices at ``request_tokens`` against a ``window``-token model (no output cap)."""
|
||||
from unittest.mock import patch
|
||||
@@ -50,7 +51,7 @@ def _context_rejection(request_tokens: int, window: int = 65_536, error="HTTP 50
|
||||
st.compression_attempts = 0
|
||||
st.agent.max_tokens = None
|
||||
st.agent.context_compressor = SimpleNamespace(context_length=window)
|
||||
st.agent.provider, st.agent.base_url, st.agent.tools = "lmstudio", "http://127.0.0.1:1234/v1", None
|
||||
st.agent.provider, st.agent.base_url, st.agent.tools = "lmstudio", base_url, None
|
||||
st.agent._buffer_vprint = st.agent._buffer_diagnostic_status = lambda *a, **k: None
|
||||
compressed = []
|
||||
st.agent._compress_context = lambda msgs, *a, **k: (compressed.append(1) or [{"role": "user", "content": "x"}], None)
|
||||
@@ -87,6 +88,14 @@ def test_context_rejection_near_the_window_still_compresses():
|
||||
assert compressed and verdict.action == "break"
|
||||
|
||||
|
||||
def test_hosted_context_rejection_far_below_the_known_window_compresses():
|
||||
"""A hosted route has no shared slot to wait out: a small request rejected there means the
|
||||
route's real window is below the one Hermes assumes, so /retry would fail forever. Compress."""
|
||||
verdict, compressed = _context_rejection(
|
||||
47_000, window=1_000_000, base_url="https://api.anthropic.com",
|
||||
error="This model's maximum context length was exceeded. Please reduce the length of the messages.",
|
||||
)
|
||||
assert compressed and verdict.action == "break"
|
||||
|
||||
|
||||
def test_empty_response_exhaustion_has_one_text_everywhere():
|
||||
|
||||
@@ -339,7 +339,7 @@ Look at the CLI startup line — it shows the detected context length (e.g., `
|
||||
|
||||
**Local servers (llama.cpp, Ollama) that go silent instead of erroring:** when a provider rejects a request as too large, Hermes compacts the conversation and rebuilds the request. Hermes re-measures the *complete* rebuilt request (system prompt + tool schemas + messages) before retrying, and runs further bounded compaction passes if it is still over the threshold. If the request still cannot fit, the turn ends with `Context length exceeded: compression could not reduce the rebuilt request below the safe threshold` rather than sending an oversized request that llama.cpp would silently truncate (`stop processing: n_tokens = 65535, truncated = 1` in the server log). If you hit that message, the fix is almost always the configured `context_length` above: make it match the server's actual `-c` / `--ctx-size`.
|
||||
|
||||
**"The model server rejected this request as too large, but this conversation is only about N tokens…":** the server said "context exceeded" without quoting any measurement, while Hermes's own estimate of the request is far below the window it knows for the model — so it does **not** compress or blame the conversation, and the turn stays retryable. On single-slot local servers (LM Studio, Ollama) this is almost always another request holding the server's context at that moment — typically a background memory review from an earlier session (`thread=bg-review` in `logs/agent.log`). Wait a moment and `/retry`. If it recurs with no other Hermes process running, the server is loading the model with a smaller window than Hermes assumes: raise the server's context setting or lower `model.context_length` to match it.
|
||||
**"The model server rejected this request as too large, but this conversation is only about N tokens…":** a local server (localhost, LAN, Tailscale) said "context exceeded" without quoting any measurement, while Hermes's own estimate of the request is far below the window it knows for the model — so it does **not** compress or blame the conversation, and the turn stays retryable. On single-slot local servers (LM Studio, Ollama) this is almost always another request holding the server's context at that moment — typically a background memory review from an earlier session (`thread=bg-review` in `logs/agent.log`). Wait a moment and `/retry`. If it recurs with no other Hermes process running, the server is loading the model with a smaller window than Hermes assumes: raise the server's context setting or lower `model.context_length` to match it. Hosted providers never get this message: they have no shared slot to wait out, so the same rejection there means the route's real window is smaller than Hermes assumes, and Hermes compresses and retries instead.
|
||||
|
||||
To fix context detection, set it explicitly:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user