feat(nous): anthropic_wire=auto decides a session's wire from its first response

Portal will serve anthropic/* from more than one upstream (OpenRouter
passthrough today; GMI/Vertex once it is back online). The native Messages
wire is the better transport but is only safe where the upstream keeps
prompt-cache routing sticky: measured false on the OpenRouter path (14-20% of
consecutive calls re-write the previous turn; #104284 moved the default to
chat), untested on GMI. Hermes cannot see the upstream in the request, only
in the response: OpenRouter stamps `provider` (chat wire) and mints
`gen-<unix>-<rand>` ids; GMI/Vertex returns Anthropic-native `msg_...` ids
and no provider.

`auto` therefore starts every session on chat (correct on both upstreams),
classifies the first response, and switches that session to native only
when the upstream is GMI AND `agent/nous_wire.py::GMI_NATIVE_WIRE_CLEARED`
is True. The switch is scheduled at response time and applied at the start
of the next iteration (turn_iteration_prep), so nothing is rebuilt while a
response is being consumed; it goes through switch_model so the client,
cache policy and _primary_runtime stay consistent. One decision per session,
call 1 only; unknown upstream never switches; a failed switch logs and stays.

GMI_NATIVE_WIRE_CLEARED is False: until the 20x6 concurrency probe
(evals/postmortem/live_ab) is clean on a GMI-served anthropic/* id on the
native wire, `auto` behaves exactly like `chat`. Flipping it is the whole
rollout once GMI is measured. Default stays `chat`.

Tests (17): classifier on real Portal response shapes from both wires and
both upstreams; chat for openrouter/unknown, GMI gated on the flag; one
decision per session, call 1 only, explicit chat/native never auto-switch,
other providers/models untouched, switch failure swallowed and final;
record_response_usage on a real AIAgent invokes the hook once.

Live (auto, real Portal, Fable 5.1): arm A, real classification
(OpenRouter today) - stays on chat through a tool loop and a second turn,
cache 97-99%. Arm B, classifier forced to gmi with the flag on - call 1 on
chat, switch applied before call 2, calls 2-3 on the native wire in the
same session, tool result and both turns correct, cache 97-99%. An earlier
shape that switched inside the response path broke call 1 (SimpleNamespace
has no .content); the scheduled apply is why.
This commit is contained in:
Teknium
2026-09-06 09:11:04 -07:00
parent 314e365363
commit 45646f4d09
7 changed files with 277 additions and 8 deletions

112
agent/nous_wire.py Normal file
View File

@@ -0,0 +1,112 @@
"""Nous Portal ``anthropic/*`` wire selection when ``nous.anthropic_wire`` is ``auto``.
Portal serves Claude two ways and Hermes cannot tell which from the request: an OpenRouter
passthrough (today, for every ``anthropic/*`` id) or GMI/Vertex (planned once GMI is back). The
native Messages wire is the better transport, but on the OpenRouter path it re-writes the previous
turn's prompt cache on 14-20% of consecutive calls in concurrent tool loops (measured 2026-09-06;
NousResearch/api#227), so the session must ride chat/completions there. On GMI that is untested,
and until it is measured ``auto`` never promotes to native.
The upstream IS visible in the first RESPONSE: OpenRouter stamps ``provider`` (chat wire) and
mints ``gen-<unix>-<rand>`` ids; GMI/Vertex responses carry neither. So ``auto`` starts every
session on chat (safe on both upstreams), reads the first response, and switches the session to
native only when the upstream is GMI and native has been cleared for GMI. One decision per
session, at call 1, before there is a cache to lose; later calls never flip.
``classify_upstream`` is pure and unit-tested; ``maybe_switch_wire_after_first_response`` is
the single hook, called from the usage recorder.
"""
from __future__ import annotations
import logging
import re
from typing import Any, Optional
logger = logging.getLogger(__name__)
# Flip to True only after the 20x6 concurrency probe (evals/postmortem/live_ab) is clean on a
# GMI-served anthropic/* id on the native wire. Until then ``auto`` is chat everywhere.
GMI_NATIVE_WIRE_CLEARED = False
_OPENROUTER_ID = re.compile(r"^gen-\d{9,}-[A-Za-z0-9_-]{8,}$")
def classify_upstream(response: Any) -> Optional[str]:
"""``"openrouter"`` / ``"gmi"`` / ``None`` (unknown) from a Portal response object.
Works on both wires: the OpenAI SDK object exposes ``.provider`` (OpenRouter's upstream name,
e.g. ``"Anthropic"``, ``"Amazon Bedrock"``) and an OpenRouter-minted ``.id``; the Anthropic SDK
object has ``.id`` only. GMI/Vertex responses have Anthropic-native ``msg_…`` ids and no
``provider``. Anything else is unknown, and unknown never triggers a switch.
"""
if response is None:
return None
if isinstance(getattr(response, "provider", None), str) and getattr(response, "provider"):
return "openrouter"
rid = getattr(response, "id", None)
if isinstance(rid, str):
if _OPENROUTER_ID.match(rid):
return "openrouter"
if rid.startswith("msg_"):
return "gmi"
return None
def wire_for_upstream(upstream: Optional[str]) -> str:
"""The api_mode ``auto`` wants once the upstream is known. Chat unless GMI and cleared."""
if upstream == "gmi" and GMI_NATIVE_WIRE_CLEARED:
return "anthropic_messages"
return "chat_completions"
def maybe_switch_wire_after_first_response(agent: Any, response: Any, api_call_count: int) -> bool:
"""Decide the session's wire from its first response; the switch itself is applied at the
start of the next iteration (``apply_pending_wire_switch``), never while a response is being
consumed. Returns True when a switch was scheduled.
Only for provider=nous, anthropic/* models, ``nous.anthropic_wire: auto``, and only on the
session's first API call.
"""
if api_call_count != 1 or getattr(agent, "_nous_wire_decided", False):
return False
if (getattr(agent, "provider", "") or "").lower() != "nous":
return False
model = str(getattr(agent, "model", "") or "")
if not model.lower().startswith("anthropic/"):
return False
try:
from hermes_cli.providers import _nous_anthropic_wire
if _nous_anthropic_wire() != "auto":
return False
except Exception:
return False
agent._nous_wire_decided = True # one decision per session, whatever it is
upstream = classify_upstream(response)
want = wire_for_upstream(upstream)
if want == getattr(agent, "api_mode", None):
logger.debug("nous wire auto: upstream=%s, staying on %s", upstream, want)
return False
agent._nous_wire_pending = (want, upstream)
return True
def apply_pending_wire_switch(agent: Any) -> bool:
"""At iteration start, with no response in flight: perform the switch scheduled by
``maybe_switch_wire_after_first_response``. Reuses ``switch_model`` (same model/provider, new
api_mode) so client rebuild, cache policy and ``_primary_runtime`` stay consistent. A failure
is logged and the session stays on its current wire."""
pending = getattr(agent, "_nous_wire_pending", None)
if not pending:
return False
agent._nous_wire_pending = None
want, upstream = pending
try:
from agent.agent_runtime_helpers import switch_model
switch_model(agent, agent.model, "nous", api_key=getattr(agent, "api_key", "") or "",
base_url=getattr(agent, "base_url", "") or "", api_mode=want)
except Exception as exc: # never let wire selection break a turn
logger.warning("nous wire auto: switch to %s failed (%s); staying on %s", want, exc, agent.api_mode)
return False
logger.info("nous wire auto: upstream=%s -> %s for the rest of session %s", upstream, want,
getattr(agent, "session_id", "?"))
return True

View File

@@ -39,6 +39,12 @@ def prepare_iteration(agent: Any,*, messages: Any, api_call_count: Any) -> Itera
_INTERRUPT_SCAFFOLD_MARKER, _maybe_inject_run_budget_wrapup
)
# nous.anthropic_wire=auto: a wire switch decided from the previous response lands here,
# before this iteration's request is built and with nothing in flight.
if getattr(agent, "_nous_wire_pending", None):
from agent.nous_wire import apply_pending_wire_switch
apply_pending_wire_switch(agent)
# Fire step_callback for gateway hooks (agent:step event).
if agent.step_callback is not None:
try:

View File

@@ -181,6 +181,11 @@ def record_response_usage(
prompt_tokens, completion_tokens, total_tokens,
api_duration, _cache_pct,
)
# nous.anthropic_wire=auto: the session's wire is decided once, from this first response.
if agent.session_api_calls == 1 and (agent.provider or "") == "nous":
with suppress(Exception):
from agent.nous_wire import maybe_switch_wire_after_first_response
maybe_switch_wire_after_first_response(agent, response, agent.session_api_calls)
# MoA: agent.model/provider are the virtual preset/"moa" with no pricing entry, silently
# dropping aggregator spend. Price at the REAL model/provider from the aggregator slot.

View File

@@ -2286,10 +2286,12 @@ DEFAULT_CONFIG = {
"keepalive_interval_seconds": 900,
# anthropic_wire: which Portal route carries anthropic/* models. "chat" =
# /v1/chat/completions (default for now); "native" = /v1/messages, the Anthropic
# Messages wire (signed thinking passthrough, native cache_control scopes). Native is the
# better wire but re-writes the previous turn's cache on 14-20% of consecutive calls in
# concurrent tool loops (measured 2026-09-06; NousResearch/api#227), so chat is the
# default until that is fixed.
# Messages wire (signed thinking passthrough, native cache_control scopes); "auto" =
# start on chat and, per session, switch to native from the first response when the
# Portal upstream serving the model is one where native is known clean. Native is the
# better wire but on the OpenRouter-served path it re-writes the previous turn's cache on
# 14-20% of consecutive calls in concurrent tool loops (measured 2026-09-06;
# NousResearch/api#227), so chat is the default until that is fixed.
"anthropic_wire": "chat",
},
# Google Vertex AI (Gemini). Auth is OAuth2 from a service-account JSON or ADC, NOT an API key;

View File

@@ -322,18 +322,21 @@ def nous_api_mode(model: str = "") -> str:
signed native blocks, and cache_control scopes are translated by the portal's adapter.
Empty/unknown model defaults to ``chat_completions`` (the historical Nous transport)."""
if str(model or "").strip().lower().startswith("anthropic/"):
# ``auto`` starts on chat too: it is safe on every upstream, and ``agent/nous_wire.py``
# promotes the session to native from the first response when the upstream allows it.
return "anthropic_messages" if _nous_anthropic_wire() == "native" else "chat_completions"
return "chat_completions"
def _nous_anthropic_wire() -> str:
"""``nous.anthropic_wire``: ``"chat"`` (default) or ``"native"``. Anything else reads as ``chat``."""
"""``nous.anthropic_wire``: ``"chat"`` (default), ``"native"``, or ``"auto"`` (chat, then per-session
promotion decided from the first response; see ``agent/nous_wire.py``). Anything else reads as ``chat``."""
try:
from hermes_cli.config import load_config_readonly
value = str(((load_config_readonly().get("nous") or {}).get("anthropic_wire")) or "chat").strip().lower()
except Exception:
return "chat"
return "native" if value == "native" else "chat"
return value if value in ("native", "auto") else "chat"
def determine_api_mode(provider: str, base_url: str = "", model: str = "") -> str:

View File

@@ -0,0 +1,139 @@
"""``nous.anthropic_wire: auto`` decides a session's wire once, from its first response.
Contracts:
- the upstream classifier reads what Portal actually returns on both wires (OpenRouter stamps
``provider`` + ``gen-…`` ids; GMI/Vertex returns Anthropic-native ``msg_…`` ids, no provider);
- ``auto`` starts on chat and never promotes to native until GMI native is cleared;
- one decision per session, only on call 1, only for nous + anthropic/*, never for chat/native;
- through a real AIAgent and the real usage recorder, the hook fires exactly once and a switch
failure never breaks the turn.
"""
from __future__ import annotations
from types import SimpleNamespace
import pytest
from agent import nous_wire
from hermes_cli import providers as _providers
def _resp(**kw):
return SimpleNamespace(**kw)
class TestClassifier:
def test_openrouter_by_provider_field_or_gen_id(self):
assert nous_wire.classify_upstream(_resp(provider="Anthropic", id="gen-1788708403-xXjDsFw")) == "openrouter"
assert nous_wire.classify_upstream(_resp(provider="Claude Platform on AWS", id="x")) == "openrouter"
# native wire: no provider field, but the id is still OpenRouter-minted
assert nous_wire.classify_upstream(_resp(id="gen-1788680469-hyks5Nj9n7rmdAbMyQ5x")) == "openrouter"
def test_gmi_by_anthropic_native_id_without_provider(self):
assert nous_wire.classify_upstream(_resp(id="msg_01XrffvmxsRUWwCbVpBsADHo")) == "gmi"
@pytest.mark.parametrize("r", [None, _resp(), _resp(id=None), _resp(id="chatcmpl-abc"), _resp(id="", provider="")])
def test_unknown_never_classifies(self, r):
assert nous_wire.classify_upstream(r) is None
class TestWireChoice:
def test_chat_for_openrouter_and_unknown(self):
assert nous_wire.wire_for_upstream("openrouter") == "chat_completions"
assert nous_wire.wire_for_upstream(None) == "chat_completions"
def test_gmi_is_chat_until_cleared_then_native(self, monkeypatch):
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", False)
assert nous_wire.wire_for_upstream("gmi") == "chat_completions"
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", True)
assert nous_wire.wire_for_upstream("gmi") == "anthropic_messages"
def _agent(**kw):
a = SimpleNamespace(provider="nous", model="anthropic/claude-fable-5.1", api_mode="chat_completions",
api_key="k", base_url="https://inference-api.nousresearch.com/v1", session_id="s")
for k, v in kw.items():
setattr(a, k, v)
return a
class TestHook:
@pytest.fixture(autouse=True)
def _auto(self, monkeypatch):
monkeypatch.setattr(_providers, "_nous_anthropic_wire", lambda: "auto")
self.switches = []
monkeypatch.setattr("agent.agent_runtime_helpers.switch_model",
lambda agent, m, p, api_key="", base_url="", api_mode="", **k: self.switches.append(api_mode) or setattr(agent, "api_mode", api_mode))
def test_gmi_cleared_schedules_once_and_applies_at_next_iteration(self, monkeypatch):
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", True)
a = _agent()
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(id="msg_01abc"), 1) is True
# decided but NOT switched yet: the response is still being consumed on the old wire
assert a.api_mode == "chat_completions" and self.switches == []
assert nous_wire.apply_pending_wire_switch(a) is True
assert a.api_mode == "anthropic_messages" and self.switches == ["anthropic_messages"]
assert nous_wire.apply_pending_wire_switch(a) is False # nothing pending twice
# a later call, even with a different-looking response, never flips again
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(provider="Anthropic", id="gen-1-x"), 2) is False
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(provider="Anthropic", id="gen-1-x"), 1) is False
assert self.switches == ["anthropic_messages"]
def test_openrouter_stays_on_chat(self):
a = _agent()
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(provider="Anthropic", id="gen-1-x"), 1) is False
assert a.api_mode == "chat_completions" and self.switches == []
def test_gmi_uncleared_stays_on_chat(self, monkeypatch):
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", False)
a = _agent()
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(id="msg_01abc"), 1) is False
assert a.api_mode == "chat_completions" and self.switches == []
@pytest.mark.parametrize("mode", ["chat", "native"])
def test_explicit_modes_never_auto_switch(self, monkeypatch, mode):
monkeypatch.setattr(_providers, "_nous_anthropic_wire", lambda: mode)
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", True)
a = _agent(api_mode="chat_completions" if mode == "chat" else "anthropic_messages")
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(id="msg_01abc"), 1) is False
assert self.switches == []
def test_other_providers_and_models_untouched(self, monkeypatch):
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", True)
assert nous_wire.maybe_switch_wire_after_first_response(_agent(provider="openrouter"), _resp(id="msg_01abc"), 1) is False
assert nous_wire.maybe_switch_wire_after_first_response(_agent(model="openai/gpt-5.6-sol"), _resp(id="msg_01abc"), 1) is False
assert self.switches == []
def test_switch_failure_is_swallowed_and_decision_is_final(self, monkeypatch):
monkeypatch.setattr(nous_wire, "GMI_NATIVE_WIRE_CLEARED", True)
def boom(*a, **k):
raise RuntimeError("no client")
monkeypatch.setattr("agent.agent_runtime_helpers.switch_model", boom)
a = _agent()
assert nous_wire.maybe_switch_wire_after_first_response(a, _resp(id="msg_01abc"), 1) is True
assert nous_wire.apply_pending_wire_switch(a) is False
assert a.api_mode == "chat_completions" and a._nous_wire_decided is True and a._nous_wire_pending is None
def test_real_agent_usage_recorder_calls_the_hook_once(tmp_path, monkeypatch):
"""The wiring: record_response_usage on a real AIAgent invokes the hook on call 1 only."""
monkeypatch.setenv("HERMES_HOME", str(tmp_path))
(tmp_path / "config.yaml").write_text("nous:\n anthropic_wire: auto\n", encoding="utf-8")
from run_agent import AIAgent
from agent import turn_usage
calls = []
monkeypatch.setattr(nous_wire, "maybe_switch_wire_after_first_response",
lambda agent, response, n: calls.append((n, nous_wire.classify_upstream(response))) or False)
a = AIAgent(api_key="jwt", base_url="https://inference-api.nousresearch.com/v1", provider="nous",
api_mode="chat_completions", model="anthropic/claude-fable-5.1", session_id="t", platform="cli",
quiet_mode=True, skip_context_files=True, skip_memory=True, save_trajectories=False, enabled_toolsets=["file"])
try:
usage = SimpleNamespace(prompt_tokens=100, completion_tokens=5, total_tokens=105, prompt_tokens_details=None, completion_tokens_details=None)
for n in (1, 2, 3):
resp = SimpleNamespace(usage=usage, id="gen-1788708403-xXjDsFwabc", provider="Anthropic", model="anthropic/claude-fable-5.1")
turn_usage.record_response_usage(a, resp, messages=[{"role": "user", "content": "hi"}], api_call_count=n,
api_duration=0.1, compression_attempts=0, max_compression_attempts=3)
finally:
a.close()
assert calls == [(1, "openrouter")]

View File

@@ -240,10 +240,12 @@ Nous Portal serves its `anthropic/*` models on two routes: OpenAI-compatible `/v
```yaml
nous:
anthropic_wire: chat # default. "native" = the Anthropic Messages wire
anthropic_wire: chat # default. "native" = the Anthropic Messages wire; "auto" = decide per session
```
`chat` is the default for now. The native wire is the better transport (signed thinking blocks pass through unchanged, native `cache_control` scopes), but in concurrent tool loops it currently re-writes the previous turn's prompt cache on 14–20% of consecutive calls, which is 15–20% of a fan-out's cache-write bill; the chat route measured 0 on the same test. Set `native` to opt back in (for example once the portal-side fix has shipped). Only `anthropic/*` models are affected; everything else on Nous already uses chat/completions.
`chat` is the default for now. The native wire is the better transport (signed thinking blocks pass through unchanged, native `cache_control` scopes), but on the Portal's OpenRouter-served path it currently re-writes the previous turn's prompt cache on 14–20% of consecutive calls in concurrent tool loops, which is 15–20% of a fan-out's cache-write bill; the chat route measured 0 on the same test. Set `native` to opt back in (for example once the portal-side fix has shipped). Only `anthropic/*` models are affected; everything else on Nous already uses chat/completions.
`auto` is for when the Portal serves the same model from more than one upstream. A session starts on chat, Hermes reads which upstream answered the first call, and switches that session to native only when the upstream is one where native is known to be clean (the switch happens between calls, so no in-flight response and no warm cache is lost). Today no upstream is cleared, so `auto` behaves exactly like `chat`; it exists so the flip can be made from a measurement rather than a config change.
## When does it take effect?