Classes: C9 provider wire-format drift, C17 prompt-cache hits, C11 real
auth / credential routing / `/models` parse.
Mocks encode our own belief about each vendor's wire format; vendor-side
schema changes, reasoning-replay rules, streaming shape changes and cache
behaviour are only observable against the real APIs. This drives the REAL
AIAgent + real adapters (chat_completions, anthropic_messages via OpenRouter,
codex_responses for xAI/OpenAI, GeminiNativeClient) in a temp HERMES_HOME
through a scripted 3-turn conversation with a deterministic registered tool:
turn 1 one forced tool call; valid JSON args; result round-trips
turn 2 two tool calls in ONE assistant message; both results round-trip
turn 3 no tools; answer from replayed history (tool results + reasoning
replay) under a cache breakpoint
Invariants per turn: no 4xx on any agent-loop request (429 excepted), the
turn completes, no <think>/DSML/control-token/raw tool JSON in user text,
the credential resolved from env is the one sent and ONLY to the provider's
host, tool results are persisted to state.db. Anthropic-family routes also
assert 1..4 well-formed breakpoints (none on role:tool / inside
tool_result.content[]) and cache-read tokens > 0 on turn 3 (warm once).
A `/models` listing check per provider uses the live fetchers without the
curated fallback (Gemini is strict-xfail on open #62259).
Spend: cheap models, <= 3 turns, max_tokens 2048, max_iterations 6, a hard
per-test token and $ guard; usage + estimated $ printed per test and
appended to $HERMES_LIVE_USAGE_FILE. Every case skips cleanly without its
key. `live` is excluded by default addopts; select with `-m live`.
(cherry picked from commit 3bd7de6991bcb03bfe5d11945de6f2317c534b7d)