opencode_transport() re-derived the wire for every OpenCode-family custom
entry, while hermes_cli/runtime_provider_custom.py::_resolve_named_custom_runtime
only re-derives when the entry declares no api_mode. Aux now mirrors that gate,
so an entry pinned to chat_completions stays a plain chat client on both paths.
Test grid trimmed to [None, one wrong mode]; the bridge fixture drops its
api_mode (the override no longer applies to it) and a pinned entry asserts the
honoured mode instead.
`resolve_provider_client("opencode-go", model="gpt-5.6-luna")` built a plain
Chat Completions client, so every auxiliary call (compression, titles,
vision, MoA) on that model went to /chat/completions and OpenCode answered
500 — while the main conversation on the same provider+model worked, because
hermes_cli/runtime_provider re-derives api_mode per model from
opencode_model_api_mode() and the auxiliary path never did.
`agent/opencode_affinity.py::opencode_transport()` is the single per-model
(api_mode, base_url) decision for OpenCode relay targets (built-in families,
`opencode-go-*` custom entries, opencode.ai hosts). `_wrap_transport` — the
transport chokepoint every resolve branch ends in — and the named-custom
branch (whose entry may have persisted the api_mode of whichever model was
selected at save time) consult it, so Responses-only models get
CodexAuxiliaryClient and Anthropic-wire models land on /v1/messages with the
/v1-stripped relay URL. A stale task/provider-level api_mode is ignored for
these targets, exactly like the main runtime.
Live: fake relay — before: OpenAI client → POST /v1/chat/completions → 500;
after: CodexAuxiliaryClient → POST /v1/responses. Control: glm-5 stays a
plain chat client on both.
Fixes#98799
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>