Commit Graph

5 Commits

Author SHA1 Message Date
teknium1
1a2bf64aa7 fix(aux): honour a custom OpenCode-family entry's declared api_mode, like the main runtime
opencode_transport() re-derived the wire for every OpenCode-family custom
entry, while hermes_cli/runtime_provider_custom.py::_resolve_named_custom_runtime
only re-derives when the entry declares no api_mode. Aux now mirrors that gate,
so an entry pinned to chat_completions stays a plain chat client on both paths.

Test grid trimmed to [None, one wrong mode]; the bridge fixture drops its
api_mode (the override no longer applies to it) and a pinned entry asserts the
honoured mode instead.
2026-09-19 09:38:39 -07:00
teknium1
37f7f00323 fix(aux): route OpenCode auxiliary clients by the model's wire, not a persisted api_mode
`resolve_provider_client("opencode-go", model="gpt-5.6-luna")` built a plain
Chat Completions client, so every auxiliary call (compression, titles,
vision, MoA) on that model went to /chat/completions and OpenCode answered
500 — while the main conversation on the same provider+model worked, because
hermes_cli/runtime_provider re-derives api_mode per model from
opencode_model_api_mode() and the auxiliary path never did.

`agent/opencode_affinity.py::opencode_transport()` is the single per-model
(api_mode, base_url) decision for OpenCode relay targets (built-in families,
`opencode-go-*` custom entries, opencode.ai hosts). `_wrap_transport` — the
transport chokepoint every resolve branch ends in — and the named-custom
branch (whose entry may have persisted the api_mode of whichever model was
selected at save time) consult it, so Responses-only models get
CodexAuxiliaryClient and Anthropic-wire models land on /v1/messages with the
/v1-stripped relay URL. A stale task/provider-level api_mode is ignored for
these targets, exactly like the main runtime.

Live: fake relay — before: OpenAI client → POST /v1/chat/completions → 500;
after: CodexAuxiliaryClient → POST /v1/responses. Control: glm-5 stays a
plain chat client on both.

Fixes #98799
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-19 09:38:39 -07:00
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
139396995a feat(opencode): send x-opencode-session on every OpenCode request for backend affinity
OpenCode pins requests sharing an x-opencode-session value to one upstream
backend, which is what keeps its prompt cache warm across a conversation.
Hermes never sent it, so cache ratios on OpenCode traffic were poor.

- agent/opencode_affinity.py: single owner of the header — target detection
  (built-in zen/go/free, custom opencode-* providers, any opencode.ai URL)
  and the key (affinity scope → conversation root → session id, cron
  timestamp stripped), same resolution as OpenRouter/xAI affinity hints.
- build_api_kwargs: merged once after the per-mode builder, so
  chat_completions, codex_responses and anthropic_messages all carry it.
- auxiliary _build_call_kwargs: same key from the runtime-main session so
  compression/title/vision calls stay on the conversation's backend; the
  aux Codex and Anthropic adapters now forward extra_headers.

Closes #81584, #81832 (deepseek-v4-flash 400 without the header).
2026-09-02 21:46:36 -07:00