Commit Graph

11 Commits

Author SHA1 Message Date
teknium1
7f9e1453a3 chore: merge origin/main (resolve agent/opencode_affinity.py, tests/hermes_cli/test_web_server_idle_proof.py) 2026-09-19 10:56:19 -07:00
ericmaddox
6ca41cbe8b fix(opencode): send ephemeral x-opencode-session header on one-shot requests
OpenCode Go strictly requires an x-opencode-session header on all requests to route requests efficiently and avoid HTTP 400 MissingSessionID. In turn chats and parented auxiliary calls, the header is derived from the conversation context or session ID. For stateless one-shot requests (commit messages, summaries, unparented auxiliary tasks), opencode_session_headers now falls back to generating an ephemeral session ID so one-shot requests to OpenCode succeed.

Closes #105841
2026-09-19 10:06:00 -07:00
teknium1
1a2bf64aa7 fix(aux): honour a custom OpenCode-family entry's declared api_mode, like the main runtime
opencode_transport() re-derived the wire for every OpenCode-family custom
entry, while hermes_cli/runtime_provider_custom.py::_resolve_named_custom_runtime
only re-derives when the entry declares no api_mode. Aux now mirrors that gate,
so an entry pinned to chat_completions stays a plain chat client on both paths.

Test grid trimmed to [None, one wrong mode]; the bridge fixture drops its
api_mode (the override no longer applies to it) and a pinned entry asserts the
honoured mode instead.
2026-09-19 09:38:39 -07:00
teknium1
37f7f00323 fix(aux): route OpenCode auxiliary clients by the model's wire, not a persisted api_mode
`resolve_provider_client("opencode-go", model="gpt-5.6-luna")` built a plain
Chat Completions client, so every auxiliary call (compression, titles,
vision, MoA) on that model went to /chat/completions and OpenCode answered
500 — while the main conversation on the same provider+model worked, because
hermes_cli/runtime_provider re-derives api_mode per model from
opencode_model_api_mode() and the auxiliary path never did.

`agent/opencode_affinity.py::opencode_transport()` is the single per-model
(api_mode, base_url) decision for OpenCode relay targets (built-in families,
`opencode-go-*` custom entries, opencode.ai hosts). `_wrap_transport` — the
transport chokepoint every resolve branch ends in — and the named-custom
branch (whose entry may have persisted the api_mode of whichever model was
selected at save time) consult it, so Responses-only models get
CodexAuxiliaryClient and Anthropic-wire models land on /v1/messages with the
/v1-stripped relay URL. A stale task/provider-level api_mode is ignored for
these targets, exactly like the main runtime.

Live: fake relay — before: OpenAI client → POST /v1/chat/completions → 500;
after: CodexAuxiliaryClient → POST /v1/responses. Control: glm-5 stays a
plain chat client on both.

Fixes #98799
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-19 09:38:39 -07:00
teknium1
11f44afbbd chore: stack on #115830 to resolve agent/opencode_affinity.py conflict
opencode_session_headers derives the key via resolve_affinity_key() (#115904) and keeps the
#115830 oneshot-<uuid> fallback so an OpenCode target never gets {} (OpenCode Go 400 MissingSessionID).
2026-09-19 03:36:23 -07:00
teknium1
2828e19a02 feat: slim the per-provider session header to the route seam and document it
Trim of the cherry-picked #104566 so it clears the salvage bar and covers #86241:

- The config getter matches entries the way every sibling per-provider knob does
  (`_entries_for_route`, route identity only) instead of a second provider-name axis;
  returns "" like `get_custom_provider_extra_headers` returns {}.
- `merge_opencode_session_headers` becomes `merge_session_affinity_headers` at both
  call sites (main `build_api_kwargs`, auxiliary `_build_call_kwargs`); the alias line
  is gone. Both sources merge (an OpenCode target that also declares a header gets both).
- Tests cut from six to two invariants: configured header carries one value per
  conversation on chat_completions, anthropic_messages and auxiliary kwargs (different
  for another session, caller-pinned wins); unconfigured → no header on any path.
- Docs: `configuring-models.md` per-provider options, `providers.md` entry key list,
  `cli-config.yaml.example` — the key is opt-in, default off, so DEFAULT_CONFIG is unchanged.

Why: a session-aware proxy classifies a request with no session id whose last message is a
tool_result as a NEW conversation and re-sends the whole history upstream (cache_write ≈
cache_read). Hermes already derives a rotation-stable conversation key for OpenCode; naming
the header per provider lets any proxy receive it without shipping an identifier by default.

Co-authored-by: 0xAlyDev <agentai891@gmail.com>
2026-09-19 01:02:31 -07:00
0xAlyDev
ece07d24fa feat(agent): per-provider session_affinity_header (#104449) 2026-09-19 00:55:34 -07:00
ericmaddox
709cc2214c fix(opencode): send ephemeral x-opencode-session header on one-shot requests
OpenCode Go strictly requires an x-opencode-session header on all requests to route requests efficiently and avoid HTTP 400 MissingSessionID. In turn chats and parented auxiliary calls, the header is derived from the conversation context or session ID. For stateless one-shot requests (commit messages, summaries, unparented auxiliary tasks), opencode_session_headers now falls back to generating an ephemeral session ID so one-shot requests to OpenCode succeed.

Closes #105841
2026-09-19 00:04:07 -07:00
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
139396995a feat(opencode): send x-opencode-session on every OpenCode request for backend affinity
OpenCode pins requests sharing an x-opencode-session value to one upstream
backend, which is what keeps its prompt cache warm across a conversation.
Hermes never sent it, so cache ratios on OpenCode traffic were poor.

- agent/opencode_affinity.py: single owner of the header — target detection
  (built-in zen/go/free, custom opencode-* providers, any opencode.ai URL)
  and the key (affinity scope → conversation root → session id, cron
  timestamp stripped), same resolution as OpenRouter/xAI affinity hints.
- build_api_kwargs: merged once after the per-mode builder, so
  chat_completions, codex_responses and anthropic_messages all carry it.
- auxiliary _build_call_kwargs: same key from the runtime-main session so
  compression/title/vision calls stay on the conversation's backend; the
  aux Codex and Anthropic adapters now forward extra_headers.

Closes #81584, #81832 (deepseek-v4-flash 400 without the header).
2026-09-02 21:46:36 -07:00