171 Commits

Author SHA1 Message Date
Robin Fernandes
749220ef00 feat(web): serve managed search through Perplexity
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.

The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.

Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
2026-09-24 16:17:16 -04:00
teknium1
9d799e0531 docs: hindsight installs from the plugin catalog
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
2026-09-23 01:16:41 -07:00
teknium1
4b8a813400 docs(portal): relative profiles link so the docs link check passes on main
f50d4b3423 (a docs revert) restored a route-style /user-guide/profiles link in
nous-portal.md (en + zh-Hans); website/scripts/check_doc_links.py rejects those, so
the Docs Site check has been red on main since. Rewritten with the script's --fix.
2026-09-21 10:51:01 -07:00
teknium1
f50d4b3423 Revert "docs(portal): the shared token store refreshes a login, it does not seed one"
This reverts commit 0269ddb7ca.
2026-09-21 09:55:40 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
teknium1
21ef1b97f9 fix(context): proxied Codex routes resolve the Codex OAuth window, not the direct-API catalog
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.

Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.

Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).

Fixes #116191
2026-09-19 12:17:05 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
Teknium
d6d9fac694 Merge pull request #115904 from NousResearch/fix/boa-res-R3-responses-reasoning-session-header
feat(providers): custom providers can carry the Hermes session id as an opt-in header (#86241, salvage #104566, supersedes #114496)
2026-09-19 11:14:31 -07:00
Teknium
699d46baea Merge pull request #115827 from NousResearch/fix/boa-vision-image-openai-opencodevision
fix(providers): opencode-go vision models attach images natively and resumed sessions stay on the Go endpoint (#96066, salvage #96116)
2026-09-19 11:09:37 -07:00
teknium1
7f9e1453a3 chore: merge origin/main (resolve agent/opencode_affinity.py, tests/hermes_cli/test_web_server_idle_proof.py) 2026-09-19 10:56:19 -07:00
teknium1
7aff381943 chore: merge origin/main (resolve tests/agent/test_output_cap_parsing.py, tests/hermes_cli/test_runtime_provider_resolution.py) 2026-09-19 10:52:01 -07:00
teknium1
e03af55edd chore: merge origin/main (resolve website/docs/integrations/providers.md) 2026-09-19 10:47:48 -07:00
teknium1
8e180c69f6 feat: label a model.context_length pin and warn once when it disagrees with the provider
model.context_length short-circuits get_model_context_length ahead of every
provider source, but nothing told the user the number they saw was their own
pin rather than provider metadata (#66168). Keep the pin semantics (custom
endpoints depend on it) and make it visible instead:

- agent/context_pin.py: is_context_pinned / context_pin_suffix for renderers,
  and warn_once_on_pin_disagreement, called from _resolve_context_length at
  agent init. The advertised value comes from LOCAL sources only (endpoint
  cache, models.dev disk cache, hardcoded catalog) so the check never adds a
  network probe or latency to startup; one warning per (model, pin) per process.
- "(pinned)" label on every CLI surface that renders the window: welcome
  banner, /model switch summary, /usage current-context line, wide status bar.
- No new config key or env var; the numeric value used is unchanged.

Fixes #66168
2026-09-19 10:28:34 -07:00
teknium1
55db0a9847 fix: local sub-64K context refusal stops assuming Ollama
When an OpenAI-compatible local server (llama.cpp, vLLM, ...) serves a
window below MINIMUM_CONTEXT_LENGTH, the startup refusal, the CLI banner
warning and the num_ctx log line were worded for Ollama (/api/show,
model.context_length advice that cannot raise a served window). Name the
served window and the server-agnostic remedies: start the server with a
>=64K context or set model.ollama_num_ctx (honoured on every local
endpoint) to the window it really serves. Hosted routes keep the
model.context_length advice. The floor itself is unchanged.

Part of #87075
2026-09-19 10:21:28 -07:00
teknium1
bb787f45f3 docs: note that session-less one-shots send an ephemeral x-opencode-session key
The providers page listed every OpenCode caller that carries the affinity header; add the
stateless one-shot case (Desktop commit-message generation with no active session), which
now sends a fresh ephemeral key instead of nothing so the relay no longer rejects it with
400 MissingSessionID (#105841).
2026-09-19 10:06:00 -07:00
teknium1
37f7f00323 fix(aux): route OpenCode auxiliary clients by the model's wire, not a persisted api_mode
`resolve_provider_client("opencode-go", model="gpt-5.6-luna")` built a plain
Chat Completions client, so every auxiliary call (compression, titles,
vision, MoA) on that model went to /chat/completions and OpenCode answered
500 — while the main conversation on the same provider+model worked, because
hermes_cli/runtime_provider re-derives api_mode per model from
opencode_model_api_mode() and the auxiliary path never did.

`agent/opencode_affinity.py::opencode_transport()` is the single per-model
(api_mode, base_url) decision for OpenCode relay targets (built-in families,
`opencode-go-*` custom entries, opencode.ai hosts). `_wrap_transport` — the
transport chokepoint every resolve branch ends in — and the named-custom
branch (whose entry may have persisted the api_mode of whichever model was
selected at save time) consult it, so Responses-only models get
CodexAuxiliaryClient and Anthropic-wire models land on /v1/messages with the
/v1-stripped relay URL. A stale task/provider-level api_mode is ignored for
these targets, exactly like the main runtime.

Live: fake relay — before: OpenAI client → POST /v1/chat/completions → 500;
after: CodexAuxiliaryClient → POST /v1/responses. Control: glm-5 stays a
plain chat client on both.

Fixes #98799
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-09-19 09:38:39 -07:00
teknium1
8f1542422b fix: fallback entries inherit a named provider's declared transport
A fallback_providers entry naming a `providers.<name>` block (or
`custom:<name>`) without its own api_mode was re-detected from the
resolved host: an Anthropic-Messages proxy on a plain host or a
Responses-only relay behind a generic gateway landed on chat_completions
while resolve_provider_client had already built the declared client
(#33062 bottom thread, #81932 provider-level `transport:`). The hint pass
now reads the named block's api_mode/transport as explicit, and entry-level
`transport:` is accepted as an alias of `api_mode` with the same
canonicalisation the providers block uses (`responses` -> codex_responses).

Fixes #33062
Fixes #81932
2026-09-19 09:27:08 -07:00
teknium1
11f44afbbd chore: stack on #115830 to resolve agent/opencode_affinity.py conflict
opencode_session_headers derives the key via resolve_affinity_key() (#115904) and keeps the
#115830 oneshot-<uuid> fallback so an OpenCode target never gets {} (OpenCode Go 400 MissingSessionID).
2026-09-19 03:36:23 -07:00
teknium1
5eb31233b4 fix: parse the "limited to N" output-cap figure so the budget steps down
is_output_cap_error already classified "max_completion_tokens is limited to
16384 for glm-5.2" as an output cap, but parse_available_output_tokens_from_error
returned None, so _recover_context_length failed fast instead of clamping.
Add the ("limited to",) signal and the direct-cap regex; the Scaleway string
now yields 16384. Also document model.key_env in the custom provider block.
2026-09-19 01:15:42 -07:00
teknium1
2828e19a02 feat: slim the per-provider session header to the route seam and document it
Trim of the cherry-picked #104566 so it clears the salvage bar and covers #86241:

- The config getter matches entries the way every sibling per-provider knob does
  (`_entries_for_route`, route identity only) instead of a second provider-name axis;
  returns "" like `get_custom_provider_extra_headers` returns {}.
- `merge_opencode_session_headers` becomes `merge_session_affinity_headers` at both
  call sites (main `build_api_kwargs`, auxiliary `_build_call_kwargs`); the alias line
  is gone. Both sources merge (an OpenCode target that also declares a header gets both).
- Tests cut from six to two invariants: configured header carries one value per
  conversation on chat_completions, anthropic_messages and auxiliary kwargs (different
  for another session, caller-pinned wins); unconfigured → no header on any path.
- Docs: `configuring-models.md` per-provider options, `providers.md` entry key list,
  `cli-config.yaml.example` — the key is opt-in, default off, so DEFAULT_CONFIG is unchanged.

Why: a session-aware proxy classifies a request with no session id whose last message is a
tool_result as a NEW conversation and re-sends the whole history upstream (cache_write ≈
cache_read). Hermes already derives a rotation-stable conversation key for OpenCode; naming
the header per provider lets any proxy receive it without shipping an identifier by default.

Co-authored-by: 0xAlyDev <agentai891@gmail.com>
2026-09-19 01:02:31 -07:00
teknium1
ba9d68a838 fix(cli): a resumed opencode-go session re-derives its wire format from the stored model
The classic CLI (`hermes --resume`, mid-chat `/resume`, oneshot resume) read the
session row's persisted api_mode/base_url verbatim in stored_session_route(), so
a row written while the session ran an anthropic_messages model on opencode-go
(MiniMax) pinned that wire onto a chat_completions model such as
deepseek-v4-flash-vision-exp after the model column moved. This is the CLI twin
of the tui_gateway _rederive_per_model_route() fix already on this branch: for
providers that pick the wire per model (model_derived_api_mode() is not None)
the route follows the stored model and the relay URL is healed; fixed-wire
providers keep honoring their row.

Probe (direct call of _restore_session_model on the PR head, no network): a row
with provider opencode-go, base_url '' and api_mode anthropic_messages for
deepseek-v4-flash-vision-exp resumed on the Go relay (never api.anthropic.com —
the empty base_url is re-resolved to https://opencode.ai/zen/go/v1 on both the
same-provider and provider-changed paths) but with api_mode anthropic_messages;
after: chat_completions on the same relay.

Docs: the OpenCode paragraph in providers.md now states both behaviours users
can observe from #96066 — per-model routing survives resume, and `*-vision*`
OpenCode ids attach images natively without a supports_vision override.

Part of #96066
2026-09-19 00:45:43 -07:00
teknium1
30c8bbe562 docs: note that session-less one-shots send an ephemeral x-opencode-session key
The providers page listed every OpenCode caller that carries the affinity header; add the
stateless one-shot case (Desktop commit-message generation with no active session), which
now sends a fresh ephemeral key instead of nothing so the relay no longer rejects it with
400 MissingSessionID (#105841).
2026-09-19 00:06:55 -07:00
teknium1
0cd13aaf74 docs: relative link to the borrowed-logins section (route-style links fail the docs check) 2026-09-18 20:55:24 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
b9d3c70303 docs(portal): name the exact per-profile sign-in path and fix the guide page too
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.

Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).

guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
2026-09-18 10:28:04 -07:00
PRATHAMESH75
0269ddb7ca docs(portal): the shared token store refreshes a login, it does not seed one
The Portal "Profile setup" section promised that every profile picks up a
Portal login automatically. The shared token store only merges fresher
tokens into a profile that already has a Nous state; a profile with no
providers.nous entry of its own fails closed at boot and is asked to run
the setup flow (hermes_cli/auth.py:1545, auth_nous.py:1056), because
profiles are independent islands (#111724). Correct the promise to
"sign in once per profile" and describe the fail-closed behavior, in both
the English and zh-Hans copies.

Fixes #114259
2026-09-18 10:28:04 -07:00
teknium1
d6a753fe91 fix(agent): aux Responses adapter emits the main transport's tool schemas and aliases
_CodexCompletionsAdapter._build_responses_kwargs (title generation, compression,
MoA aggregation) built its own chat->Responses tool list: raw names, no
``strict: False``. On api.perplexity.ai and opencode.ai routes the main loop
already aliased reserved names (hermes_<name>), so every auxiliary call 400ed
while chat looked healthy and MoA silently fell back.

- tools go through codex_responses_adapter._responses_tools (strict: False) and
  agent/transports/codex.py::_alias_wire_tools — one reserved-name table for
  both paths (OpenCode, Perplexity, xAI)
- replayed history tool_calls are renamed to the alias this request declares
- the alias map rides on the request payload (``_wire_aliases``), popped in
  create() and mapped back onto the parsed tool_calls before dispatch; never
  instance state (aux adapters are cached and shared)

Ported from #114457's auxiliary half; design per the #114260 report.

Co-authored-by: Tyler Lyon <lyonrt@icloud.com>
2026-09-18 09:54:13 -07:00
teknium1
d7f2644141 fix(codex): profile declarations follow the host and never override OpenAI's own ladder
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes #114249 but has two
edges the transport must keep:

- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
  non-None custom declaration short-circuited the host lookup: a
  ``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
  Router catalog clamp and would send ``max`` to a gateway that 400s on it.
  Resolve by endpoint host first, then by name — the host is the endpoint's
  truth; the config-entry name is only a label (the existing Router test now
  uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
  codex_responses) is still OpenAI: its per-model ladder is known, so skip
  profile declarations on the official origin and the Codex backend.
  ``_is_openai_api_origin`` is the shared exact-host check;
  ``_is_official_openai_responses_route`` reuses it.

Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.

Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
2026-09-18 09:40:31 -07:00
teknium1
227384e332 fix(auth): drop dead code from the Codex login retry; trim to two invariant tests
Follow-up to the salvaged #114614 pick:

- `_codex_login_post`: the `for` loop ended in an unreachable `raise … # pragma: no cover`
  (needed only to satisfy the return type). A `while True` with the terminal condition
  folded into the except branch has no dead path and the same three-attempt bound.
- `_codex_poll_authorization_code`: the pick duplicated the `except KeyboardInterrupt`
  handler; the second copy was unreachable.
- Tests: six change-detector tests collapsed into two invariants (poll survives blips but
  never retries a non-transport exception; one-shot POST retries once and keeps the typed
  AuthError + TLS hint + cause chain at the cap). The existing
  test_codex_device_login_ssl_hint.py still pins the poll's terminal hint path.
- Docs: providers.md notes that a single dropped connection during device login is no
  longer fatal.
2026-09-18 09:18:00 -07:00
Ritesh Patel
9927199998 docs: drop opencode-free references (provider removed)
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
2026-09-18 15:40:37 +05:30
teknium1
931b5ff9e7 fix(opencode): every credential-resolution surface keys off the model it will send; family heal only for built-in providers
Closes the remaining atoms of #112600.

A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
   its siblings still resolved credentials against config's `default`: the CLI
   auth-fallback rung, `--resume` credential re-resolution, the gateway
   provider-override helper (channel overrides, persisted /model switches,
   API-server provider refresh), the gateway fallback chain, the TUI /model
   switch-from runtime and ACP agent construction. With a `*-free` default the
   OpenCode free-tier rung fired first and a Go-only model was built against
   the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
   the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
   optional `target_model` and the two test stubs of it accept the kwarg.

B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
   provider matched by opencode_provider_family, including custom providers
   merely named after a family (`opencode-go-bridge`, #85589) whose relay the
   user declared explicitly in `providers:`. The family heal now applies to the
   built-in canonical providers only; custom prefix-named providers keep their
   per-model api_mode routing and /v1 handling. Documented in the providers
   guide.

C) Same function: the official-host check uses parsed.hostname (a port no
   longer defeats the heal) and only the path is edited, so query/fragment
   round-trip instead of being dropped.

Fixes #112600
2026-09-17 08:59:22 -07:00
teknium1
d1c6786d1e fix(models): custom providers can inherit a vendor catalog; unknown models get no synthesized output cap
A metadata-only model_overrides entry on a custom provider (only context_window set) took the
unknown-model branch of get_model_capabilities and reported max_output_tokens=8192 from the
_UNKNOWN_MODEL_BASE template; /model info showed "Max output: 8,192 tokens" for a model nobody
knows. The output limit is now left unknown (ModelCapabilities.max_output_tokens=None,
ModelInfo.max_output=0), matching the fail-open treatment #113146 gave supports_vision and
supports_reasoning. Consumers already handle a missing value: the dashboard only renders a
positive limit and the /model card only prints a truthy max_output.

There was also no way for a custom provider that resells a catalogued vendor's models (a
gateway/proxy) to reach _BUILTIN_MODEL_METADATA or the models.dev catalog, which is what forced
the metadata-only override in the first place. providers.<name>.catalog_provider (also honoured
on legacy custom_providers[] rows) names that vendor; models_dev resolves the catalog id through
it for capabilities, context, model info and provider info. It affects metadata lookups only,
never routing or credentials, and an explicit model_overrides entry still wins.

Closes the remaining atoms of #112649 (docs: providers page).
2026-09-17 08:53:48 -07:00
teknium1
6f8389d67b fix(aux): keep the /btw and title snapshots duck-typed; make the out-of-turn header tests discriminating
Follow-up on the salvaged #112721 commits (@fangliquanflq):

- agent/turn_context.py, tui_gateway/methods_prompt.py: add ``session_id`` to the explicit
  snapshot dicts instead of switching them to ``agent._current_main_runtime()``. The titling
  prologue is duck-typed (its tests drive a minimal stub) and the snapshot values feed the
  ``runtime_validator`` equality checks, where ``_current_main_runtime()``'s ``"" `` for a
  missing attribute would no longer match the live ``None``. Same outcome — the background
  request inherits the conversation's ``x-opencode-session`` — without changing what the
  callers read.
- tests/agent/test_opencode_session_affinity.py: the salvaged title test passed on an
  unfixed tree because the titler thread republishes the conversation contextvar
  (``set_conversation_context``) and the affinity header falls back to it. Replace it with
  two invariant tests (sync + async ``call_llm(main_runtime=...)`` with EVERY ambient source
  unset via the ``out_of_turn`` fixture): red on origin/main, green here; both also pin
  that the explicit binding does not leak past the call.
- website/docs/integrations/providers.md: name the background/out-of-turn auxiliary calls
  the header now covers.

Fixes #112717
2026-09-16 17:22:17 -07:00
teknium1
d7c791c64c docs: metadata-only model_overrides leave vision/reasoning unknown
The config_defaults comment still described the unknown-model template as
"vision/reasoning off", which is exactly the synthesized-False behaviour
#112649 reports and the preceding commit removes. State the new contract
where users read it (config_defaults, providers docs): a context_window-only
override for an uncatalogued model keeps supports_vision / supports_reasoning
unknown so vision_analyze, video_analyze and the reasoning-effort picker stay
available; only an explicit false in the override marks the model text-only.
2026-09-16 17:09:50 -07:00
teknium1
f7ea39481a fix(kanban): dashboard estimate calls declare a relay-affinity key too
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.

Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
2026-09-15 18:33:43 -07:00
teknium1
1a8d922003 docs(providers): note the per-task x-opencode-session key for headless Kanban aux calls 2026-09-15 18:33:43 -07:00
teknium1
b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
Teknium
0c2e66ea8b test(auth): OpenRouter PKCE invariants, fake-authority A/B harness, docs, contributor map
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
2026-09-12 22:07:41 -07:00
Justin Bennington
8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington
d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
teknium1
d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
KoNit-K
ef903617ea docs(providers): note the OpenSSL 3.5 PQ-group middlebox failure on Codex device login
Docs hunk carried from PR #106389 (hermes_cli changes superseded by #106394's
smaller equivalent); the reporter's exact openssl.cnf snippet follows in a
maintainer commit.

Refs #106384
(cherry picked from commit abb323d86ced829859477ed72d59f409c8d8b895, docs hunk only)
2026-09-09 10:14:58 -07:00
Teknium
bf53ff00a7 fix(config): one bounded backups/config/ dir replaces four config.yaml.bak schemes
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.

hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.

Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
2026-09-09 02:36:00 -07:00
Teknium
520e63661c fix: keep command-auth model discovery lazy across config and setup 2026-09-07 21:22:49 -07:00
Teknium
bbbccd3935 fix: failed probe credentials cannot fall through to configured auth 2026-09-07 08:08:04 -07:00
Teknium
7d44fe9c74 fix: capability probes send minted credentials instead of callable representations
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:08:04 -07:00
Teknium
bdf45abd7c fix: prefer owned Anthropic grants and bind auxiliary refresh to request credentials
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>

Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
2026-09-07 08:07:26 -07:00