Commit Graph

514 Commits

Author SHA1 Message Date
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
b04de1202a fix(models): filter retired Zen ids on every picker path, not only the live merge
The offline (no key) path serves the curated floor merged with models.dev, and both still list
x-preview-f-free; the setup-flow path goes through merge_profile_catalog with an empty live list.
One helper, _drop_delisted_opencode_models, runs on the FINAL rows of merge_profile_catalog and of
the static/models.dev return in provider_model_ids, so no path can offer a slug the relay 401s.
Adds the offline invariant test.
2026-09-20 10:14:01 -07:00
KeyArgo
4541fd5561 fix(models): delist retired x-preview-f-free from the OpenCode Zen picker
The Zen relay retired x-preview-f-free (the picker-facing id for Ox Alpha),
but the exclusion set named the wrong slug (ox-alpha-free) and the live-first
merge only filtered the live half, so the curated floor resurrected the
delisted id and the picker kept offering a model that 401s.

- add x-preview-f-free to _OPENCODE_FREE_EXCLUDED_MODELS
- filter the MERGED result (curated floor is merged back in as secondary half)

Closes #115496
2026-09-20 10:14:01 -07:00
kshitijk4poor
b916b3c130 refactor(models): cache-only OpenRouter read parses the disk copy once; test nits 2026-09-20 15:17:14 +05:30
kshitijk4poor
4243abd633 fix(gateway): /model listing also skips the OpenRouter catalog GET and saved-endpoint /models probes
Gate review: two live sockets survived the first commit on the chat /model reply path —
fetch_openrouter_models() re-downloaded the curated catalog once the disk copy passed its
TTL, and probe_custom_providers defaulted to True so every saved custom endpoint with a key
was probed. fetch_openrouter_models gains cache_only (memory → stale disk → in-repo snapshot,
never a socket); list_picker_providers forwards cache_only and the probe_* flags; the gateway
passes the same read-path flags the GUI picker uses. Tests fold into the existing /model
harness files and record every live probe seam instead of only the prefetch entry.
2026-09-20 15:17:14 +05:30
kshitijk4poor
6cfbc5f891 refactor(credits): throttle the re-warm, scope the seed thread, one subscription predicate
- rewarm_pricing_before_depleted_notice: a failed fetch caches {} for
  _FAILED_CATALOG_TTL_SECONDS and the peek reads that as cold, so every
  header in that window spawned a thread that read the auth store and hit
  the cached {}. Remember when the last warm started and decide inline
  until the window passes. Drop the dead try/except around the pure peek.
- _bg_seed now runs under spawn_context_thread: the warm it gained reads
  the profile's auth store, so the thread must carry the profile scope.
- _rerun_notice_policy replaces the idiom copied at three sites.
- _is_subscription_billed: the free-tier default filtered on any truthy
  billing_mode while _is_model_free keyed on == 'subscription'.
- The no-respawn guard test counts warm calls instead of enumerating
  finished threads (which always read 0).
2026-09-20 14:02:17 +05:30
Robin Fernandes
a76e491660 feat(nous): unlock subscription-billed models for free-tier accounts
The Nous gateway can bill a catalog row to a subscription the account
holds instead of to credits, and marks such rows with
`billing_mode: "subscription"` on GET /v1/models. A free-tier account
can run them, but the picker locked every row not priced at $0.

Carry the marker into the Nous pricing entry and count it in
`_is_model_free`, which already feeds the tier partition and the
credits-depleted notice. The silent default for a free-tier account
still prefers a genuinely free model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 14:02:17 +05:30
kshitijk4poor
7fec7cf594 fix(setup): profile-owned catalog keeps the curated list when the live fetch fails
The profile branch added for #116408 returned ``merge_profile_catalog(...) or []``, so a
built-in API-key provider with a short curated row and a profile declaring no
fallback_models (minimax, minimax-cn, gemini, kilocode, arcee, stepfun, xiaomi) offered
zero models at first-time setup whenever the live catalog was unreachable — the generic
probe it replaced returned the curated list. Fall back to ``curated``, and share the
key-gated probe with the ``/model`` picker (``models.probe_profile_catalog``) so setup
neither sends a keyless request nor diverges from the picker's rows.
2026-09-20 12:17:52 +05:30
teknium1
e22e33c189 fix(cli): first-time setup resolves plugin catalogs like the /model picker
`_api_key_provider_model_list` now routes a registered profile through the
same merge the picker uses (`fetch_models()` curated-first with
`fallback_models`; `fallback_models` alone when the fetch returns None or
raises) via a shared `models.merge_profile_catalog`, extracted from
`_profile_live_catalog` so setup and switching cannot drift. The picker
path also treats a raising `fetch_models` override as an empty catalog
instead of dropping to `[]`.

Why: salvage #116437 returned the live list only when it was at least as
long as the curated one and let a raising catalog abort setup, so setup
and `/model` could still offer different rows for the same profile.

Part of #116408
2026-09-19 20:56:31 -07:00
Robin Fernandes
5fa01f2fe0 fix(models): reduce repeated Nous recommendation traffic 2026-09-19 20:55:20 -07:00
teknium1
22beb95427 fix(models): external_process catalog fetch degrades to fallback_models on error; test mirrors registry via public types
A raising fetch_models() on the external_process branch escaped to
provider_model_ids() outer try and lost the profile fallback_models, unlike the
api_key branch. The admission test mirrored through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead.
2026-09-19 20:54:06 -07:00
teknium1
6f12165a94 fix(models): admit every plugin provider to the picker by slug; catalogs fall back to the profile
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.

- _profile_live_catalog: external_process profiles use fetch_models(),
  then fallback_models; every other non-api-key profile returns its
  fallback_models instead of None (in-tree ones declare none, so the
  built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
  key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
  the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
  picker rows and the authenticated flag are derived.

Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
2026-09-19 20:54:06 -07:00
tobenwarrior
cc2ec588f2 fix(agent): key ACP launch kwargs on provider profile, not vendor slug
_explicit_client_kwargs hardcoded copilot-acp for command/args launch
kwargs; out-of-tree external_process plugin providers got no launch path
and failed at client construction. Key on the provider profile's
auth_type instead, so every ACP/subprocess provider launches the same
way (same approach as #111194, folded in with credit).

Also harden _profile_live_catalog: a signature-strict external_process
profile (fetch_models requiring keyword-only api_key/base_url) now falls
back to credential kwargs on TypeError instead of crashing discovery.

Tests proven red on base for both behaviors.
2026-09-19 20:54:06 -07:00
tobenwarrior
de9e45f508 fix(models): route external_process provider catalogs through profile.fetch_models
_profile_live_catalog gated on auth_type=='api_key', so even
picker-admitted ACP providers had no live catalog and fell to the
self-named single-model fallback. External-process profiles supply their
own catalog via fetch_models (subprocess-owned); route them through it.
Verified live: kiro-acp lists its real 19 models (claude-opus-5,
gpt-5.6-sol/terra/luna, deepseek-3.2, glm-5, ...) in model.options.
2026-09-19 20:54:06 -07:00
teknium1
ff4399a0d7 fix: route a slug shared by several catalogs to the provider the user can use
`detect_provider_for_model` took the FIRST static-catalog hit as the only
guess. `gpt-5.6-luna` (and the rest of the gpt-5.6 family) is listed by both
`openai-api` and `openai-codex`, so a user with a Codex OAuth grant and no
OPENAI_API_KEY was routed to a keyless openai-api on a fresh (`auto`)
session, or — after the credential gate — left on the current provider with
the request silently ignored, while the grant they hold was never
considered.

`_static_catalog_matches` now yields every catalog that lists the slug in
ladder order; `detect_provider_for_model` keeps its existing credential gate
and takes the first sibling the user actually has credentials for. A fresh
session with no usable provider anywhere still fails loudly on the first
guess, and a user holding both keys keeps today's openai-api routing.

Fixes #102775
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-19 10:22:01 -07:00
teknium1
9a403e52f4 fix(models): azure-foundry disk-cache fingerprint tracks model.base_url
The wizard writes only model.base_url; the fingerprint hashed the env vars alone, so
switching resource under the same key served the previous resource's catalog for the
TTL/SWR window. Fold the config-resolved base URL in, as openai does with effective_base.
Part of #27989.
2026-09-19 09:42:49 -07:00
xxxigm
1d62b6ac6d fix(models): the /model picker lists the configured Azure Foundry resource's models
`provider_model_ids("azure-foundry")` fell through to the static catalog
`_PROVIDER_MODELS["azure-foundry"] = []`, so `/model azure-foundry` showed "0 models"
even on resources with many deployments. The plugin profile ships `base_url=""`
(per-resource), which is exactly why the generic `_profile_live_catalog` never
fires for it.

Register an `azure-foundry` entry in `_PROVIDER_CATALOG_FETCHERS` that resolves
the endpoint and credential through the runtime resolver
(`_resolve_azure_foundry_runtime`: `model.base_url` / `AZURE_FOUNDRY_BASE_URL`,
API key or Entra token-provider callable) and reuses the wizard's
`azure_detect._probe_openai_models` (api-version fallbacks, never raises).
Anthropic-style `/anthropic` routes have no `/models`; the probe fails soft and
the picker keeps the static `[]`.

Reapplied onto the fetcher-table layout from #28006; reformat/bloat stripped.

Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
2026-09-19 09:42:49 -07:00
teknium1
29bc6343d3 fix(auth): read-only Codex reads take no store lock; /model picker never refreshes
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.

`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.

`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
2026-09-19 09:24:25 -07:00
kshitijk4poor
b2d2cf4e24 fix(cli): the startup route decodes custom:<name> from the caller's providers, not a second config read
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
2026-09-19 01:44:53 +05:30
teknium1
8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00
finn763
cdacc2bcf1 fix(picker): never wait on provider catalog probes in the model-options read path
Opening the desktop model picker could sit on skeleton placeholders for 70s+
because a normal open (refresh=False) ran live provider catalog probes inline:
the serial row pass fetched each stale provider's /v1/models itself, and the
parallel prefetch joined every worker, so one degraded provider (hanging
endpoint, failed auth probe) held the whole response.

A normal open is now a read path:

* cached_provider_model_ids(non_blocking=True) serves the same-credentials
  disk entry of any age and refreshes it in a daemon thread; a cold row
  returns [] so the row keeps its curated list.
* list_authenticated_providers(non_blocking_catalogs=True) skips the joining
  prefetch and reads every row cache-only; build_model_options_payload turns
  it on for refresh=False (api-server / dashboard / TUI model.options).
* rows whose catalog is still warming carry catalog_pending, so a GUI can
  tell "not resolved yet" from "that is the provider's catalog".
* Ollama Cloud's 8s probe becomes a cached read + background warm; the
  loopback LM Studio probe stays (1.5s, cannot be a degraded remote).
* The SWR write now takes the cache lock: the read path spawns one warm per
  stale provider, and concurrent load-modify-save dropped rows.

An explicit refresh (Refresh Models) still probes live and blocking.

(cherry picked from commit 0f2c2e7a3631f2c8f3b5140642306eaf2d688682)
2026-09-18 11:01:52 -07:00
kshitijk4poor
debfc7420b refactor(providers): finish the keyless cleanup
`get_api_key_provider_status` pre-populated configured/logged_in/base_url/
key_source purely so the deleted keyless short-circuit could return them;
every one is now recomputed before the return, so build the snapshot once
instead of overwriting four dead placeholders. Same keys, same order, same
values.

The `_OPENCODE_FREE_EXCLUDED_MODELS` comment also still implied an
anonymous path that no longer exists; it now names the two ids it holds
and why.
2026-09-18 15:40:37 +05:30
kshitijk4poor
66a29c3e9d docs: correct the free-tier narrative the removal left behind
The keyless free tier is gone, but five comments and two test docstrings
still described its routing rung: the resolve_runtime_provider ladder
docstring listed a step that no longer exists, and both target_model call
sites plus their regression tests explained themselves in terms of a
`*-free` default being routed to the keyless Zen relay. They now state
what the code actually does (the model-keyed rungs pick the relay and
api_mode).

Also drops the two comments that only narrated the removal
(_OPENCODE_FREE_EXCLUDED_MODELS' history, and an orphan note in
auxiliary_client._resolve_api_key_branch).
2026-09-18 15:40:37 +05:30
kshitijk4poor
b388a48a2e refactor(providers): drop the now-dead keyless provider plumbing
The keyless OpenCode free tier was the only provider that ever set
`HermesOverlay.keyless`, so the flag and everything keyed off it is now
unreachable: the `keyless=` field on `HermesOverlay` and
`ProviderDescriptor`, the `_overlay_has_creds` early return, both
`_provider_is_keyless` copies (auth.py and inventory.py), the
`get_api_key_provider_status` keyless short-circuit and its
`key_source: "keyless"` placeholder, and the empty
`_KEYLESS_STABLE_CACHE_PROVIDERS` set whose `_credential_fingerprint`
branch could never match.

Dropping the two catalog-derived test exemptions follows: they computed
the empty set.
2026-09-18 15:40:37 +05:30
Ritesh Patel
ce6c45d335 fix(models): restore _OPENCODE_FREE_EXCLUDED_MODELS used by the Zen/Go live pickers
The removal dropped the constant but _profile_live_catalog() still filters
live-first Zen/Go listings through it (relay advertises delisted *-free slugs
that 400/403 on POST). Without it, any keyed live picker hit a NameError.
Restored with the original contents (ox-alpha-free, deepseek-v4-flash-free)
and an updated comment; extension point kept for future relay delistings.
2026-09-18 15:40:37 +05:30
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
teknium1
931b5ff9e7 fix(opencode): every credential-resolution surface keys off the model it will send; family heal only for built-in providers
Closes the remaining atoms of #112600.

A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
   its siblings still resolved credentials against config's `default`: the CLI
   auth-fallback rung, `--resume` credential re-resolution, the gateway
   provider-override helper (channel overrides, persisted /model switches,
   API-server provider refresh), the gateway fallback chain, the TUI /model
   switch-from runtime and ACP agent construction. With a `*-free` default the
   OpenCode free-tier rung fired first and a Go-only model was built against
   the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
   the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
   optional `target_model` and the two test stubs of it accept the kwarg.

B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
   provider matched by opencode_provider_family, including custom providers
   merely named after a family (`opencode-go-bridge`, #85589) whose relay the
   user declared explicitly in `providers:`. The family heal now applies to the
   built-in canonical providers only; custom prefix-named providers keep their
   per-model api_mode routing and /v1 handling. Documented in the providers
   guide.

C) Same function: the official-host check uses parsed.hostname (a port no
   longer defeats the heal) and only the path is edited, so query/fragment
   round-trip instead of being dropped.

Fixes #112600
2026-09-17 08:59:22 -07:00
teknium1
d6d6565d90 fix(picker): native Ollama cache rows use the 300s TTL and never stale-serve empty catalogs
Routing the native /api/tags probe through cached_fetch_api_models made the CURRENT endpoint's
probe cache-first with the generic 1h TTL (+7d SWR), so a model pulled after the first picker
open stayed invisible for up to an hour across restarts; the built-in `ollama` slug clamps to
_OLLAMA_LOCAL_MODELS_CACHE_TTL in cached_provider_model_ids but this path did not. Pass the
300s native TTL for the native admission.

An authoritative EMPTY native catalog was also persisted with native_catalog:true and served
back through the whole stale window, so an Ollama that was model-less at first open kept an
empty row after models were pulled. Mirror cached_provider_model_ids: empty native rows are
valid only inside the TTL, never stale-served, and not resurrected when the live probe fails
(the caller falls through to the generic /v1/models fallback instead).
2026-09-16 17:06:26 -07:00
teknium1
9b0c53b43a refactor(opencode): fold the family-path heal into normalize_opencode_base_url as a table
The salvaged fix added a 30-line `_heal_opencode_family_path` helper plus six tests
for one behaviour. The relay path per family is a three-entry table
(`_OPENCODE_FAMILY_PATHS`), so the heal is one `re.fullmatch` on `/zen(/go)?(/v1)?`
inside the normalizer that already owns the opencode.ai host check and the
anthropic `/v1` symmetry — no second urlparse, no separate helper.

Tests trimmed to two invariants: the family path follows the resolved provider
(both directions, both api modes, custom proxy and non-/zen path controls) and the
end-to-end `resolve_runtime_provider(requested="opencode-go")` with a Zen-pinned
`model.base_url` — the reporter's path. The mirror runtime test duplicated the
unit parametrization.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 16:54:57 -07:00
KeyArgo
723b462b12 fix(cli): heal the OpenCode family path in a carried-over base_url
normalize_opencode_base_url() healed only the /v1 suffix, so a model.base_url
pinned to the Zen relay (https://opencode.ai/zen/v1) survived a switch to an
opencode-go model. The two relays serve different model sets, so every request
went to Zen and 401'd ("Model mimo-v2.5 is not supported") with the provider
already correctly resolved to opencode-go.

The family path segment (/zen vs /zen/go) is now healed to the resolved
provider family inside the same normalizer every resolution path already
funnels through (_finalize_base_url, model_switch, custom family providers).
Only /zen-rooted paths on opencode.ai hosts are rewritten; custom
OPENCODE_*_BASE_URL proxies and unrelated paths are left alone.

Refs #112600
2026-09-16 16:54:57 -07:00
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
KoNit-K
91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
teknium1
9aaa71367b fix(catalog): apply the opencode free exclusion on the live-first keyed Zen/Go picker 2026-09-15 18:23:29 -07:00
teknium1
bb5d745ab1 fix(catalog): keep the delisted deepseek-v4-flash-free out of the live OpenCode picker too
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).

Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
2026-09-15 18:23:29 -07:00
Louis.Anson
c008cfe920 fix(copilot): pass two-slash custom-model ids through unchanged
Copilot enterprise custom models (BYOK) expose catalog ids shaped
owner/sub/model. normalize_copilot_model_id() only tried the full id and the
id minus its FIRST segment, and returned the stripped form when neither
matched the (unreachable) catalog - corrupting those ids into sub/model,
which the CLI's second normalization pass then stripped again to model. The
Copilot API answered HTTP 400 model_not_supported on every call.

Only accept the strip guess when the remainder is itself a flat id: a result
that still contains "/" cannot be a Copilot id.

Fixes #110597
2026-09-15 04:47:37 -07:00
teknium1
fdd5995ecb fix(tui_gateway): serve fails closed when hosting a second profile home; profile RPCs bind the full runtime scope
`hermes serve` / the Desktop backend hosted many profile homes (session
profile_home, the `profile` RPC param, hosted rooms) but never called
agent.secret_scope.set_multiplex_active, so every unscoped get_secret read for
a secondary silently returned the LAUNCH profile's os.environ value, and
@_profile_scoped bound only HERMES_HOME: `config.get full` for profile B
expanded B's `${VAR}` refs to the default profile's plaintext credentials,
model.options listed the default's env-keyed providers, llm.oneshot billed the
default's auxiliary key.

- tui_gateway/launch_profile_policy.py (was launch_terminal_policy.py): the
  first time _profile_home registers a non-launch home the process freezes
  the launch env and flips get_secret to fail closed
  (activate_multi_profile_hosting); launch_secret_scope composes the launch
  profile's .env + external sources over that frozen env so systemd / op-run
  injection survives the flip while a secondary never sees it.
- model_switch._profile_runtime_scope_tokens is the ONE composer for
  home + secret + terminal scope: a named profile binds its own files; the
  launch profile binds its frozen-env scope once multiplexing is active and
  stays unscoped in a single-profile process (legacy os.environ precedence).
  _profile_scoped, _profile_scoped_rpc, _session_profile_runtime_scope,
  _bind_build_profile_scopes and _prepare_turn_input all go through it.
- Hosted-room / Group Chat turns for a DEFAULT-profile member in a
  `multiplex_profiles: true` gateway no longer die at agent build with
  UnscopedSecretError: `profile_home is None` was treated as "no scope"
  in _start_agent_build._build and _prepare_turn_input.
- llm.oneshot runs under the session's (or params.profile's) scope;
  _lap_builtin_rows / _overlay_has_creds / _provider_has_credentials read
  provider keys through _scoped_key_env instead of raw os.environ;
  methods_groups._profile_execution_policy resolves the hosted-room policy
  (which reads provider credentials) under the profile's full scope.

Live repro (real `hermes serve`, two homes, config.get {key: full, profile: b}):
base  a_ref: <A_VALUE>  b_ref: ${B_ONLY_TOKEN}  env_ref: <ENV_INJECTED>
head  a_ref: ${A_ONLY_TOKEN}  b_ref: <B_VALUE>  env_ref: ${ENV_INJECTED_TOKEN}
Control (one home, --single): launch config still resolves env_ref from os.environ.
2026-09-15 03:46:29 -07:00
Teknium
d26dbb2f77 Port from OpenHands/OpenHands#16758: Anthropic model catalogs no longer stop at the first page
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.

- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
  follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
  and it now honors the base_url argument instead of hardcoding
  api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
  incl. single-page and stuck-cursor termination; updated the two URL-pinning
  pool-discovery tests for the ?limit=1000 contract
2026-09-13 21:07:35 -07:00
kshitijk4poor
db66b21ae4 fix(copilot-acp): honor the fetch_models None-on-failure contract; bound the probe budget; refresh-proof the memo
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
  missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
  hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
  return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
  timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
  session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
  the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
2026-09-13 19:30:48 +05:30
kshitijk4poor
d5ceff958d fix(models): memoize the Copilot ACP session probe; source it from the provider profile
`/model <x>` onto copilot-acp validates through `models_validate._static_catalog`, which
reads `provider_model_ids` with no disk cache. After the session probe landed, every such
switch spawned `copilot --acp`, ran the handshake, and killed it (1-3 s; up to the 15 s
probe timeout when the CLI is installed but the session stalls). The GitHub-API tier that
path used before sat behind a 5-minute in-memory memo; the ACP tier now has the same memo,
and it remembers failures too so a broken CLI is not re-spawned per switch.

The probe itself moves to `CopilotACPProfile.fetch_models` — the slot that already said
"model listing is handled by the ACP subprocess" and returned None — so hermes_cli/models.py
no longer hand-builds `CopilotACPClient` kwargs that `profile.create_client` owns.
Discovery failures are logged at debug instead of swallowed.

Tests: the two picker wiring tests collapse into one parametrized contract; a new test
proves three consecutive switch validations pay one probe and a failed probe is not retried
(fails when the memo read is removed).
2026-09-13 19:17:06 +05:30
Gille
0cd897286a fix(models): discover Copilot ACP session catalog 2026-09-13 19:17:06 +05:30
teknium1
d364473620 refactor(model-providers): fetch_models stubs drop to the ABC; commandcode and AI Gateway catalog GETs go through open_credentialed_url
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
2026-09-13 05:19:48 -07:00
Teknium
5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
Teknium
30e4673776 refactor(models): import openrouter_variant_base from its defining module
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
2026-09-11 16:50:20 -07:00
jakobdylanc
284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Siddharth Balyan
04a76c4109 fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller

Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a
sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host
serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs
after the account is persisted: a config on the free tier's route moves to the account's inference
host and the recommended default for the account's plan, through the same config write a plain
Nous login uses; a config on the user's own model is left alone. The pick is the one
GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model
so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default.

* fix(auth): a sign-in completion with no eligible recommendation leaves no default model

The static provider-wide default is not narrowed by the account's plan or org policy, so writing it
as a fallback could persist a model the account may not use. When the recommendation cannot yield a
model, the route still moves to the account's host but model.default is left unset; the CLI says so
and points at `hermes model`.

* docs(free-tier): say what happens when no recommendation is available after sign-in

* fix(auth): sign-in completion moves the host and clears the default in one config write

Two writes could fail between them and leave the account host paired with nous/welcome.
_update_config_for_provider gains clear_default so the caller with no model to offer removes
model.default in the same atomic write that sets the host.
2026-09-11 03:45:31 +05:30
Teknium
9a3a15c87a fix(models): a vendor's own id on its first-party provider never re-routes
215fd0ecb9 made the current provider's live catalog outrank static
guesses, but a live catalog that could not be fetched (Codex outage, cold
cache) looked identical to "not served", so `/model gpt-6-astra` on
openai-codex still walked to OpenRouter (which relists every vendor)
whenever an OpenRouter key existed. Same shape for grok-* on xai-oauth
and claude-* on anthropic.

detect_provider_for_model now treats a vendor id typed on that vendor's
single-vendor first-party provider as a selection: stay, let the vendor
accept or reject it (the existing "not found in listing" note still
fires). Cross-vendor remaps (claude id on Codex -> keyed OpenRouter) and
aggregator / custom / multi-vendor reseller sessions are unchanged.
2026-09-10 09:11:00 -07:00
Teknium
2466684db5 fix(models): never auto-switch to a provider the user has no credentials for
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).

detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.

Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
2026-09-10 03:58:37 -07:00
Teknium
215fd0ecb9 fix(models): a model the current provider serves never re-routes to OpenRouter
detect_provider_for_model() walked static catalogs and then the OpenRouter
catalog. Providers whose live catalog outruns the static list (Codex
early-access ids, Nous Portal slugs, Ollama Cloud) had nothing to stop the
ladder, so `/model gpt-6-astra` from a Codex session, or `/model glm-5.3-flash`
from a Nous-sub-only session, silently rebuilt the agent on metered OpenRouter
(or on a vendor the user has no key for).

The current provider's disk-cached live catalog (cached_provider_model_ids, 1h
TTL + SWR) is now consulted first: an exact or bare-name hit stays put and
resolves to the provider's own spelling. Fixes the whole class at the shared
choke point, covering /model, ACP, oneshot and the dashboard.

Refs #97487 (sj-unit72 diagnosed the ollama-cloud instance of this).
2026-09-10 02:44:26 -07:00
teknium1
53b9615ea6 test(hermes_cli): keep one opencode-go delisting invariant; record the delisting beside the keyed-suffix set
Drop the floor-pin test from #106290 (a snapshot of the static table); the
end-to-end merge test through provider_model_ids with the real curated floor
already fails if the slug is re-added. Note beside
_OPENCODE_FREE_KEYED_SUFFIX_MODELS why ox-alpha-free stays excluded from the
keyless catalog even though it left the curated floor.
2026-09-09 11:51:00 -07:00
kshitijk4poor
e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30