Commit Graph

496 Commits

Author SHA1 Message Date
kshitijk4poor
b2d2cf4e24 fix(cli): the startup route decodes custom:<name> from the caller's providers, not a second config read
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
2026-09-19 01:44:53 +05:30
teknium1
8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00
finn763
cdacc2bcf1 fix(picker): never wait on provider catalog probes in the model-options read path
Opening the desktop model picker could sit on skeleton placeholders for 70s+
because a normal open (refresh=False) ran live provider catalog probes inline:
the serial row pass fetched each stale provider's /v1/models itself, and the
parallel prefetch joined every worker, so one degraded provider (hanging
endpoint, failed auth probe) held the whole response.

A normal open is now a read path:

* cached_provider_model_ids(non_blocking=True) serves the same-credentials
  disk entry of any age and refreshes it in a daemon thread; a cold row
  returns [] so the row keeps its curated list.
* list_authenticated_providers(non_blocking_catalogs=True) skips the joining
  prefetch and reads every row cache-only; build_model_options_payload turns
  it on for refresh=False (api-server / dashboard / TUI model.options).
* rows whose catalog is still warming carry catalog_pending, so a GUI can
  tell "not resolved yet" from "that is the provider's catalog".
* Ollama Cloud's 8s probe becomes a cached read + background warm; the
  loopback LM Studio probe stays (1.5s, cannot be a degraded remote).
* The SWR write now takes the cache lock: the read path spawns one warm per
  stale provider, and concurrent load-modify-save dropped rows.

An explicit refresh (Refresh Models) still probes live and blocking.

(cherry picked from commit 0f2c2e7a3631f2c8f3b5140642306eaf2d688682)
2026-09-18 11:01:52 -07:00
kshitijk4poor
debfc7420b refactor(providers): finish the keyless cleanup
`get_api_key_provider_status` pre-populated configured/logged_in/base_url/
key_source purely so the deleted keyless short-circuit could return them;
every one is now recomputed before the return, so build the snapshot once
instead of overwriting four dead placeholders. Same keys, same order, same
values.

The `_OPENCODE_FREE_EXCLUDED_MODELS` comment also still implied an
anonymous path that no longer exists; it now names the two ids it holds
and why.
2026-09-18 15:40:37 +05:30
kshitijk4poor
66a29c3e9d docs: correct the free-tier narrative the removal left behind
The keyless free tier is gone, but five comments and two test docstrings
still described its routing rung: the resolve_runtime_provider ladder
docstring listed a step that no longer exists, and both target_model call
sites plus their regression tests explained themselves in terms of a
`*-free` default being routed to the keyless Zen relay. They now state
what the code actually does (the model-keyed rungs pick the relay and
api_mode).

Also drops the two comments that only narrated the removal
(_OPENCODE_FREE_EXCLUDED_MODELS' history, and an orphan note in
auxiliary_client._resolve_api_key_branch).
2026-09-18 15:40:37 +05:30
kshitijk4poor
b388a48a2e refactor(providers): drop the now-dead keyless provider plumbing
The keyless OpenCode free tier was the only provider that ever set
`HermesOverlay.keyless`, so the flag and everything keyed off it is now
unreachable: the `keyless=` field on `HermesOverlay` and
`ProviderDescriptor`, the `_overlay_has_creds` early return, both
`_provider_is_keyless` copies (auth.py and inventory.py), the
`get_api_key_provider_status` keyless short-circuit and its
`key_source: "keyless"` placeholder, and the empty
`_KEYLESS_STABLE_CACHE_PROVIDERS` set whose `_credential_fingerprint`
branch could never match.

Dropping the two catalog-derived test exemptions follows: they computed
the empty set.
2026-09-18 15:40:37 +05:30
Ritesh Patel
ce6c45d335 fix(models): restore _OPENCODE_FREE_EXCLUDED_MODELS used by the Zen/Go live pickers
The removal dropped the constant but _profile_live_catalog() still filters
live-first Zen/Go listings through it (relay advertises delisted *-free slugs
that 400/403 on POST). Without it, any keyed live picker hit a NameError.
Restored with the original contents (ox-alpha-free, deepseek-v4-flash-free)
and an updated comment; extension point kept for future relay delistings.
2026-09-18 15:40:37 +05:30
Ritesh Patel
998f614c7f feat(providers): remove the keyless opencode-free tier
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:

- drop the opencode-free provider row, aliases (free/opencode_free), model
  catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
  clear error naming the removal

Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
2026-09-18 15:40:37 +05:30
teknium1
931b5ff9e7 fix(opencode): every credential-resolution surface keys off the model it will send; family heal only for built-in providers
Closes the remaining atoms of #112600.

A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
   its siblings still resolved credentials against config's `default`: the CLI
   auth-fallback rung, `--resume` credential re-resolution, the gateway
   provider-override helper (channel overrides, persisted /model switches,
   API-server provider refresh), the gateway fallback chain, the TUI /model
   switch-from runtime and ACP agent construction. With a `*-free` default the
   OpenCode free-tier rung fired first and a Go-only model was built against
   the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
   the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
   optional `target_model` and the two test stubs of it accept the kwarg.

B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
   provider matched by opencode_provider_family, including custom providers
   merely named after a family (`opencode-go-bridge`, #85589) whose relay the
   user declared explicitly in `providers:`. The family heal now applies to the
   built-in canonical providers only; custom prefix-named providers keep their
   per-model api_mode routing and /v1 handling. Documented in the providers
   guide.

C) Same function: the official-host check uses parsed.hostname (a port no
   longer defeats the heal) and only the path is edited, so query/fragment
   round-trip instead of being dropped.

Fixes #112600
2026-09-17 08:59:22 -07:00
teknium1
d6d6565d90 fix(picker): native Ollama cache rows use the 300s TTL and never stale-serve empty catalogs
Routing the native /api/tags probe through cached_fetch_api_models made the CURRENT endpoint's
probe cache-first with the generic 1h TTL (+7d SWR), so a model pulled after the first picker
open stayed invisible for up to an hour across restarts; the built-in `ollama` slug clamps to
_OLLAMA_LOCAL_MODELS_CACHE_TTL in cached_provider_model_ids but this path did not. Pass the
300s native TTL for the native admission.

An authoritative EMPTY native catalog was also persisted with native_catalog:true and served
back through the whole stale window, so an Ollama that was model-less at first open kept an
empty row after models were pulled. Mirror cached_provider_model_ids: empty native rows are
valid only inside the TTL, never stale-served, and not resurrected when the live probe fails
(the caller falls through to the generic /v1/models fallback instead).
2026-09-16 17:06:26 -07:00
teknium1
9b0c53b43a refactor(opencode): fold the family-path heal into normalize_opencode_base_url as a table
The salvaged fix added a 30-line `_heal_opencode_family_path` helper plus six tests
for one behaviour. The relay path per family is a three-entry table
(`_OPENCODE_FAMILY_PATHS`), so the heal is one `re.fullmatch` on `/zen(/go)?(/v1)?`
inside the normalizer that already owns the opencode.ai host check and the
anthropic `/v1` symmetry — no second urlparse, no separate helper.

Tests trimmed to two invariants: the family path follows the resolved provider
(both directions, both api modes, custom proxy and non-/zen path controls) and the
end-to-end `resolve_runtime_provider(requested="opencode-go")` with a Zen-pinned
`model.base_url` — the reporter's path. The mirror runtime test duplicated the
unit parametrization.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 16:54:57 -07:00
KeyArgo
723b462b12 fix(cli): heal the OpenCode family path in a carried-over base_url
normalize_opencode_base_url() healed only the /v1 suffix, so a model.base_url
pinned to the Zen relay (https://opencode.ai/zen/v1) survived a switch to an
opencode-go model. The two relays serve different model sets, so every request
went to Zen and 401'd ("Model mimo-v2.5 is not supported") with the provider
already correctly resolved to opencode-go.

The family path segment (/zen vs /zen/go) is now healed to the resolved
provider family inside the same normalizer every resolution path already
funnels through (_finalize_base_url, model_switch, custom family providers).
Only /zen-rooted paths on opencode.ai hosts are rewritten; custom
OPENCODE_*_BASE_URL proxies and unrelated paths are left alone.

Refs #112600
2026-09-16 16:54:57 -07:00
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
KoNit-K
91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
teknium1
9aaa71367b fix(catalog): apply the opencode free exclusion on the live-first keyed Zen/Go picker 2026-09-15 18:23:29 -07:00
teknium1
bb5d745ab1 fix(catalog): keep the delisted deepseek-v4-flash-free out of the live OpenCode picker too
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).

Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
2026-09-15 18:23:29 -07:00
Louis.Anson
c008cfe920 fix(copilot): pass two-slash custom-model ids through unchanged
Copilot enterprise custom models (BYOK) expose catalog ids shaped
owner/sub/model. normalize_copilot_model_id() only tried the full id and the
id minus its FIRST segment, and returned the stripped form when neither
matched the (unreachable) catalog - corrupting those ids into sub/model,
which the CLI's second normalization pass then stripped again to model. The
Copilot API answered HTTP 400 model_not_supported on every call.

Only accept the strip guess when the remainder is itself a flat id: a result
that still contains "/" cannot be a Copilot id.

Fixes #110597
2026-09-15 04:47:37 -07:00
teknium1
fdd5995ecb fix(tui_gateway): serve fails closed when hosting a second profile home; profile RPCs bind the full runtime scope
`hermes serve` / the Desktop backend hosted many profile homes (session
profile_home, the `profile` RPC param, hosted rooms) but never called
agent.secret_scope.set_multiplex_active, so every unscoped get_secret read for
a secondary silently returned the LAUNCH profile's os.environ value, and
@_profile_scoped bound only HERMES_HOME: `config.get full` for profile B
expanded B's `${VAR}` refs to the default profile's plaintext credentials,
model.options listed the default's env-keyed providers, llm.oneshot billed the
default's auxiliary key.

- tui_gateway/launch_profile_policy.py (was launch_terminal_policy.py): the
  first time _profile_home registers a non-launch home the process freezes
  the launch env and flips get_secret to fail closed
  (activate_multi_profile_hosting); launch_secret_scope composes the launch
  profile's .env + external sources over that frozen env so systemd / op-run
  injection survives the flip while a secondary never sees it.
- model_switch._profile_runtime_scope_tokens is the ONE composer for
  home + secret + terminal scope: a named profile binds its own files; the
  launch profile binds its frozen-env scope once multiplexing is active and
  stays unscoped in a single-profile process (legacy os.environ precedence).
  _profile_scoped, _profile_scoped_rpc, _session_profile_runtime_scope,
  _bind_build_profile_scopes and _prepare_turn_input all go through it.
- Hosted-room / Group Chat turns for a DEFAULT-profile member in a
  `multiplex_profiles: true` gateway no longer die at agent build with
  UnscopedSecretError: `profile_home is None` was treated as "no scope"
  in _start_agent_build._build and _prepare_turn_input.
- llm.oneshot runs under the session's (or params.profile's) scope;
  _lap_builtin_rows / _overlay_has_creds / _provider_has_credentials read
  provider keys through _scoped_key_env instead of raw os.environ;
  methods_groups._profile_execution_policy resolves the hosted-room policy
  (which reads provider credentials) under the profile's full scope.

Live repro (real `hermes serve`, two homes, config.get {key: full, profile: b}):
base  a_ref: <A_VALUE>  b_ref: ${B_ONLY_TOKEN}  env_ref: <ENV_INJECTED>
head  a_ref: ${A_ONLY_TOKEN}  b_ref: <B_VALUE>  env_ref: ${ENV_INJECTED_TOKEN}
Control (one home, --single): launch config still resolves env_ref from os.environ.
2026-09-15 03:46:29 -07:00
Teknium
d26dbb2f77 Port from OpenHands/OpenHands#16758: Anthropic model catalogs no longer stop at the first page
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.

- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
  follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
  and it now honors the base_url argument instead of hardcoding
  api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
  incl. single-page and stuck-cursor termination; updated the two URL-pinning
  pool-discovery tests for the ?limit=1000 contract
2026-09-13 21:07:35 -07:00
kshitijk4poor
db66b21ae4 fix(copilot-acp): honor the fetch_models None-on-failure contract; bound the probe budget; refresh-proof the memo
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
  missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
  hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
  return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
  timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
  session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
  the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
2026-09-13 19:30:48 +05:30
kshitijk4poor
d5ceff958d fix(models): memoize the Copilot ACP session probe; source it from the provider profile
`/model <x>` onto copilot-acp validates through `models_validate._static_catalog`, which
reads `provider_model_ids` with no disk cache. After the session probe landed, every such
switch spawned `copilot --acp`, ran the handshake, and killed it (1-3 s; up to the 15 s
probe timeout when the CLI is installed but the session stalls). The GitHub-API tier that
path used before sat behind a 5-minute in-memory memo; the ACP tier now has the same memo,
and it remembers failures too so a broken CLI is not re-spawned per switch.

The probe itself moves to `CopilotACPProfile.fetch_models` — the slot that already said
"model listing is handled by the ACP subprocess" and returned None — so hermes_cli/models.py
no longer hand-builds `CopilotACPClient` kwargs that `profile.create_client` owns.
Discovery failures are logged at debug instead of swallowed.

Tests: the two picker wiring tests collapse into one parametrized contract; a new test
proves three consecutive switch validations pay one probe and a failed probe is not retried
(fails when the memo read is removed).
2026-09-13 19:17:06 +05:30
Gille
0cd897286a fix(models): discover Copilot ACP session catalog 2026-09-13 19:17:06 +05:30
teknium1
d364473620 refactor(model-providers): fetch_models stubs drop to the ABC; commandcode and AI Gateway catalog GETs go through open_credentialed_url
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
2026-09-13 05:19:48 -07:00
Teknium
5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
Teknium
30e4673776 refactor(models): import openrouter_variant_base from its defining module
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
2026-09-11 16:50:20 -07:00
jakobdylanc
284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Siddharth Balyan
04a76c4109 fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller

Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a
sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host
serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs
after the account is persisted: a config on the free tier's route moves to the account's inference
host and the recommended default for the account's plan, through the same config write a plain
Nous login uses; a config on the user's own model is left alone. The pick is the one
GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model
so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default.

* fix(auth): a sign-in completion with no eligible recommendation leaves no default model

The static provider-wide default is not narrowed by the account's plan or org policy, so writing it
as a fallback could persist a model the account may not use. When the recommendation cannot yield a
model, the route still moves to the account's host but model.default is left unset; the CLI says so
and points at `hermes model`.

* docs(free-tier): say what happens when no recommendation is available after sign-in

* fix(auth): sign-in completion moves the host and clears the default in one config write

Two writes could fail between them and leave the account host paired with nous/welcome.
_update_config_for_provider gains clear_default so the caller with no model to offer removes
model.default in the same atomic write that sets the host.
2026-09-11 03:45:31 +05:30
Teknium
9a3a15c87a fix(models): a vendor's own id on its first-party provider never re-routes
215fd0ecb9 made the current provider's live catalog outrank static
guesses, but a live catalog that could not be fetched (Codex outage, cold
cache) looked identical to "not served", so `/model gpt-6-astra` on
openai-codex still walked to OpenRouter (which relists every vendor)
whenever an OpenRouter key existed. Same shape for grok-* on xai-oauth
and claude-* on anthropic.

detect_provider_for_model now treats a vendor id typed on that vendor's
single-vendor first-party provider as a selection: stay, let the vendor
accept or reject it (the existing "not found in listing" note still
fires). Cross-vendor remaps (claude id on Codex -> keyed OpenRouter) and
aggregator / custom / multi-vendor reseller sessions are unchanged.
2026-09-10 09:11:00 -07:00
Teknium
2466684db5 fix(models): never auto-switch to a provider the user has no credentials for
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).

detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.

Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
2026-09-10 03:58:37 -07:00
Teknium
215fd0ecb9 fix(models): a model the current provider serves never re-routes to OpenRouter
detect_provider_for_model() walked static catalogs and then the OpenRouter
catalog. Providers whose live catalog outruns the static list (Codex
early-access ids, Nous Portal slugs, Ollama Cloud) had nothing to stop the
ladder, so `/model gpt-6-astra` from a Codex session, or `/model glm-5.3-flash`
from a Nous-sub-only session, silently rebuilt the agent on metered OpenRouter
(or on a vendor the user has no key for).

The current provider's disk-cached live catalog (cached_provider_model_ids, 1h
TTL + SWR) is now consulted first: an exact or bare-name hit stays put and
resolves to the provider's own spelling. Fixes the whole class at the shared
choke point, covering /model, ACP, oneshot and the dashboard.

Refs #97487 (sj-unit72 diagnosed the ollama-cloud instance of this).
2026-09-10 02:44:26 -07:00
teknium1
53b9615ea6 test(hermes_cli): keep one opencode-go delisting invariant; record the delisting beside the keyed-suffix set
Drop the floor-pin test from #106290 (a snapshot of the static table); the
end-to-end merge test through provider_model_ids with the real curated floor
already fails if the slug is re-added. Note beside
_OPENCODE_FREE_KEYED_SUFFIX_MODELS why ox-alpha-free stays excluded from the
keyless catalog even though it left the curated floor.
2026-09-09 11:51:00 -07:00
kshitijk4poor
e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30
kshitijk4poor
4e2a871adb fix(models): clear probe negative-cache entry on a successful probe 2026-09-09 21:16:48 +05:30
finn763
48b8528e7c fix(desktop): stop UI freeze on unreachable provider Closes #81123 2026-09-09 21:16:48 +05:30
kshitijk4poor
1eaeb73839 fix(models): OpenRouter disk snapshot follows the configured catalog TTL, not the default constant
refresh_interval_seconds() honours model_catalog.ttl_minutes / legacy ttl_hours;
reading DEFAULT_TTL_MINUTES would let the snapshot and the manifest it is
filtered from go stale on different clocks for anyone who changed the TTL.
2026-09-09 21:16:28 +05:30
kshitijk4poor
8b66bc88df refactor(models): OpenRouter disk cache reuses catalog TTL and atomic JSON cache helpers 2026-09-09 21:16:28 +05:30
Hyperion
4c003069f3 fix(models): persist curated OpenRouter catalog to disk so picker opens fast
Re-derived from PR #96099 (f127ec4e) on current main; the :nitro/:floor validate hunk is omitted because main already handles routing suffixes in hermes_cli/models_validate.py.
2026-09-09 21:16:28 +05:30
Teknium
734461d213 fix(models): same-URL custom endpoints stop evicting each other's cached catalog; no-probe picker opens revalidate
Two picker-freshness defects in cached_fetch_api_models():

1. The disk cache row was keyed on base_url only, with the credential
   fingerprint stored inside the row. N custom_providers entries sharing
   one proxy URL with different keys (#106184) took turns overwriting the
   single slot; every sibling then failed the fingerprint check, got an
   empty catalog, and disappeared from the Desktop pickers (which hide
   zero-model rows). Key on url#fingerprint so each credential owns a row.

2. cache_only opens (Desktop model.options without refresh) served a
   past-TTL row for up to 7 days without ever revalidating, so a model
   loaded on a non-current local endpoint stayed invisible until the user
   found "Refresh Models". Serve the stale row AND spawn the same
   off-thread SWR refresh the blocking path uses; the caller still never
   waits on the network.

Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
  GUI no-probe open  before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
                     after  {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
2026-09-09 03:33:06 -07:00
Teknium
520e63661c fix: keep command-auth model discovery lazy across config and setup 2026-09-07 21:22:49 -07:00
kshitijk4poor
e583d68c14 perf(models): stop revalidating a fresh provider cache just because it lists Astra
``gated_cache`` bypassed both the fresh-hit and the stale-while-revalidate branches of
``cached_provider_model_ids`` whenever the disk entry contained Astra. For an entitled
account that is every entry: live discovery rewrites the entry with Astra → next call is gated
again → a blocking /models round-trip on every picker open, for exactly the providers the user
is most likely on. The parallel prefetch's staleness check didn't know about the gate either, so
it skipped the slug and the blocking call landed in the serial picker loop the prefetch exists to
avoid.

Gate on provenance, not contents: only live discovery ever writes Astra into a same-credential
entry (the static/offline paths filter it out), so a fresh entry IS the entitlement record. The
one filter that matters stays — a stale entry served because the refresh failed drops Astra, and
the entry itself is left intact so the next successful fetch restores it.

The regression test now pins both halves: fresh entry served with Astra and zero live calls;
failed refresh past the fresh window serves the entry minus Astra.
2026-09-07 21:43:54 +05:30
kshitijk4poor
b6852995ed refactor(openai): one is_astra_model predicate; gate Astra at the untrusted Codex inputs
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.

``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.

``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.
2026-09-07 21:43:54 +05:30
Eva
fdb76ce643 fix(openai): include canonical API picker identity in Astra discovery gate
(cherry picked from commit de57e601193a955a69201e1a274d5ccab83a7634)
2026-09-07 21:43:54 +05:30
Eva
328542d807 fix(openai): revalidate Astra at cached and saved-model picker boundaries
(cherry picked from commit f46e2f4c0609c1543d5636c780c6da0d5ce09a56)
2026-09-07 21:43:54 +05:30
Eva
2c315ff59b feat(openai): add GPT-6 Astra baseline support
(cherry picked from commit a8c53d20c6b16cc35745e364e16bb7259166a1d3)
2026-09-07 21:43:54 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium
c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium
7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium
d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00