The whitespace reject ran for every provider, so a self-hosted catalog or a
user-configured base_url could not select an id that legitimately contains
spaces. Cloud providers still reject. Picker payloads drop an id that
reject will still refuse.
Fixes#43140
Fold _profile_owns_catalog and _profile_owned_catalog into _profile_catalog ->
(catalog, authoritative) so _validate_live_listing has a single owned-catalog
branch; drop the redundant catalog-hit test (covered by the owns-catalog accept
tests in test_setup_provider_catalog.py).
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The validation ladder had no llamacpp branch: a switch to a freshly downloaded
local model fell through to the generic live-listing probe, which hard-rejects
when the spawn-only /v1/models listing has not learned the new file yet — so the
Local Models Use button and the composer picker could never succeed for a
non-catalog model. The managed runtime now validates against the staged library
on disk (the source of truth for what the user downloaded), case-insensitively;
ids that were never staged keep the live-listing verdict.
The activate flow's self-heal had the same blind spot: its rescan only ran when
this process supervised the server, but ensure_local_runtime returns None when
another process owns it — precisely the desktop situation after a download job
bounced the router once. Probe the live listing through the persisted endpoint
and bounce the router when it lacks the model.
A plugin whose `fetch_models` (or `models_url`) IS its catalog was validated
against the generic `{base_url}/v1/models` listing; when that endpoint 200s
with a different product catalog (a relay serving a subscription SKU set,
#101705) every model the `/model` picker offered was rejected with
suggestions from another vendor. `_validate_live_listing` now accepts an
exact match against `provider_model_ids()` — the picker's list — before the
generic probe, only for profiles that own their catalog; everything else
falls through unchanged.
Part of #116408
Part of #101705
A composer, script or older client can pin a model override the selected
provider does not serve (gpt-5.5 on anthropic; the incident pair
deepseek/deepseek-v4-flash-0731 on openai-codex). The gateway minted the
session anyway and the FIRST turn died with the provider's 404, leaving a
dead chat the user had to diagnose and recreate.
session.create now checks the pair before any session state exists and
answers JSON-RPC -32602 naming the model, the provider and up to five
closest models from that provider's curated catalog (error.data carries
them structured). The check is offline and refuses only what Hermes knows
belongs elsewhere: a foreign-family name another native vendor's catalog
lists, or any foreign-family name on the strict OAuth catalogs
(openai-codex / xai-oauth). Custom endpoints, aggregators, same-family
names the curated list lacks and names no catalog lists stay permissive.
Without an explicit provider the pair is judged against the provider the
session would build with (profile config, then env) under the profile's
scope.
Direction and reject-before-side-effects shape from #96845 by
@victorftrdba, trimmed to the offline catalog rule.
Fixes#96817
Two diagnostics diverged after a session-only `/model` switch (#111436).
/status: `_status_model_route` only took `context_total` from a live/cached
compressor or the raw `model.context_length` pin, so between turns (no
compressor yet) it fell to the occupancy-only line ("Context: ~79,455
tokens") while /context resolved the 1M window for the same session. /status
now runs the same resolver /context uses (`_resolve_gateway_model_context`,
off the event loop — it can probe /models), fed the WINNING route's
provider/base_url/api_key so the lookup targets the endpoint that serves the
displayed model, never a losing route's endpoint. The raw config pin moves
into the resolver, which already drops it when the route no longer matches
the configured one — a session switch must not inherit the default model's
pin. A window the resolver merely invented (unknown model →
DEFAULT_FALLBACK_CONTEXT) is grounded via a catalog match: `context_source`
is "default" only when no catalog entry matches, and /status keeps the honest
occupancy-only line for that case (catalog-listed 256K models still display).
Validator: `_validate_anthropic_messages` used one soft-accept message for
both "listing unreachable" and "listing answered 200 but lacks the slug", so
a reachable endpoint was described as one that "does not implement GET
/v1/models". The two cases now get distinct wording; the reachable case
matches case-insensitively and surfaces alias candidates at similarity 0.4
(`kimi-k3` vs `k3` ≈ 0.44 sits below the default 0.5 cutoff).
Slimmer redo of #111458 by @KoNit-K, which resolved only the override route
(not persisted/DB routes) and displayed the fallback window unconditionally.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:
- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
model_switch applied it. Version bumps, dated snapshots and qualifiers
all sit inside 0.9 of a sibling, so a newer release the listing lacked
was swapped for the older one under the user's label. The matcher now
does exact membership -> suggestion text only; the id goes to the wire
verbatim and a genuine typo is refused with the listed siblings named.
Every branch that carried the correction (live listing, static catalog,
curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
loses it in one place.
- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
bare key (the slug Desktop picker rows carry) validated as a built-in
and hit the hard-rejecting live-listing branch; the same endpoint as
`custom:<key>` soft-accepted. Both spellings now validate as the user's
custom endpoint.
- apps/desktop: `manualPickRemoved` (composer reseed) and
`reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
sticky pick to the profile default / the row's first model whenever the
provider row did not list it. Rows are hints (discovered, curated,
capped); the gateway's switch result is the only authority on a pick.
Both helpers are removed; the pick stays put.
Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
Follow-up to the salvaged #96379 commits: the fallback verdict is computed once
(`accepted = api_mode in chat modes`), the warning says what actually happened
("accepted without verification" vs "was not saved") instead of promising a
save it then refused, and the contributor's ten regression tests collapse to two
parametrized invariants (chat modes persist unverified; other modes still reject;
a reachable catalog stays authoritative).
Allow custom chat-completions endpoints without a usable model catalog to persist explicitly requested model IDs with the existing verification warning.
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.