Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
Only the --provider path handed user_providers/custom_providers to
resolve_alias, so the provider-identity comparison behind the
current-provider preference could not see legacy custom_providers
entries on the implicit path (/model <id>) or the authenticated-provider
fallback. On custom:corp-llm, /model shared-model picked another
provider's alias for the same model id and switched to its base_url.
Pass both provider maps on every resolve_alias call site. The two
existing resolve_alias stubs in test_ollama_cloud_auth.py accept the
extra positional args; their assertions are unchanged.
The --provider ownership check compared normalize_provider() spellings, but a
legacy custom_providers entry resolves to id "custom:<name>". An alias with
"provider: corp-llm" under "--provider corp-llm" was therefore dropped and the
switch fell back to the provider's own base_url/key instead of the alias's.
Compare resolve_provider_full(...).id on both sides (normalize_provider when
unresolvable), in the ownership check and in resolve_alias's reverse-lookup
preference loop, which now receives the user/custom provider config.
An explicit `/model <id> --provider X` switch could report target_provider=X
while base_url still pointed at the previous provider's endpoint: resolve_alias
finds a direct alias by reverse model-ID lookup, and the base_url override
applied that alias unconditionally even when the alias belonged to a different
provider. Requests would then be sent to the old provider's URL under the new
provider's identity.
Two narrow changes:
- resolve_alias's reverse lookup now prefers an alias whose provider matches
current_provider, falling back to first-match only when none does. Insertion
order is not a routing decision.
- The direct-alias base_url override now ignores (and stops reporting) an alias
whose provider differs from target_provider.
Exact alias-name lookup stays provider-agnostic, and the first-match fallback is
preserved, both pinned by tests.
[salvaged onto the refactored switch pipeline: the ownership guard now sits in
_route_explicit_provider, where resolved_alias is settled, so both the runtime
resolve (explicit_base_url) and the direct-alias override see the filtered alias]
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.
Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.
/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The /model -> local/ flow offers the 'llamacpp' row from staged GGUFs alone,
but resolve_provider_full only admitted that id behind a live endpoint, so
selecting a staged model died with "Unknown provider 'llamacpp'" before the
runtime seam could start or attach a server.
- providers.py: LLAMACPP_PROVIDER_ID/ALIASES are the single definition; the
llamacpp rung resolves when a server is reachable OR a model is staged.
- inventory.py: the picker row's slug comes from that definition.
- endpoint.py: honour local_runtime.detect_ports (every provider caller
invoked the resolver bare, leaving the documented knob dead).
- model_switch.py: a local-runtime alias failure surfaces the seam's own
message instead of an API-key hint.
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.
`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.
Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).
Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
get_authenticated_provider_slugs only needs slugs but ran the full listing,
which force-refreshed every stale catalog (6 probes on a cold cache with 6
keys) during alias-fallback routing of a model switch. Pass
non_blocking_catalogs=True: same slug set, 0 forced probes.
#116138 taught the bare-`custom` switched arm to keep the session endpoint when the credential
ladder answered with the `OPENROUTER_BASE_URL` mirror. That test was host/URL-only, so it also
fired when the ladder's answer came from an endpoint configured *ahead* of the OpenRouter rung —
`CUSTOM_BASE_URL`, or a trusted `model.base_url` holding the same URL as the mirror (two env vars
aimed at one proxy). In that configuration the switch paired the new model with the previous
provider's host and key again, i.e. the #73680 shape this arm exists to prevent.
The guard now asks the ladder's own question: the mirror counts as a fallback only when
`CUSTOM_BASE_URL` and a trusted `model.base_url` are both absent, so a URL that equals the mirror
because the user configured it there still wins.
Second act of `test_switch_to_bare_custom_ignores_an_openrouter_mirror` pins it: with both env
vars pointed at the same proxy the switch adopts that endpoint; without the precision check the
test fails (`api.anthropic.com` kept instead of the configured proxy).
The switched-provider bare-`custom` arm adopts the resolved endpoint unless it is the
OpenRouter default. `_fell_back_to_openrouter_default` only knew the built-in
`openrouter.ai` host, so an `OPENROUTER_BASE_URL` mirror — the credential ladder's last
rung, reached whenever no trusted `model.base_url`/`CUSTOM_BASE_URL` exists — looked like a
configured custom endpoint: a session on another provider switching to bare `custom` landed
on the mirror with the `no-key-required` placeholder, dropping the key it was using.
The helper now also recognizes that mirror (read the way the resolver reads it, through the
profile's secret scope), so the arm keeps the session's endpoint and key instead. Sharing the
helper also covers the same-provider #74143 path: a `custom`/`local` session on a session-only
endpoint no longer adopts the mirror either — the same "resolution landed on OpenRouter"
case that guard exists for.
The mirror read is a GUARD read, not a credential fetch: a scope failure (unscoped under
multiplexing) leaves the mirror undetected so the caller keeps the session endpoint, instead
of raising out of `switch_model` where the resolver's own read of the same name is suppressed.
`CUSTOM_BASE_URL` and a trusted `model.base_url` still win: those are endpoints configured for
`custom`, which is the point of the arm.
- model_switch: the host-mandated wire mode (codex_responses on api.openai.com)
corrects STALE modes; it must not overwrite the resolver's codex_app_server opt-in.
- runtime_provider: docstring ladder gains step 9 (the overlay); one logger.info when
the overlay replaces a rung's api_mode, since that rung's credential/endpoint is not used.
- gateway/run_agent_cache.py::_rehydrate_session_model_override re-resolved api_mode from the
target model but kept the persisted base_url verbatim, so a /v1-stripped or other-family
relay URL written by an older build stayed on a chat_completions model. Run
normalize_opencode_base_url after re-resolution for opencode-family providers (test drives
the production rehydration entry: red before, green after).
- hermes_cli/model_switch.py::_build_switch_result now uses model_derived_api_mode() instead
of its own exact-key table lookup, so a custom provider extending a family slug
(opencode-go-bridge) gets the same per-model wire on live /model as on resume.
- tui_gateway/server.py::_resolve_agent_model_runtime: the re-derivation guard is
any(overrides.values()) — a model_override dict always carries the three keys, so
'if overrides' was always true.
Part of #96066
Two independent gaps hit a user on provider opencode-go with
deepseek-v4-flash-vision-exp (#96066):
- Vision: the model is absent from models.dev, so get_model_capabilities()
returned None and image_input_mode: auto detoured images through the lossy
text describe path. OpenCode Zen/Go resell vendor previews before the catalog
indexes them; the id's `-vision` token is the vendor's own capability marker,
so _apply_overrides() now fills the catalog gap with a vision-only base for
any OpenCode-family provider (built-in or family-slug custom). Catalog entries,
explicit and _default overrides stay authoritative; reasoning stays unknown.
- Route: _resolve_agent_model_runtime() applied a resumed row's persisted
api_mode/base_url verbatim. A row written while the session ran an
anthropic_messages model (MiniMax on Go, or an older build's /zen/v1 relay)
pinned that wire onto a chat_completions model, surfacing as 401s from the
Anthropic transport or the Zen relay. Providers that pick the wire per model
(OpenCode families, Copilot, Nous — the existing _PROVIDER_API_MODE_OVERRIDES
table, now reachable via model_derived_api_mode()) re-derive api_mode from the
target model on resume and heal the relay URL; fixed-wire providers keep
honoring their row.
Salvages #96116 (@Finn763): same two atoms, redone slim (data-driven table
instead of an opencode-only hostname guard in the tui_gateway facade).
The switched-provider comment implied the guard catches the resolver raise;
it is suppress() leaving st.* at the session values that does. The
stay-on-provider comment now states the one semantic widening of moving to
the shared helper: an empty resolver result keeps the session endpoint on
any host, not only off openrouter.ai.
Gate finding: resolve_runtime_provider(requested="custom") never returns an
empty base_url — with no model.base_url but an OpenRouter key present it
lands on OpenRouter's default host, so the "keep the current endpoint"
fallback was dead and an Anthropic session adopting `provider: custom`
hopped to openrouter.ai. The guard _creds_for_current_provider already
carries for this (#74143) is now a shared helper used by both branches;
the branch assigns once instead of assign-then-undo; the import sits with
the other from-imports; the new test's write_text carries an encoding.
Second invariant test pins the no-endpoint case.
The bare-custom credential branch kept the session's current base_url and
key whenever the target was `custom`, regardless of where the session was.
That is right for a bare-custom session picking another model on its own
endpoint, but when the TUI's per-turn config sync adopts `provider: custom`
from an OpenRouter (or any other) session, the new model was paired with
OpenRouter's URL and key and every request 400/404'd (Bug 2 in issue 73680).
Arriving from another provider now resolves the configured custom endpoint
through resolve_runtime_provider; with none configured the current endpoint
is kept, so direct aliases still supply their own URL afterwards and the
bare-custom same-endpoint case is unchanged.
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
Third entry path with the same egress: `_resolve_startup_runtime` seeded the model
from HERMES_INFERENCE_MODEL and went straight to static provider detection, so the
qualified string reached the configured default. Route through
`resolve_startup_model_route` first, as HermesCLI and oneshot now do.
Also: an unconfigured `custom:<typo>:<model>` now logs a warning and leaves the
route undecided instead of steering a garbage model id at the bare custom endpoint,
and the redundant `qualified_model != raw` guard is gone.
`hermes -m custom:jetson-vllm:nemotron-nano-30b` (and `-z`) kept the configured default
provider and sent the UNSPLIT string as the model name, so the whole prompt reached
api.anthropic.com before a 404 — a local endpoint the user picked never saw the request.
`parse_model_input` already decodes the form; nothing on the startup path called it.
`resolve_startup_model_route`, the owner of startup provider/model decoding, now consumes
the colon-qualified form through `parse_model_input` ahead of the `provider/model` branch,
so HermesCLI picks it up unchanged. The oneshot resolver routes through the same owner
before provider auto-detection instead of only after a direct-alias miss.
Live, config `providers.jetson-vllm` at 127.0.0.1:8000 and `model.provider: anthropic`:
before — HTTP 401 from Anthropic, local sink idle; after — sink receives
POST /v1/chat/completions model=nemotron-nano-30b, Anthropic untouched.
Closes#73943. First reported fix: PR #74082 by @Drexuxux (HermesCLI.__init__ decode); the
seam moved to the startup route owner during the Sep-2026 cli.py decomposition.
Co-authored-by: Drexuxux <drexux0@gmail.com>
The explicit_base_url hunk in _creds_for_switched_provider passed every
URL-bearing direct alias's endpoint to resolve_runtime_provider, so a
built-in label (anthropic, openai) resolved its vendor key AGAINST the
alias's foreign host; _apply_direct_alias_endpoint then saw a same-origin
credential and kept it. Gate on _resolves_to_custom: only ollama/vllm-style
labels, the ones the missing-endpoint guard refuses, get the alias URL.
CI: tests/hermes_cli/test_model_alias_credentials.py
::test_builtin_label_does_not_pull_that_providers_key_to_a_foreign_host
[anthropic] red on ce4e9d580ac, green here; direct-ollama-alias /model
tests stay green.
Follow-up on the picked #113768 commit so it clears the constraint that got
e9a54c48f2 reverted (a9fabe43c4): `/model <direct-alias>` resolves the alias
LABEL first and applied the alias endpoint only afterwards, so any guard that
fires on the label alone turns a working switch into "ollama is not connected"
(tests/hermes_cli/test_models.py::TestLocalOllamaModelDiscovery went red again
with the PR as-is).
- hermes_cli/runtime_provider.py::_raise_if_local_alias_missing_endpoint keys
on the class (anything auth.resolve_provider maps to `custom` without a rung
of its own; llamacpp keeps its managed-server fail-fast) instead of a second
hardcoded alias set, requires the providers.<alias> block to actually carry a
base_url (an entry without one falls through to OpenRouter too), and names
the alias plus where to set its endpoint. OPENROUTER_BASE_URL never counts as
the alias endpoint; an explicit api_key does not lift the guard.
- hermes_cli/model_switch.py::_creds_for_switched_provider hands a URL-bearing
direct alias's base_url to the resolver as explicit_base_url, so the alias
has an endpoint at the moment the guard runs (the same URL
_apply_direct_alias_endpoint installs later).
- Tests trimmed to two invariants (raise-and-name incl. the OPENROUTER_BASE_URL
/ explicit-api_key non-lifts + bare-custom control; every endpoint source
resolves to its URL). Docs: troubleshooting entry in the local Ollama guide.
Cron replay snapshots that store the resolved `custom` instead of the alias
are #109765's atom and untouched here.
Top-level `model_aliases:` entries carried credentials into DirectAlias but
the nested `model.aliases:` dict form built it without `api_key`/`key_env`,
so an alias with its own base_url got no credential (401) even though the
URL resolved correctly. Pass both fields through, as the top-level path does.
Source hunk from PR #109834 (earliest filer); invariant tests from the
identical later PR #114472. Nested-alias credential atom of #114471.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Dropping an identity-equal legacy `custom_providers` entry from the candidate list also dropped the
models only that entry declared: `/model gpt-5.4-mini` (declared by the legacy row but not by
providers.relay) fell through to the current provider (openrouter) instead of routing to the shared
endpoint — a fail-open flip versus the pre-#112788 behaviour, which routed it to custom:relay.
Keep every entry as a candidate and credit a duplicate's hits to the providers row it duplicates, so
the collapse loses no declaration. `_duplicates_configured_row` now returns the owning slug.
Also require credential equality (not just endpoint) in the raw-list `provider_key` shortcut: a
hand-written entry stamped `provider_key: relay` on the same URL but a different key_env is a
distinct route and stays ambiguous, matching the issue's "preserve ambiguity when credential differs"
criterion.
`_configured_provider_matches` now drops a `custom_providers` entry when it is a second view of a
`providers.<slug>` row: either that row's compat projection (provider_key names the slug AND the
endpoint matches) or a legacy duplicate with the same provider identity (normalized name, endpoint,
credential identity, api_mode — read through the same normalizer that builds the compat view).
Any difference in endpoint, credential or wire protocol keeps two rows distinct.
Why: the first pass (#113103) only filtered entries whose provider_key equalled a providers slug.
A hand-migrated config that kept the same endpoint in both sections was still reported as
"declared by multiple configured providers (custom:relay, relay)", and on the raw-list fallback
(gateway/agent_init pass cfg['custom_providers'] verbatim when the compat view fails) a hand-written
provider_key pointing at a different endpoint was silently hidden instead of ambiguous.
`_route_configured_provider` treats `custom:<name>` of a projected row as an alias of its slug when
checking whether the session is already on the matched provider, so a session running on
`custom:relay` that types a model providers.relay declares stays on `custom:relay` instead of
flipping target_provider/explicit_provider to `relay`.
Closes the remaining atoms of #112788. Identity-tuple approach ported from the
`_configured_provider_identity` hunk of #112794.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The gateway, Ink TUI and classic CLI hand switch_model() both the modern
``providers:`` map and ``get_compatible_custom_providers(cfg)``, which
re-lists every ``providers.<slug>`` row a second time as a ``custom:<name>``
entry stamped with ``provider_key: <slug>``. ``_configured_provider_matches``
only collapsed identical slugs, so a model declared by exactly one configured
endpoint matched twice (``relay`` and ``custom:relay``) and
``_route_configured_provider`` rejected the switch as "declared by multiple
configured providers" instead of routing it (#112788).
Skip a custom entry whose ``provider_key`` is already a ``providers`` slug in
the candidate list: it is the projection of that very row, not a second
declaration. Hand-written ``custom_providers:`` rows carry no provider_key and
remain independent candidates, so two genuinely distinct endpoints declaring
the same model still stop the switch with the ambiguity message.
Slim redo of #112794 (identity-tuple dedup helper) using the marker the
projection already carries.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The cherry-picked commit extends the config/profile/MCP/managed-scope/completer/
OAuth/skills-manifest signatures. This commit finishes the class and trims it:
- `file_signature()` lives in `utils.py` next to the other stat/metadata helpers
instead of `hermes_cli.managed_scope` (gateway/ and agent/ callers no longer
reach into the managed-scope module for a generic stat helper).
- `hermes_cli/config_effective.py` was left comparing 2-/4-wide prefixes against
the widened `_RAW_CONFIG_CACHE` / `_load_config_cache_sig` records, so
`load_user_config_effective()` re-parsed on every call (3 parses for 3 calls on
an unchanged file, 1 before); index by the new widths.
- Sibling caches keyed on the same (mtime, size) shape and reading the SAME files
now use the helper: `load_env()` memo, `agent/skill_utils` raw-config and
external-dirs caches, `hermes_cli/model_switch` alias identity, `agent/moa_loop`
preset stamp, `hermes_cli/auth` global auth-store memo.
- Tests trimmed to one invariant each (pinned-mtime replacement invalidates; an
unchanged file still hits), both red on origin/main.
Left alone on purpose: `tools/registry.py`, `tools/skills_tool_dedup.py`,
`gateway/status.py`, `hermes_cli/banner.py`, `hermes_cli/main.py`,
`hermes_cli/session_recovery.py` — those fingerprint source files, PID/lock files
or write to persisted on-disk caches shared across processes, where an inode/ctime
key would churn on every checkout/copy rather than catch a replaced config.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
`model.key_env` is not custom-only: the Desktop settings UI stores REGISTRY
provider keys there (e.g. HERMES_CUSTOM_LMSTUDIO_API_KEY with provider
lmstudio, #106336) and auth._model_level_key_env honours it. The previous
predicate (`not custom or route_changed`) therefore wiped that pointer on a
same-provider same-base_url model re-pick and silently broke the user's
credential. The pointer now clears ONLY when provider or base_url changed;
the inline api_key/api rule is unchanged.
Custom-endpoint activation (model_setup_flows_custom) popped base_url /
api_key but never key_env, so a stale pointer from a previous endpoint
outranked the credential it had just written — pop it alongside.
Tests: the same-route re-pick case now covers a registry provider (red on
the old predicate), and the two clear_model_endpoint_credentials tests are
folded into one invariant.
Main no longer clears credentials in web_server's _apply_main_model_assignment;
every /model surface (CLI, gateway, TUI, dashboard) persists through
model_selection_config_updates(), which only dropped api_key/api on a route
change. Custom-endpoint activation writes model.key_env with NO inline key, so
a pointer-only model block survived every provider switch and routed the new
provider's requests to the old endpoint's env var (the PR's Bug 5) — live
repro on main: activate custom_myep -> /api/model/set openrouter left
key_env: CUSTOM_MYEP_API_KEY on disk.
Port the PR's web_server hunk to the one shape function: key_env/api_key_env
clear under the same route-changed rule as api_key (same-route re-pick keeps
them). The dashboard's _resolve_assignment_credentials re-adds the TARGET
provider's own pointer after this, so custom->custom still ends with the new
endpoint's key_env. One invariant test in the one-shape suite.
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.
One request now carries a model pick AND its effort on every surface:
- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
<level>` (validated against `parse_reasoning_effort`; unknown level ->
`MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
`ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
high` applies the effort AFTER the agent swap (`switch_model` re-resolves
`reasoning_config` from config.yaml, so an earlier write is clobbered) with the
pick's scope (session; config on `--global`; `--once` snapshots and restores it).
The `/model` picker gains a third stage, "Reasoning effort for <model>", built
from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
`config.set model "X --reasoning high"` applies after the swap; session pin
(`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
one-turn restore carries `reasoning_config`; re-emits `session_info` so the
status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
`<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
`_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
the Copilot-only inline prompt; Copilot keeps its per-model level set via
`github_model_reasoning_efforts`, other routes get the ladder, catalog
`supports_reasoning=False` skips it) plus a "Reasoning effort for the current
model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
with the same step (+ "Provider default"), stored as
`auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
the task list ("openrouter · model · high"), cleared by "Reset all to auto";
tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
it.
Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
"Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
errors "Model names cannot contain spaces"; after switches and `config.get
reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
"reasoning: high", status bar "fable 5.1 high".
`/model` fed `validate_requested_model()` a key resolved through
`agent.secret_scope.get_secret`, which (multiplexing off) reads only
`os.environ`. Hermes does not export `$HERMES_HOME/.env` into the process
environment, so a `custom_providers` entry whose `key_env` lives only in `.env`
probed `/v1/models` unauthenticated, got 401 and printed a spurious "could not
reach this custom endpoint's model listing" note while chat worked fine.
Resolve through `get_env_prefer_dotenv` — the chain `client_lifecycle` uses for
the real request — when no profile scope is installed. With a scope installed or
multiplexing active the scope stays authoritative: a scoped miss still returns
"" and never borrows another profile's `.env`/process value.
Slimmed from the contributor's two commits (same mechanism, fewer branches,
tests trimmed to two invariants).
Fixes#109315
model_selection_config_updates cleared model.api_key/api only when the target
was not a custom provider, so custom:legacy-box -> custom:other-box (or the
same 'custom' provider on a different base_url) carried endpoint A's inline
secret to endpoint B. The old dashboard writer cleared on any provider change.
The inline key now survives only a same-route re-pick (same provider AND same
normalized base_url); an explicitly submitted key is re-added by the dashboard
after this shape is applied, as before.
One `/model --global` produced four config.yaml shapes. CLI wrote
default/provider/base_url/api_mode and cleared the context pin on a route
change; the gateway rewrote the whole `model:` block (whole-file save_config)
and only set api_mode for `custom`; the TUI wrote three keys and never
touched api_mode, so a switch off an Anthropic-wire endpoint left a stale
`api_mode: anthropic_messages` in config; the dashboard main slot had its own
switched-provider logic, wrote `base_url: ""` and always dropped
context_length. ACP `session/set_model` and `POST /api/model/set` accepted
any model string (parse_model_input + detect_provider_for_model) so a model
no catalog knows, or a provider with no credentials, was handed to the
session / persisted and only failed at inference time.
Canonical: `hermes_cli.model_switch.model_selection_config_updates` (the
shape) + `persist_model_selection(result, config_path=None)` (targeted
per-key `atomic_roundtrip_yaml_update` writes, so sibling
`model_slots`/`model_fallback` keys survive; explicit path for the
multiplexed gateway's profile config) + `apply_model_selection` (same shape
applied to an in-memory `model:` dict for callers that save a whole
document). `atomic_roundtrip_yaml_update(value=None)` now REMOVES the key
instead of writing `key: null`, so per-key and whole-document writers land
the same file. Shape = CLI/gateway semantics: default, provider, base_url
(cleared when the target has none), api_mode (cleared when unresolved),
context_length cleared only when `should_clear_context_pin` says the route
identity changed, inline api_key/api cleared for non-custom targets.
Sites -> canonical:
hermes_cli/cli_model_switch_mixin.py::_persist_global_switch -> deleted; _commit_model_switch calls persist_model_selection
hermes_cli/cli_model_switch_mixin.py::_clear_persisted_context_for_model_switch -> deleted (folded into the shape)
gateway/slash_commands_model.py::_persist_model_switch_to_config -> to_thread forwarder: persist_model_selection(result, ctx.config_path)
tui_gateway/model_switch.py::_persist_model_switch -> deleted; _apply_model_switch calls persist_model_selection
hermes_cli/web_server_config.py::_apply_main_model_assignment -> apply_model_selection(result) (+ explicit custom api_key)
hermes_cli/web_server_config.py::_validated_main_model_selection -> NEW: switch_model(--provider) gate; rejection -> HTTP 400
hermes_cli/web_routers/{models,profiles,config_env}.py main-slot paths -> through _validated_main_model_selection
acp_adapter/server.py::_resolve_model_selection -> deleted; _switch_model calls switch_model (provider:model -> --provider), rejection -> ValueError
Behavior changes: TUI --global now writes/clears model.api_mode and clears a
route-changed context pin; gateway --global no longer rewrites the whole
model block (sibling keys survive) and clears api_mode for every target;
dashboard main slot / profile-create model / custom-endpoint activate now
reject unknown/uncredentialed/unlisted models (HTTP 400) and persist the
resolved base_url/api_mode instead of `base_url: ""`; ACP rejects the same
(ValueError surfaced by the command/protocol handler). Gateway persist runs
on a worker thread against the routed profile's config_path (multiplex-safe).
Cleared keys are removed from config.yaml rather than left as `null`. ACP
still never persists.
Kept `_normalize_main_model_assignment`: switch_model rejects a vendor name
posing as a provider (`moonshotai` -> "Unknown provider"), so the
vendor->aggregator repair is not a duplicate; E2E verified both branches.
No config migration: readers already coalesce `base_url: ""` to absent
(`_config_base_url_for_provider`) and gate api_mode on provider match
(`_provider_supports_explicit_api_mode`), so no stale-shape reader bug.
Tests: tests/hermes_cli/test_model_persist_one_shape.py (four surfaces land
one block; same-route re-pick keeps the pin), tests/acp_adapter/
test_acp_dashboard_model_switch_validation.py (rejection + explicit
provider prefix). Replaces test_acp_set_model_explicit_provider.py and the
two TUI-only persist tests; tests that intercepted the old per-surface seams
(`cli.save_config_value`, `load_config_readonly`, `tui_gateway.server.
_persist_model_switch`) now intercept the canonical seam. Each fix
sabotage-verified red.
A user who picked `deepseek-v4.1-flash` on their own custom endpoint kept
landing on `deepseek-v4-flash-0731`. Three sites each "helped" by diffing
the pick against a catalog and moving it:
- hermes_cli/models_validate.py: the shared catalog matcher auto-corrected
any id within difflib ratio 0.9 of a listed one (`corrected_model`), and
model_switch applied it. Version bumps, dated snapshots and qualifiers
all sit inside 0.9 of a sibling, so a newer release the listing lacked
was swapped for the older one under the user's label. The matcher now
does exact membership -> suggestion text only; the id goes to the wire
verbatim and a genuine typo is refused with the listed siblings named.
Every branch that carried the correction (live listing, static catalog,
curated fallback, MiniMax, Anthropic, custom, OpenRouter preset base)
loses it in one place.
- hermes_cli/model_switch.py: a `providers.<key>` endpoint reached by its
bare key (the slug Desktop picker rows carry) validated as a built-in
and hit the hard-rejecting live-listing branch; the same endpoint as
`custom:<key>` soft-accepted. Both spellings now validate as the user's
custom endpoint.
- apps/desktop: `manualPickRemoved` (composer reseed) and
`reconcileSelectionAfterCatalogRefresh` (Refresh Models) retargeted a
sticky pick to the profile default / the row's first model whenever the
provider row did not list it. Rows are hints (discovered, curated,
capped); the gateway's switch result is the only authority on a pick.
Both helpers are removed; the pick stays put.
Tests: change-detectors pinning the swap are rewritten as invariants
(never `corrected_model`; unlisted id on a user endpoint is kept and
warned; typo is refused with a suggestion); proven red on origin/main.
Step c converted `vendor:model` to `vendor/model` only while the current
provider was an aggregator. On a direct provider (`alibaba`),
`/model Alibaba:qwen3.6-plus` skipped the conversion and went into the
catalog lookup as an unknown id, while `Alibaba/qwen3.6-plus` worked.
Convert on any provider when the left side names a provider Hermes knows
(built-in id/alias or a configured `providers:` entry). Ollama-style tags
(`qwen3.5:4b`) have no provider on the left and stay intact; aggregators
keep the unconditional conversion.
Fixes#9748
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop
* feat(gateway): /signin signs the free tier into a Nous account from a DM
* feat(cli): chat surfaces name /signin as the sign-in verb
* fix(auth): review follow-ups for the shared sign-in flow and /signin
* fix(i18n): carry the /status free-tier line in every locale catalog
* refactor(cli): the chat sign-in command is /login
* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping
A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.
* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin
* docs(user-guide): free tier and signing in
New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.
* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account
Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.
* fix(model): switching off the free tier points at signing in, never hops providers
* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once
* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token
* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check
Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.
* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors
The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.
* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off
* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check
The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.
* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure
A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.
* fix(auth): decide the Nous model together with the route on every credential-pool swap
The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.
* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start
* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died
The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.
* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential
The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.
* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only
A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
A bare `custom`/`local` session whose base_url is not the trusted config
`model.base_url` re-resolved credentials from config on a same-provider
switch and fell through to the OpenRouter default: the next turn hit
openrouter.ai with an empty (or the custom) key. Keep the session endpoint
and key when resolution comes back empty or on OpenRouter; a config-backed
custom URL still wins so key/endpoint rotation is not pinned to a stale
session.
Diagnosed by fangliquanflq in #74143 / #71693; this is the minimal form of
that fix at the same-provider credential step.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.