get_default_model_for_provider('bedrock') returns the static list's first
entry, so putting Opus 5.5 at [0] silently moved unconfigured Bedrock
users, gateway/API-server runs with no model, and '/model bedrock' from
Sonnet 5 to the priciest flagship. Opus 5.5 now sits second.
The static Bedrock list is what the picker shows when live discovery
(ListFoundationModels + ListInferenceProfiles) is unavailable; it had no
Opus 5.5. Its 1M window comes from the anthropic.claude-opus-5 key in
BEDROCK_CONTEXT_LENGTHS (longest-substring match).
Salvaged from #119444 (catalog hunk only).
(cherry picked from commit 22b7e61f6b)
The native Anthropic curated list had no Opus 5.5 even though the
OpenRouter table in the same file already lists anthropic/claude-opus-5.5,
so the native picker showed it only as a live-discovered extra at the
bottom. Uses the hyphenated id that /v1/models returns.
Salvaged from #119444 (catalog hunk only; the thinking/OAuth/tool_choice
halves are covered by #121016 and #106414).
(cherry picked from commit 757ebf625ae0fee45aae8df29b9d140c16c27a8c)
Mapping the llamacpp aliases to custom in hermes_cli.models sent the
managed local runtime's /model validation down the custom branch before
the staged-library check, so a downloaded-but-not-running GGUF lost its
recognized verdict and an unstaged one was accepted. Keep local and vllm
on custom there, leave the llamacpp aliases on their runtime id, and
move the orphaned 'local' display label to 'custom'.
The parity test now reads the aliases from the custom provider profile
instead of a hand-written list.
Local OpenAI-compatible server aliases (local, vllm, llamacpp, llama.cpp,
llama-cpp) were normalized inconsistently across the three provider-name
tables: hermes_cli.providers mapped them to the orphan "local" id,
hermes_cli.models left them unmapped, while hermes_cli.auth mapped them to
"custom". Align all three on the generic "custom" provider so routing,
the model picker, and credential resolution agree, matching what
resolve_provider already did for vllm/llamacpp. Add a parity contract test.
Two cross-file test leaks in tests/tui_gateway/, found by pairwise bisection after a broad
-k selection went red on pristine main. One was the harness, one was a product bug the
harness had been hiding.
1. hermes_cli.models_catalog_static admitted plugin providers into CANONICAL_PROVIDERS once,
at import. A profile registered after that import — a plugin whose own imports pull
hermes_cli.models in mid-_discover_providers(), or a runtime register_provider() — never
reached list_available_providers / _PROVIDER_LABELS, so the picker, hermes model, /model
and Desktop model.options omitted it until restart. hermes_cli.auth already closes this
window for PROVIDER_REGISTRY (#102123); the catalog snapshot never did. providers.
_sync_auth_registry now re-admits into both snapshots via sync_plugin_provider_catalog(),
idempotent by slug, built-in rows untouched. Exposed as test_auto_continue.py (imports
hermes_cli.models) then test_external_process_picker.py (registers a profile) failing.
2. test_release_resets_every_scope_when_one_reset_fails swaps reset_terminal_scope for one
that raises — that IS the scenario — and left profile B's terminal scope bound on the main
thread for every later test (test_profile_terminal_scope_entrypoints asserts it is None).
The test now unbinds with the saved real reset after the release; the assertion it makes
about the other scopes is unchanged.
Red on main / green here: tests/hermes_cli/test_models_catalog_late_plugin_provider.py, and
the two polluter+victim pairs (3 failed → 30 passed). tests/tui_gateway serial: 2041 passed,
4 failed that pass isolated and in their own files (640/640) — bisected separately.
#119410 registered gpt-6-terra alongside sol and luna from the request text, and the Codex
forward-compat synthesis put it (and -900k) in the live /model picker on every surface.
Nothing serves it: not the Codex account catalog (astra, sol, luna), not OpenRouter (same
three, plus -pro), and OpenAI's model page 404s. Forward-compat is for published tiers the
account catalog has not listed yet, not for guessed names. Removed from the Codex fallback
list and template chain, the context/900k tables, the effort ladder prefixes, the aux-client
family list, and the Nous/OpenRouter static catalogs; catalog JSON regenerated. A contract
test pins the synthesized GPT-6 set to the published tiers.
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.
Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.
/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
Anthropic released Claude Opus 5.5. Both the OpenRouter catalog and the
Nous Portal /v1/models endpoint serve it (verified with a live max_tokens=16
completion on each route — echoed model matches, usage.cost billed).
- hermes_cli/models_catalog_static.py: opus-5.5 in OPENROUTER_MODELS, above
opus-5 and below the fable-5 flagship pair. The nous list is derived from
the OpenRouter tuple, so the Portal picker picks it up from the same edit.
- website/static/api/model-catalog.json: regenerated via
scripts/build_model_catalog.py.
Provider-agnostic metadata already resolves for the new slug — no edits
needed: DEFAULT_CONTEXT_LENGTHS key claude-opus-5 substring-matches to
1,000,000 (matches live OpenRouter metadata), and the claude-opus-5
reasoning-timeout prefix matches through the "." separator to the 240s
floor. Both routes bill via official_models_api live pricing, so no
_OFFICIAL_DOCS_PRICING snapshot entry.
a1b1817adf added grok-4.7 to the generated website/static/api/model-catalog.json
by hand without adding it to _PROVIDER_MODELS["nous"], so
tests/hermes_cli/test_model_catalog.py::test_in_repo_lists_match_manifest has been
red on main since. Pin the id first in the nous list (the manifest already had it
first) and regenerate the manifest with scripts/build_model_catalog.py.
Both routes serve the slug (tools supported, 1,048,576 context, $0.37/$1.25 per M).
The Nous list is derived from OPENROUTER_MODELS so one tuple edit covers both; the docs
manifest is regenerated in the same commit.
DEFAULT_CONTEXT_LENGTHS gets its own key: substring matching would otherwise land the
slug on the glm-5.3-flash entry (1,310,720) and overstate the window by 25%.
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.
- _profile_live_catalog: external_process profiles use fetch_models(),
then fallback_models; every other non-api-key profile returns its
fallback_models instead of None (in-tree ones declare none, so the
built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
picker rows and the authenticated flag are derived.
Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
The CANONICAL_PROVIDERS auto-extend skipped plugin auth_type=
"external_process" providers ("non-api-key flows need bespoke picker UX"),
so plugin ACP backends never reached /model, hermes model, or
list_available_providers — while the in-tree copilot-acp provider of the
same auth_type does, and selection already rides the standard
model_switch pipeline with credentials reported via
auth.get_external_process_provider_status. The skip predates that status
path; admission is now honest (configured = binary resolves) and the
bespoke-UX classes (oauth/device/aws_sdk/copilot/vertex) stay excluded.
Fixes#113865
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.
Part of #95794
The keyless free tier is gone, but five comments and two test docstrings
still described its routing rung: the resolve_runtime_provider ladder
docstring listed a step that no longer exists, and both target_model call
sites plus their regression tests explained themselves in terms of a
`*-free` default being routed to the keyless Zen relay. They now state
what the code actually does (the model-keyed rungs pick the relay and
api_mode).
Also drops the two comments that only narrated the removal
(_OPENCODE_FREE_EXCLUDED_MODELS' history, and an orphan note in
auxiliary_client._resolve_api_key_branch).
The keyless OpenCode free tier was the only provider that ever set
`HermesOverlay.keyless`, so the flag and everything keyed off it is now
unreachable: the `keyless=` field on `HermesOverlay` and
`ProviderDescriptor`, the `_overlay_has_creds` early return, both
`_provider_is_keyless` copies (auth.py and inventory.py), the
`get_api_key_provider_status` keyless short-circuit and its
`key_source: "keyless"` placeholder, and the empty
`_KEYLESS_STABLE_CACHE_PROVIDERS` set whose `_credential_fingerprint`
branch could never match.
Dropping the two catalog-derived test exemptions follows: they computed
the empty set.
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:
- drop the opencode-free provider row, aliases (free/opencode_free), model
catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
clear error naming the removal
Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
Union Alpha is OpenRouter's new $0 stealth model (262,144 ctx, tools +
tool_choice supported, text+image in). It carries no ":free" suffix, so it
needs an explicit "free, stealth model" description and an _OPENROUTER_ONLY
entry to keep it out of the derived Nous Portal list, plus a "union-alpha"
DEFAULT_CONTEXT_LENGTHS key so the offline resolver returns 262144 instead
of the 256K catch-all. Docs manifest regenerated in the same commit.
Live gate: max_tokens=64 tool-calling completion on origin/main key returned
200, echoed model stealth/union-alpha, usage.cost=0, finish_reason=tool_calls.
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).
Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).
- hermes_cli/models_catalog_static.py: remove both slugs from the
opencode-free offline floor and the opencode-zen discovery floor;
document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
invariant to cover both.
The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.
`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:
openai/gpt-5.5:nitro -> 256K (real 1.05M)
x-ai/grok-4.6:nitro -> 131K (generic "grok" catch-all)
anthropic/claude-opus-4.6:nitro-> 200K (generic "claude" catch-all)
The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.
Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.
`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.
The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.
Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.
Builds on YipTszkwan's #107126 (earliest fix in the cluster).
The Go relay (GET /zen/go/v1/models) delisted ox-alpha-free 2026-09-09, but
_PROVIDER_MODELS["opencode-go"] still carried it. _profile_live_catalog merges
the curated floor into the live list (live-first for opencode-go), so the
model picker kept offering a model that now 401s — the same failure class as
#95914 (opencode-free / x-preview-f-free).
Remove it from the floor and add two regression tests: an end-to-end merge
test through provider_model_ids with the real floor (fails if a stale floor
resurrects it) and a floor-pin test asserting the known-delisted model stays
out of the offline fallback.
(cherry picked from commit 091fd85865caa09e828928f93ee985dae7e9580d)
OpenRouter and Nous already ship these ids, but the native Anthropic
curated list still stopped at Fable 5 / Sonnet 5. Live /v1/models often
lags or 401s on subscription tokens, so the picker fell back to that
stale list and hid models that already work when addressed directly.
- Generalize Gemini 3 thinking config model prefix match to gemini-3*
- Add pricing snapshot entries for gemini-3.7-flash and gemini-3.8-flash
- Add model identifiers to Google, OpenRouter, and Vertex CLI lists
- Add test coverage for gemini-3.8-flash thinkingLevel configuration
The bare deepseek/deepseek-v4-flash slug is the pre-snapshot release; both
aggregators now carry deepseek-v4-flash-0731 (and Nous exposes the rolling
~deepseek/deepseek-v4-flash-latest alias). Keeping both rows in the curated
picker just duplicates the flash tier.
- OPENROUTER_MODELS (derives the nous list): drop deepseek/deepseek-v4-flash
- model-catalog.json regenerated
Deliberately KEPT: alibaba-token-plan / opencode-go / commandcode / deepseek
direct plugin fallback_models + default_aux_model (bare id is the wire slug
those providers serve), DEFAULT_CONTEXT_LENGTHS / reasoning floor / pricing
snapshot entries (manually-typed id still behaves), and the
deepseek-chat/deepseek-reasoner -> deepseek-v4-flash alias normalization.
Alibaba shipped qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), the upgraded
snapshot of Qwen3.8-Max: 1M context / 131K output, $2 in / $6 out. Both
aggregators now serve it and dropped the bare slug from their catalogs
(verified 2026-09-02: OpenRouter /v1/models + a 200 completion echoing the
slug; Nous Portal /v1/models).
- OPENROUTER_MODELS (feeds the nous list too): qwen/qwen3.8-max -> qwen/qwen3.8-max-0902
- _ALIBABA_TOKEN_PLAN_MODELS: qwen3.8-max-preview -> qwen3.8-max-0902
- test_empty_model_fallback fixture follows the nous catalog slug
- model-catalog.json regenerated
No new metadata: DEFAULT_CONTEXT_LENGTHS substring-matches qwen3.8-max (1M),
reasoning floor prefix fires (180s), both aggregator routes bill via
official_models_api. alibaba/alibaba-cn/opencode-go/setup.py left unchanged
(out of scope).
Six slugs land in the nous and openrouter curated lists, above the gpt-5.6 line:
openai/gpt-6-astra{,-fast,-flex} and openai/gpt-6-astra-pro{,-fast,-flex}.
Nous Portal serves the tiers as distinct slugs (verified live: each echoes its id, service_tier
default/priority/flex, cost 1x/2x/0.5x). OpenRouter serves them as ENDPOINTS of the base model
(tags openai/fast, openai/flex) and silently routes an unknown suffix to the standard tier at
standard price, so the OpenRouter profile rewrites a tier slug to its base wire model and pins
provider.only to that tier's endpoints (OPENROUTER_ENDPOINT_PINS). The base slug is pinned to
openai/azure/azure-us so default routing never lands on a flex or fast endpoint.
Provider-agnostic metadata: one DEFAULT_CONTEXT_LENGTHS entry (gpt-6-astra: 1,050,000, live on
OpenRouter for both models; substring-matches -pro and the tier suffixes). Pricing is skipped:
both routes bill via official_models_api. Reasoning floor not added (no evidence of long thinks).