Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
The fast-mode docs list Claude Opus 4.8, Opus 5 and Opus 5.5. The old
"opus-5" substring matched any future Opus 5.x, and the API answers
`speed` on a model it doesn't list with an error.
One exact list in agent/model_metadata.py now backs both the wire gate
(anthropic_adapter) and the /fast toggle (hermes_cli/models.py). It
accepts vendor-prefixed, dotted and dated ids.
_disk_serve_tier owns the ollama clamp now, and get_model_capabilities /
get_model_info already treat config=None as load-it-yourself.
Six config=config pass-through ternaries in agent/models_dev.py whose callees
have no test double (_configured_catalog_provider, _models_dev_id,
_get_provider_models) collapse to a single call; the _cfg_get and
_load_model_overrides ternaries stay because tests replace those functions.
The parallel prefetch classified any cache entry past
_PROVIDER_MODELS_CACHE_TTL as needing a fetch and blocked the picker on
the thread pool until every one returned. But cached_provider_model_ids()
has two non-blocking tiers, not one: past the TTL and inside
_PROVIDER_MODELS_STALE_SERVE_MAX it still returns the cached list
immediately and revalidates off-thread. Prefetching those slugs traded a
non-blocking serial call for a blocking parallel one.
Since _PROVIDER_MODELS_STALE_SERVE_MAX is far longer than
_PROVIDER_MODELS_CACHE_TTL, every picker open more than a TTL after the
previous one paid for it: locally, 26 providers with a cache 60s past TTL
took 3.1-4.3s, bounded only by the slowest provider's round-trip. Gating
on usability instead of freshness brings that to 0.45s.
The gate now keys on the row the serial call actually reads via
_normalized_cache_slug: a bare "ollama" stays its own cache key rather
than folding into "custom", because the local native catalog has its own
TTL and its own empty-is-authoritative rule. And an empty catalog counts
as servable only for ollama inside _OLLAMA_LOCAL_MODELS_CACHE_TTL —
cached_provider_model_ids returns it with no round-trip, so prefetching
it is redundant work the picker waits on. Past that TTL an empty row gets
no stale-serve window and really does block, so it stays in the prefetch.
Entries the serial path genuinely cannot serve — missing, fingerprint
mismatched, or past the stale-serve window — still prefetch in parallel,
so a cold cache is unaffected. refresh=True already skips the prefetch
entirely, so explicit refresh still forces every provider.
python -m pytest tests/hermes_cli/test_model_cache_parallel_prefetch.py -q
19 passed
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit fb8954a37816002bc5db9a592ebc2d8f1af87673)
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
When a provider's live catalog fetch failed, provider_model_ids() degraded to the
curated static list and cached_provider_model_ids() wrote that list to
provider_models_cache.json with a fresh 1h TTL, exactly as if it were the
account's real catalog. The "only non-empty results are cached" guard never
fired because the fallback is non-empty. A transient Copilot outage therefore
replaced an account's 10 enabled models with the 17-model static list on every
picker surface until the TTL lapsed, and a same-credentials restart, re-auth or
Refresh Models could not shake it (#107391).
Mark the curated list as CuratedFallbackModels at the sites that serve it for a
missing live catalog (the Copilot fetcher, the generic profile merge, and the
static tail when a live fetcher declined). The cache layer then treats it as a
placeholder: it never replaces a same-credentials live row (the account's real
catalog is served instead), it is stored flagged with a 60s TTL when there is
nothing better, and it is never served through the stale-while-revalidate
window. A provider with no live source at all is unaffected: its static list is
its catalog and caches as before.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Native `deepseek` is curated-only: `deepseek-flash` / `deepseek-v4-pro`, in that
order. models.dev still indexes the retired `deepseek-v4-flash*` ids, and with
`deepseek` in `_MODELS_DEV_PREFERRED` the registry union re-added them
registry-first whenever the live /models fetch was unavailable. The Desktop
label for `deepseek-flash` now reads "DeepSeek V4.1 Flash".
Trim of the contributor diff: the deepseek-specific branch inside
`_merge_with_models_dev` was unreachable once deepseek left the preferred set,
so it is dropped; the test drives `provider_model_ids("deepseek")` instead of
the helper.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The offline (no key) path serves the curated floor merged with models.dev, and both still list
x-preview-f-free; the setup-flow path goes through merge_profile_catalog with an empty live list.
One helper, _drop_delisted_opencode_models, runs on the FINAL rows of merge_profile_catalog and of
the static/models.dev return in provider_model_ids, so no path can offer a slug the relay 401s.
Adds the offline invariant test.
The Zen relay retired x-preview-f-free (the picker-facing id for Ox Alpha),
but the exclusion set named the wrong slug (ox-alpha-free) and the live-first
merge only filtered the live half, so the curated floor resurrected the
delisted id and the picker kept offering a model that 401s.
- add x-preview-f-free to _OPENCODE_FREE_EXCLUDED_MODELS
- filter the MERGED result (curated floor is merged back in as secondary half)
Closes#115496
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Gate review: two live sockets survived the first commit on the chat /model reply path —
fetch_openrouter_models() re-downloaded the curated catalog once the disk copy passed its
TTL, and probe_custom_providers defaulted to True so every saved custom endpoint with a key
was probed. fetch_openrouter_models gains cache_only (memory → stale disk → in-repo snapshot,
never a socket); list_picker_providers forwards cache_only and the probe_* flags; the gateway
passes the same read-path flags the GUI picker uses. Tests fold into the existing /model
harness files and record every live probe seam instead of only the prefetch entry.
- rewarm_pricing_before_depleted_notice: a failed fetch caches {} for
_FAILED_CATALOG_TTL_SECONDS and the peek reads that as cold, so every
header in that window spawned a thread that read the auth store and hit
the cached {}. Remember when the last warm started and decide inline
until the window passes. Drop the dead try/except around the pure peek.
- _bg_seed now runs under spawn_context_thread: the warm it gained reads
the profile's auth store, so the thread must carry the profile scope.
- _rerun_notice_policy replaces the idiom copied at three sites.
- _is_subscription_billed: the free-tier default filtered on any truthy
billing_mode while _is_model_free keyed on == 'subscription'.
- The no-respawn guard test counts warm calls instead of enumerating
finished threads (which always read 0).
The Nous gateway can bill a catalog row to a subscription the account
holds instead of to credits, and marks such rows with
`billing_mode: "subscription"` on GET /v1/models. A free-tier account
can run them, but the picker locked every row not priced at $0.
Carry the marker into the Nous pricing entry and count it in
`_is_model_free`, which already feeds the tier partition and the
credits-depleted notice. The silent default for a free-tier account
still prefers a genuinely free model.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The profile branch added for #116408 returned ``merge_profile_catalog(...) or []``, so a
built-in API-key provider with a short curated row and a profile declaring no
fallback_models (minimax, minimax-cn, gemini, kilocode, arcee, stepfun, xiaomi) offered
zero models at first-time setup whenever the live catalog was unreachable — the generic
probe it replaced returned the curated list. Fall back to ``curated``, and share the
key-gated probe with the ``/model`` picker (``models.probe_profile_catalog``) so setup
neither sends a keyless request nor diverges from the picker's rows.
`_api_key_provider_model_list` now routes a registered profile through the
same merge the picker uses (`fetch_models()` curated-first with
`fallback_models`; `fallback_models` alone when the fetch returns None or
raises) via a shared `models.merge_profile_catalog`, extracted from
`_profile_live_catalog` so setup and switching cannot drift. The picker
path also treats a raising `fetch_models` override as an empty catalog
instead of dropping to `[]`.
Why: salvage #116437 returned the live list only when it was at least as
long as the curated one and let a raising catalog abort setup, so setup
and `/model` could still offer different rows for the same profile.
Part of #116408
A raising fetch_models() on the external_process branch escaped to
provider_model_ids() outer try and lost the profile fallback_models, unlike the
api_key branch. The admission test mirrored through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead.
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.
- _profile_live_catalog: external_process profiles use fetch_models(),
then fallback_models; every other non-api-key profile returns its
fallback_models instead of None (in-tree ones declare none, so the
built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
picker rows and the authenticated flag are derived.
Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
_explicit_client_kwargs hardcoded copilot-acp for command/args launch
kwargs; out-of-tree external_process plugin providers got no launch path
and failed at client construction. Key on the provider profile's
auth_type instead, so every ACP/subprocess provider launches the same
way (same approach as #111194, folded in with credit).
Also harden _profile_live_catalog: a signature-strict external_process
profile (fetch_models requiring keyword-only api_key/base_url) now falls
back to credential kwargs on TypeError instead of crashing discovery.
Tests proven red on base for both behaviors.
_profile_live_catalog gated on auth_type=='api_key', so even
picker-admitted ACP providers had no live catalog and fell to the
self-named single-model fallback. External-process profiles supply their
own catalog via fetch_models (subprocess-owned); route them through it.
Verified live: kiro-acp lists its real 19 models (claude-opus-5,
gpt-5.6-sol/terra/luna, deepseek-3.2, glm-5, ...) in model.options.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
`detect_provider_for_model` took the FIRST static-catalog hit as the only
guess. `gpt-5.6-luna` (and the rest of the gpt-5.6 family) is listed by both
`openai-api` and `openai-codex`, so a user with a Codex OAuth grant and no
OPENAI_API_KEY was routed to a keyless openai-api on a fresh (`auto`)
session, or — after the credential gate — left on the current provider with
the request silently ignored, while the grant they hold was never
considered.
`_static_catalog_matches` now yields every catalog that lists the slug in
ladder order; `detect_provider_for_model` keeps its existing credential gate
and takes the first sibling the user actually has credentials for. A fresh
session with no usable provider anywhere still fails loudly on the first
guess, and a user holding both keys keeps today's openai-api routing.
Fixes#102775
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
The wizard writes only model.base_url; the fingerprint hashed the env vars alone, so
switching resource under the same key served the previous resource's catalog for the
TTL/SWR window. Fold the config-resolved base URL in, as openai does with effective_base.
Part of #27989.
`provider_model_ids("azure-foundry")` fell through to the static catalog
`_PROVIDER_MODELS["azure-foundry"] = []`, so `/model azure-foundry` showed "0 models"
even on resources with many deployments. The plugin profile ships `base_url=""`
(per-resource), which is exactly why the generic `_profile_live_catalog` never
fires for it.
Register an `azure-foundry` entry in `_PROVIDER_CATALOG_FETCHERS` that resolves
the endpoint and credential through the runtime resolver
(`_resolve_azure_foundry_runtime`: `model.base_url` / `AZURE_FOUNDRY_BASE_URL`,
API key or Entra token-provider callable) and reuses the wizard's
`azure_detect._probe_openai_models` (api-version fallbacks, never raises).
Anthropic-style `/anthropic` routes have no `/models`; the probe fails soft and
the picker keeps the static `[]`.
Reapplied onto the fetcher-table layout from #28006; reformat/bloat stripped.
Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.
`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.
`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
Salvage follow-up to the previous commit (#114397 by @Finn763):
- Codex/Copilot rows went through cached_provider_model_ids directly, so a
cold cache on the non-blocking read path rendered an EMPTY Copilot row
(live repro: copilot:0). Route them through _live_or_curated_ids like
every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
_mark_catalogs_pending: no surface consumes it and it would have needed
a gateway contract regen. Drop the _spawn_background_warm wrapper: the
ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
every endpoint re-ran four chat-completion probes on every
credential-pool load (load_pool("zai") runs several times per picker
open; the reporter's logs show exactly these repeated POSTs). Memoize
the failure in-process for 5 minutes. Copilot already has the same
negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
open + row still renders; explicit refresh still probes) plus one for
the Z.AI negative cache; a rigid test fake gains **kw for the widened
cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.
Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
Opening the desktop model picker could sit on skeleton placeholders for 70s+
because a normal open (refresh=False) ran live provider catalog probes inline:
the serial row pass fetched each stale provider's /v1/models itself, and the
parallel prefetch joined every worker, so one degraded provider (hanging
endpoint, failed auth probe) held the whole response.
A normal open is now a read path:
* cached_provider_model_ids(non_blocking=True) serves the same-credentials
disk entry of any age and refreshes it in a daemon thread; a cold row
returns [] so the row keeps its curated list.
* list_authenticated_providers(non_blocking_catalogs=True) skips the joining
prefetch and reads every row cache-only; build_model_options_payload turns
it on for refresh=False (api-server / dashboard / TUI model.options).
* rows whose catalog is still warming carry catalog_pending, so a GUI can
tell "not resolved yet" from "that is the provider's catalog".
* Ollama Cloud's 8s probe becomes a cached read + background warm; the
loopback LM Studio probe stays (1.5s, cannot be a degraded remote).
* The SWR write now takes the cache lock: the read path spawns one warm per
stale provider, and concurrent load-modify-save dropped rows.
An explicit refresh (Refresh Models) still probes live and blocking.
(cherry picked from commit 0f2c2e7a3631f2c8f3b5140642306eaf2d688682)
`get_api_key_provider_status` pre-populated configured/logged_in/base_url/
key_source purely so the deleted keyless short-circuit could return them;
every one is now recomputed before the return, so build the snapshot once
instead of overwriting four dead placeholders. Same keys, same order, same
values.
The `_OPENCODE_FREE_EXCLUDED_MODELS` comment also still implied an
anonymous path that no longer exists; it now names the two ids it holds
and why.
The keyless free tier is gone, but five comments and two test docstrings
still described its routing rung: the resolve_runtime_provider ladder
docstring listed a step that no longer exists, and both target_model call
sites plus their regression tests explained themselves in terms of a
`*-free` default being routed to the keyless Zen relay. They now state
what the code actually does (the model-keyed rungs pick the relay and
api_mode).
Also drops the two comments that only narrated the removal
(_OPENCODE_FREE_EXCLUDED_MODELS' history, and an orphan note in
auxiliary_client._resolve_api_key_branch).
The keyless OpenCode free tier was the only provider that ever set
`HermesOverlay.keyless`, so the flag and everything keyed off it is now
unreachable: the `keyless=` field on `HermesOverlay` and
`ProviderDescriptor`, the `_overlay_has_creds` early return, both
`_provider_is_keyless` copies (auth.py and inventory.py), the
`get_api_key_provider_status` keyless short-circuit and its
`key_source: "keyless"` placeholder, and the empty
`_KEYLESS_STABLE_CACHE_PROVIDERS` set whose `_credential_fingerprint`
branch could never match.
Dropping the two catalog-derived test exemptions follows: they computed
the empty set.
The removal dropped the constant but _profile_live_catalog() still filters
live-first Zen/Go listings through it (relay advertises delisted *-free slugs
that 400/403 on POST). Without it, any keyed live picker hit a NameError.
Restored with the original contents (ox-alpha-free, deepseek-v4-flash-free)
and an updated comment; extension point kept for future relay delistings.
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:
- drop the opencode-free provider row, aliases (free/opencode_free), model
catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
clear error naming the removal
Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Closes the remaining atoms of #112600.
A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
its siblings still resolved credentials against config's `default`: the CLI
auth-fallback rung, `--resume` credential re-resolution, the gateway
provider-override helper (channel overrides, persisted /model switches,
API-server provider refresh), the gateway fallback chain, the TUI /model
switch-from runtime and ACP agent construction. With a `*-free` default the
OpenCode free-tier rung fired first and a Go-only model was built against
the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
optional `target_model` and the two test stubs of it accept the kwarg.
B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
provider matched by opencode_provider_family, including custom providers
merely named after a family (`opencode-go-bridge`, #85589) whose relay the
user declared explicitly in `providers:`. The family heal now applies to the
built-in canonical providers only; custom prefix-named providers keep their
per-model api_mode routing and /v1 handling. Documented in the providers
guide.
C) Same function: the official-host check uses parsed.hostname (a port no
longer defeats the heal) and only the path is edited, so query/fragment
round-trip instead of being dropped.
Fixes#112600
Routing the native /api/tags probe through cached_fetch_api_models made the CURRENT endpoint's
probe cache-first with the generic 1h TTL (+7d SWR), so a model pulled after the first picker
open stayed invisible for up to an hour across restarts; the built-in `ollama` slug clamps to
_OLLAMA_LOCAL_MODELS_CACHE_TTL in cached_provider_model_ids but this path did not. Pass the
300s native TTL for the native admission.
An authoritative EMPTY native catalog was also persisted with native_catalog:true and served
back through the whole stale window, so an Ollama that was model-less at first open kept an
empty row after models were pulled. Mirror cached_provider_model_ids: empty native rows are
valid only inside the TTL, never stale-served, and not resurrected when the live probe fails
(the caller falls through to the generic /v1/models fallback instead).
The salvaged fix added a 30-line `_heal_opencode_family_path` helper plus six tests
for one behaviour. The relay path per family is a three-entry table
(`_OPENCODE_FAMILY_PATHS`), so the heal is one `re.fullmatch` on `/zen(/go)?(/v1)?`
inside the normalizer that already owns the opencode.ai host check and the
anthropic `/v1` symmetry — no second urlparse, no separate helper.
Tests trimmed to two invariants: the family path follows the resolved provider
(both directions, both api modes, custom proxy and non-/zen path controls) and the
end-to-end `resolve_runtime_provider(requested="opencode-go")` with a Zen-pinned
`model.base_url` — the reporter's path. The mirror runtime test duplicated the
unit parametrization.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
normalize_opencode_base_url() healed only the /v1 suffix, so a model.base_url
pinned to the Zen relay (https://opencode.ai/zen/v1) survived a switch to an
opencode-go model. The two relays serve different model sets, so every request
went to Zen and 401'd ("Model mimo-v2.5 is not supported") with the provider
already correctly resolved to opencode-go.
The family path segment (/zen vs /zen/go) is now healed to the resolved
provider family inside the same normalizer every resolution path already
funnels through (_finalize_base_url, model_switch, custom family providers).
Only /zen-rooted paths on opencode.ai hosts are rewritten; custom
OPENCODE_*_BASE_URL proxies and unrelated paths are left alone.
Refs #112600