Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
Follow-up to the #123592 salvage. The canonical endpoint now comes from the
provider profile, which also covers OpenRouter (absent from PROVIDER_REGISTRY),
so pinning https://openrouter.ai/api/v1 keeps the native catalog too. Both
sides go through normalize_route_base_url, the helper the rest of the route
comparisons use, so scheme/host case and default ports no longer read as a
relay. The provider is normalized once.
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
Addresses two P1 review findings on #121614.
1. `provider_model_ids`: a failed or empty relay probe fell through to the
canonical per-provider fetcher, sending the provider credential to exactly
the vendor host the user routed away from — recreating #121387 on the
failure path. A configured `model.base_url` relay is now TERMINAL for live
catalog egress and degrades to the local curated list instead. The curated
tail is extracted as `_static_catalog` and shared by both paths.
Fetchers that already resolve `model.base_url` themselves and degrade
locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
simple api-key fetchers) are excluded from interception via
`_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
produce a better-merged catalog.
2. `_try_anthropic`: an `explicit_base_url` that failed
`_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
the ambient/canonical host and continuing with the explicit credential —
a silent retarget of authority, not the refusal the PR body claimed. It now
returns unavailable before client construction.
Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
provider_model_ids() probed the provider's canonical host even when
model.base_url pointed the configured provider at a relay or proxy, so
the picker listed the vendor's models and sent the listing request to a
host the user never chose. Discovery now probes the same endpoint
inference uses, and any probe failure falls through to the canonical
fetchers unchanged.
Fixes#121387
The fast-mode docs list Claude Opus 4.8, Opus 5 and Opus 5.5. The old
"opus-5" substring matched any future Opus 5.x, and the API answers
`speed` on a model it doesn't list with an error.
One exact list in agent/model_metadata.py now backs both the wire gate
(anthropic_adapter) and the /fast toggle (hermes_cli/models.py). It
accepts vendor-prefixed, dotted and dated ids.
_disk_serve_tier owns the ollama clamp now, and get_model_capabilities /
get_model_info already treat config=None as load-it-yourself.
Six config=config pass-through ternaries in agent/models_dev.py whose callees
have no test double (_configured_catalog_provider, _models_dev_id,
_get_provider_models) collapse to a single call; the _cfg_get and
_load_model_overrides ternaries stay because tests replace those functions.
The parallel prefetch classified any cache entry past
_PROVIDER_MODELS_CACHE_TTL as needing a fetch and blocked the picker on
the thread pool until every one returned. But cached_provider_model_ids()
has two non-blocking tiers, not one: past the TTL and inside
_PROVIDER_MODELS_STALE_SERVE_MAX it still returns the cached list
immediately and revalidates off-thread. Prefetching those slugs traded a
non-blocking serial call for a blocking parallel one.
Since _PROVIDER_MODELS_STALE_SERVE_MAX is far longer than
_PROVIDER_MODELS_CACHE_TTL, every picker open more than a TTL after the
previous one paid for it: locally, 26 providers with a cache 60s past TTL
took 3.1-4.3s, bounded only by the slowest provider's round-trip. Gating
on usability instead of freshness brings that to 0.45s.
The gate now keys on the row the serial call actually reads via
_normalized_cache_slug: a bare "ollama" stays its own cache key rather
than folding into "custom", because the local native catalog has its own
TTL and its own empty-is-authoritative rule. And an empty catalog counts
as servable only for ollama inside _OLLAMA_LOCAL_MODELS_CACHE_TTL —
cached_provider_model_ids returns it with no round-trip, so prefetching
it is redundant work the picker waits on. Past that TTL an empty row gets
no stale-serve window and really does block, so it stays in the prefetch.
Entries the serial path genuinely cannot serve — missing, fingerprint
mismatched, or past the stale-serve window — still prefetch in parallel,
so a cold cache is unaffected. refresh=True already skips the prefetch
entirely, so explicit refresh still forces every provider.
python -m pytest tests/hermes_cli/test_model_cache_parallel_prefetch.py -q
19 passed
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit fb8954a37816002bc5db9a592ebc2d8f1af87673)
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
When a provider's live catalog fetch failed, provider_model_ids() degraded to the
curated static list and cached_provider_model_ids() wrote that list to
provider_models_cache.json with a fresh 1h TTL, exactly as if it were the
account's real catalog. The "only non-empty results are cached" guard never
fired because the fallback is non-empty. A transient Copilot outage therefore
replaced an account's 10 enabled models with the 17-model static list on every
picker surface until the TTL lapsed, and a same-credentials restart, re-auth or
Refresh Models could not shake it (#107391).
Mark the curated list as CuratedFallbackModels at the sites that serve it for a
missing live catalog (the Copilot fetcher, the generic profile merge, and the
static tail when a live fetcher declined). The cache layer then treats it as a
placeholder: it never replaces a same-credentials live row (the account's real
catalog is served instead), it is stored flagged with a 60s TTL when there is
nothing better, and it is never served through the stale-while-revalidate
window. A provider with no live source at all is unaffected: its static list is
its catalog and caches as before.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Native `deepseek` is curated-only: `deepseek-flash` / `deepseek-v4-pro`, in that
order. models.dev still indexes the retired `deepseek-v4-flash*` ids, and with
`deepseek` in `_MODELS_DEV_PREFERRED` the registry union re-added them
registry-first whenever the live /models fetch was unavailable. The Desktop
label for `deepseek-flash` now reads "DeepSeek V4.1 Flash".
Trim of the contributor diff: the deepseek-specific branch inside
`_merge_with_models_dev` was unreachable once deepseek left the preferred set,
so it is dropped; the test drives `provider_model_ids("deepseek")` instead of
the helper.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The offline (no key) path serves the curated floor merged with models.dev, and both still list
x-preview-f-free; the setup-flow path goes through merge_profile_catalog with an empty live list.
One helper, _drop_delisted_opencode_models, runs on the FINAL rows of merge_profile_catalog and of
the static/models.dev return in provider_model_ids, so no path can offer a slug the relay 401s.
Adds the offline invariant test.
The Zen relay retired x-preview-f-free (the picker-facing id for Ox Alpha),
but the exclusion set named the wrong slug (ox-alpha-free) and the live-first
merge only filtered the live half, so the curated floor resurrected the
delisted id and the picker kept offering a model that 401s.
- add x-preview-f-free to _OPENCODE_FREE_EXCLUDED_MODELS
- filter the MERGED result (curated floor is merged back in as secondary half)
Closes#115496
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Gate review: two live sockets survived the first commit on the chat /model reply path —
fetch_openrouter_models() re-downloaded the curated catalog once the disk copy passed its
TTL, and probe_custom_providers defaulted to True so every saved custom endpoint with a key
was probed. fetch_openrouter_models gains cache_only (memory → stale disk → in-repo snapshot,
never a socket); list_picker_providers forwards cache_only and the probe_* flags; the gateway
passes the same read-path flags the GUI picker uses. Tests fold into the existing /model
harness files and record every live probe seam instead of only the prefetch entry.
- rewarm_pricing_before_depleted_notice: a failed fetch caches {} for
_FAILED_CATALOG_TTL_SECONDS and the peek reads that as cold, so every
header in that window spawned a thread that read the auth store and hit
the cached {}. Remember when the last warm started and decide inline
until the window passes. Drop the dead try/except around the pure peek.
- _bg_seed now runs under spawn_context_thread: the warm it gained reads
the profile's auth store, so the thread must carry the profile scope.
- _rerun_notice_policy replaces the idiom copied at three sites.
- _is_subscription_billed: the free-tier default filtered on any truthy
billing_mode while _is_model_free keyed on == 'subscription'.
- The no-respawn guard test counts warm calls instead of enumerating
finished threads (which always read 0).
The Nous gateway can bill a catalog row to a subscription the account
holds instead of to credits, and marks such rows with
`billing_mode: "subscription"` on GET /v1/models. A free-tier account
can run them, but the picker locked every row not priced at $0.
Carry the marker into the Nous pricing entry and count it in
`_is_model_free`, which already feeds the tier partition and the
credits-depleted notice. The silent default for a free-tier account
still prefers a genuinely free model.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The profile branch added for #116408 returned ``merge_profile_catalog(...) or []``, so a
built-in API-key provider with a short curated row and a profile declaring no
fallback_models (minimax, minimax-cn, gemini, kilocode, arcee, stepfun, xiaomi) offered
zero models at first-time setup whenever the live catalog was unreachable — the generic
probe it replaced returned the curated list. Fall back to ``curated``, and share the
key-gated probe with the ``/model`` picker (``models.probe_profile_catalog``) so setup
neither sends a keyless request nor diverges from the picker's rows.
`_api_key_provider_model_list` now routes a registered profile through the
same merge the picker uses (`fetch_models()` curated-first with
`fallback_models`; `fallback_models` alone when the fetch returns None or
raises) via a shared `models.merge_profile_catalog`, extracted from
`_profile_live_catalog` so setup and switching cannot drift. The picker
path also treats a raising `fetch_models` override as an empty catalog
instead of dropping to `[]`.
Why: salvage #116437 returned the live list only when it was at least as
long as the curated one and let a raising catalog abort setup, so setup
and `/model` could still offer different rows for the same profile.
Part of #116408
A raising fetch_models() on the external_process branch escaped to
provider_model_ids() outer try and lost the profile fallback_models, unlike the
api_key branch. The admission test mirrored through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead.
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.
- _profile_live_catalog: external_process profiles use fetch_models(),
then fallback_models; every other non-api-key profile returns its
fallback_models instead of None (in-tree ones declare none, so the
built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
picker rows and the authenticated flag are derived.
Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
_explicit_client_kwargs hardcoded copilot-acp for command/args launch
kwargs; out-of-tree external_process plugin providers got no launch path
and failed at client construction. Key on the provider profile's
auth_type instead, so every ACP/subprocess provider launches the same
way (same approach as #111194, folded in with credit).
Also harden _profile_live_catalog: a signature-strict external_process
profile (fetch_models requiring keyword-only api_key/base_url) now falls
back to credential kwargs on TypeError instead of crashing discovery.
Tests proven red on base for both behaviors.
_profile_live_catalog gated on auth_type=='api_key', so even
picker-admitted ACP providers had no live catalog and fell to the
self-named single-model fallback. External-process profiles supply their
own catalog via fetch_models (subprocess-owned); route them through it.
Verified live: kiro-acp lists its real 19 models (claude-opus-5,
gpt-5.6-sol/terra/luna, deepseek-3.2, glm-5, ...) in model.options.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
`detect_provider_for_model` took the FIRST static-catalog hit as the only
guess. `gpt-5.6-luna` (and the rest of the gpt-5.6 family) is listed by both
`openai-api` and `openai-codex`, so a user with a Codex OAuth grant and no
OPENAI_API_KEY was routed to a keyless openai-api on a fresh (`auto`)
session, or — after the credential gate — left on the current provider with
the request silently ignored, while the grant they hold was never
considered.
`_static_catalog_matches` now yields every catalog that lists the slug in
ladder order; `detect_provider_for_model` keeps its existing credential gate
and takes the first sibling the user actually has credentials for. A fresh
session with no usable provider anywhere still fails loudly on the first
guess, and a user holding both keys keeps today's openai-api routing.
Fixes#102775
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
The wizard writes only model.base_url; the fingerprint hashed the env vars alone, so
switching resource under the same key served the previous resource's catalog for the
TTL/SWR window. Fold the config-resolved base URL in, as openai does with effective_base.
Part of #27989.
`provider_model_ids("azure-foundry")` fell through to the static catalog
`_PROVIDER_MODELS["azure-foundry"] = []`, so `/model azure-foundry` showed "0 models"
even on resources with many deployments. The plugin profile ships `base_url=""`
(per-resource), which is exactly why the generic `_profile_live_catalog` never
fires for it.
Register an `azure-foundry` entry in `_PROVIDER_CATALOG_FETCHERS` that resolves
the endpoint and credential through the runtime resolver
(`_resolve_azure_foundry_runtime`: `model.base_url` / `AZURE_FOUNDRY_BASE_URL`,
API key or Entra token-provider callable) and reuses the wizard's
`azure_detect._probe_openai_models` (api-version fallbacks, never raises).
Anthropic-style `/anthropic` routes have no `/models`; the probe fails soft and
the picker keeps the static `[]`.
Reapplied onto the fetcher-table layout from #28006; reformat/bloat stripped.
Co-authored-by: kshitij <82637225+kshitijk4poor@users.noreply.github.com>
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.
`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.
`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
Review findings folded: (1) `parse_model_input` resolved the configured
`custom:<name>` ids through its own `load_config()`, so the `user_providers`/
`custom_providers` the three startup callers pass were ignored for this branch
— two config sources for one decision. It now takes `custom_ids` and the route
builds them from its arguments with `custom_provider_slug`, the same identity
`providers:` entries carry everywhere else. (2) The "typo" guard was wrong: it
fired on legitimate bare-custom tags (`custom:qwen3.5:4b`) and its `None` fell
straight back into the default-provider egress the fix exists to prevent. A
bare `custom` route with an unknown id now 404s on the user's own endpoint,
matching `/model`. Docstring lists the `provider:model` form.
Salvage follow-up to the previous commit (#114397 by @Finn763):
- Codex/Copilot rows went through cached_provider_model_ids directly, so a
cold cache on the non-blocking read path rendered an EMPTY Copilot row
(live repro: copilot:0). Route them through _live_or_curated_ids like
every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
_mark_catalogs_pending: no surface consumes it and it would have needed
a gateway contract regen. Drop the _spawn_background_warm wrapper: the
ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
every endpoint re-ran four chat-completion probes on every
credential-pool load (load_pool("zai") runs several times per picker
open; the reporter's logs show exactly these repeated POSTs). Memoize
the failure in-process for 5 minutes. Copilot already has the same
negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
open + row still renders; explicit refresh still probes) plus one for
the Z.AI negative cache; a rigid test fake gains **kw for the widened
cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.
Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
Opening the desktop model picker could sit on skeleton placeholders for 70s+
because a normal open (refresh=False) ran live provider catalog probes inline:
the serial row pass fetched each stale provider's /v1/models itself, and the
parallel prefetch joined every worker, so one degraded provider (hanging
endpoint, failed auth probe) held the whole response.
A normal open is now a read path:
* cached_provider_model_ids(non_blocking=True) serves the same-credentials
disk entry of any age and refreshes it in a daemon thread; a cold row
returns [] so the row keeps its curated list.
* list_authenticated_providers(non_blocking_catalogs=True) skips the joining
prefetch and reads every row cache-only; build_model_options_payload turns
it on for refresh=False (api-server / dashboard / TUI model.options).
* rows whose catalog is still warming carry catalog_pending, so a GUI can
tell "not resolved yet" from "that is the provider's catalog".
* Ollama Cloud's 8s probe becomes a cached read + background warm; the
loopback LM Studio probe stays (1.5s, cannot be a degraded remote).
* The SWR write now takes the cache lock: the read path spawns one warm per
stale provider, and concurrent load-modify-save dropped rows.
An explicit refresh (Refresh Models) still probes live and blocking.
(cherry picked from commit 0f2c2e7a3631f2c8f3b5140642306eaf2d688682)