Salvage follow-up to the previous commit (#114397 by @Finn763):
- Codex/Copilot rows went through cached_provider_model_ids directly, so a
cold cache on the non-blocking read path rendered an EMPTY Copilot row
(live repro: copilot:0). Route them through _live_or_curated_ids like
every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
_mark_catalogs_pending: no surface consumes it and it would have needed
a gateway contract regen. Drop the _spawn_background_warm wrapper: the
ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
every endpoint re-ran four chat-completion probes on every
credential-pool load (load_pool("zai") runs several times per picker
open; the reporter's logs show exactly these repeated POSTs). Memoize
the failure in-process for 5 minutes. Copilot already has the same
negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
open + row still renders; explicit refresh still probes) plus one for
the Z.AI negative cache; a rigid test fake gains **kw for the widened
cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.
Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.