Files
hermes-agent/website/docs/reference/model-catalog.md
teknium1 8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00

6.5 KiB

sidebar_position, title, description
sidebar_position title description
11 Model Catalog Remotely-hosted manifest driving curated model picker lists for OpenRouter and Nous Portal.

Model Catalog

Hermes fetches curated model lists for OpenRouter and Nous Portal from a JSON manifest hosted alongside the docs site. This lets maintainers update picker lists without shipping a new hermes-agent release.

When the manifest is unreachable (offline, network blocked, hosting failure), Hermes silently falls back to the in-repo snapshot that ships with the CLI. The manifest never breaks the picker — worst case you see whatever list was bundled with your installed version.

Live manifest URL

https://hermes-agent.nousresearch.com/docs/api/model-catalog.json

Published on every merge to main via the existing deploy-site.yml GitHub Pages pipeline. The source of truth lives in the repo at website/static/api/model-catalog.json.

Schema

{
  "version": 1,
  "updated_at": "2026-04-25T22:00:00Z",
  "metadata": {},
  "providers": {
    "openrouter": {
      "metadata": {},
      "models": [
        {"id": "z-ai/glm-5.2",         "description": "default", "default": true},
        {"id": "moonshotai/kimi-k3",   "description": "recommended", "metadata": {}},
        {"id": "openai/gpt-5.4",       "description": ""}
      ]
    },
    "nous": {
      "metadata": {},
      "models": [
        {"id": "z-ai/glm-5.2", "default": true},
        {"id": "anthropic/claude-opus-4.7"},
        {"id": "moonshotai/kimi-k3"}
      ]
    }
  }
}

Field notes:

  • version — integer schema version. Future schemas bump this; Hermes refuses manifests with versions it doesn't understand and falls back to the hardcoded snapshot.
  • metadata — free-form dict at the manifest, provider, and model level. Any keys. Hermes ignores unknown fields, so you can annotate entries ("tier": "paid", "tags": [...], etc.) without coordinating a schema change.
  • description — OpenRouter-only. Drives picker badge text ("recommended", "free", "default", or empty). Nous Portal doesn't use this.
  • default — exactly one entry per provider may carry "default": true. That model is the silent default: what Hermes lands on when the user never selected a model (GUI onboarding confirm card, provider configured with no model, empty model.default). Read cache-only at runtime (get_default_model_from_cache) so hot resolution paths never hit the network; when no cached manifest exists, Hermes falls back to the in-repo PREFERRED_SILENT_DEFAULT_MODEL constant, which must match the labeled entry. This lets maintainers rotate the silent default without shipping a release. It is deliberately a capable low-cost model, never the priciest flagship.
  • Pricing and context length are NOT in the manifest. Those come from live provider APIs (/v1/models endpoints, models.dev) at fetch time.

Fetch behavior

When What happens
/model or hermes model Fetches if disk cache is stale, else uses cache
Gateway running Background refresh every ttl_minutes (default 20), so the picker never lags the published manifest by more than one window
Disk cache fresh (< TTL) No network hit
Network failure with cache Silent fallback to cache, one log line
Network failure, no cache Silent fallback to in-repo snapshot
Manifest fails schema validation Treated as unreachable

Cache location: ~/.hermes/cache/model_catalog.json.

Per-provider model lists in the GUI picker

The Desktop, TUI and dashboard pickers (model.options) build each provider's row from the disk-cached live catalog (~/.hermes/provider_models_cache.json) or, when nothing is cached yet, the curated list. Opening the picker never waits on a provider's /v1/models probe or on an auth probe: stale or missing catalogs are refreshed in a background thread and land on the next open, so one slow, rate-limited or unreachable provider cannot hold the whole picker on its loading state. Refresh models (or /model --refresh) is the explicit action that busts the cache and probes every provider live.

Config

model_catalog:
  enabled: true
  url: https://hermes-agent.nousresearch.com/docs/api/model-catalog.json
  ttl_minutes: 20
  providers: {}

Set enabled: false to disable remote fetch entirely and always use the in-repo snapshot (this also disables the gateway's background refresh). ttl_minutes sets both the cache lifetime and the gateway refresh cadence; the legacy ttl_hours key is still honoured if you set it explicitly.

Per-provider override URLs

Third parties can self-host their own curation list using the same schema. Point a provider at a custom URL:

model_catalog:
  providers:
    openrouter:
      url: https://example.com/my-openrouter-curation.json

The overriding manifest only needs to populate the provider block(s) it cares about. Other providers continue to resolve against the master URL.

Hiding providers from the picker

excluded_providers lets you hide specific providers from the /model picker even when valid credentials exist. Useful when credentials are present for legacy or testing providers that shouldn't appear in normal use (e.g. an old Copilot or OpenRouter token still cached in auth.json or discovered via the gh CLI).

model_catalog:
  excluded_providers:
    - copilot
    - openrouter
    - openai

The exclusion is matched case-insensitively against every key a provider can surface under — the Hermes id and models.dev id (built-in mapped providers), the overlay pid and resolved Hermes slug (overlay providers), and the canonical slug (canonical providers) — so a single entry like copilot hides the provider regardless of which section emits it. It is honored by every /model picker surface: the gateway interactive/text pickers, the TUI picker, and the interactive hermes model CLI picker. An empty list (or omitting the key) has no effect.

Updating the manifest

Maintainers:

# Re-generate from the in-repo hardcoded lists (keeps manifest in sync after
# editing OPENROUTER_MODELS or _PROVIDER_MODELS["nous"] in hermes_cli/models.py).
python scripts/build_model_catalog.py

Then PR the resulting change to website/static/api/model-catalog.json to main. The docs site auto-deploys on merge and the new manifest is live within a few minutes.

You can also hand-edit the JSON directly for fine-grained metadata changes that don't belong in the in-repo snapshot — the generator script is a convenience, not the single source of truth.