vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
providers/
Registry and ABC for every inference provider Hermes knows about.
Each provider is declared once as a ProviderProfile. Every other layer —
auth resolution, transport kwargs, model listing, runtime routing — reads from
these profiles instead of maintaining its own parallel data.
Layout
providers/
├── base.py ProviderProfile dataclass + OMIT_TEMPERATURE sentinel
├── __init__.py Registry: register_provider(), get_provider_profile(), list_providers()
└── README.md This file
The profiles themselves live as plugins under
plugins/model-providers/<name>/ (bundled in this repo) and
$HERMES_HOME/plugins/model-providers/<name>/ (per-user overrides). The
registry in providers/__init__.py lazily discovers them the first time any
consumer calls get_provider_profile() or list_providers(). See
plugins/model-providers/README.md for the plugin contract and examples.
How it wires in
The registry is populated on first access. After that, every downstream layer reads from it:
hermes_cli/auth.pyextendsPROVIDER_REGISTRYwith every api-key profile it sees (skippingcopilot,kimi-coding,kimi-coding-cn,zai,openrouter,custom— those need bespoke token resolution).hermes_cli/models.pyextendsCANONICAL_PROVIDERSand callsprofile.fetch_models()insideprovider_model_ids().hermes_cli/doctor.pyadds a/modelshealth check for eachauth_type="api_key"profile.hermes_cli/config.pyinjects everyenv_varintoOPTIONAL_ENV_VARSso the setup wizard knows about it.hermes_cli/runtime_provider.pyreadsprofile.api_modeas a fallback when URL detection finds nothing.agent/model_metadata.pymaps hostname → provider viaprofile.get_hostname().agent/auxiliary_client.pyreadsprofile.default_aux_modelfirst before falling back to the legacy hardcoded dict.agent/transports/chat_completions.py::_build_kwargs_from_profile()invokesprofile.prepare_messages(),profile.build_extra_body(), andprofile.build_api_kwargs_extras()on every call.run_agent.pypassesprovider_profile=<ProviderProfile>so the transport takes the profile path instead of the legacy flag path.
Adding a provider
See plugins/model-providers/README.md — drop a new directory there (or
under $HERMES_HOME/plugins/model-providers/ for a private plugin).
Hooks you can override on ProviderProfile
| Hook | Purpose |
|---|---|
get_hostname() |
URL-based detection — default derives from base_url. |
prepare_messages(msgs) |
Provider-specific message preprocessing (Qwen normalises to list-of-parts, injects cache_control). |
build_extra_body(**ctx) |
Provider-specific extra_body (OpenRouter provider prefs, Gemini thinking_config). |
build_api_kwargs_extras(**ctx) |
(extra_body_additions, top_level_kwargs) — Kimi puts reasoning_effort top-level, Qwen splits enable_thinking/thinking_budget. |
supported_reasoning_efforts(model) |
Declared per-model reasoning-effort vocabulary for gateways that 400 on unknown levels (Ramp Router reads its live catalog). None = defer to transport defaults, () = model takes no reasoning params, tuple = clamp target. Must be cache-only — called on the request hot path. |
fetch_models(*, api_key) |
Live catalog fetch — default hits {models_url or base_url}/models with Bearer auth. Override for no-REST providers (Bedrock), OAuth catalogs (Anthropic), or public catalogs (OpenRouter). |
Configuration fields
Full reference in providers/base.py dataclass definition.