Ramp Router is an OpenAI Responses-compatible LLM gateway at https://api.router.com/v1 that routes each request across upstream providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side fallbacks and spend controls. Nous asked for a PR adding it as a provider, so: - plugins/model-providers/router/: RouterProfile plugin — api_mode=codex_responses, RAMP_ROUTER_API_KEY auth, RAMP_ROUTER_BASE_URL override, live account-scoped catalog via GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and Router's docs mandate runtime catalog reads). - hermes_cli/providers.host_mandated_api_mode + runtime_provider._detect_api_mode_for_url: api.router.com -> codex_responses. The host is Responses-only — POST /v1/chat/completions does not exist and 404s — so this is a genuine host mandate (exact hostname match per #32243, mirroring the api.meta.ai precedent). - providers/base.py: new overrideable supported_reasoning_efforts(model) hook (tri-state: None=defer, ()=model takes no reasoning params, tuple=clamp set). Router validates reasoning.effort per model and returns HTTP 400 invalid-argument on levels outside the model's published vocabulary, and 400 unsupported_parameter when a non-reasoning model receives any reasoning field (both verified live). The profile answers from a cached copy of the catalog's router.capabilities.reasoning block: cache-only on the hot path, seeded for free by fetch_models(), disk-mirrored across processes (/cache/router_catalog.json), background-warmed when cold — same design as the OpenRouter reasoning-caps clamp on the chat path. - agent/transports/codex.py: consult the profile-declared vocabulary in the generic effort-clamp branch (xai/actual/github branches untouched; profiles that do not override the hook see no behavior change). - cli-config.yaml.example + adding-providers.md + providers/README.md: document the provider, the host mandate, and the new hook. - tests: behavior contracts for the host mandate/URL detection/spoof rejection, profile registration + auth auto-registry wiring, catalog parsing, and transport clamp/suppression/fallback paths. Verified live against api.router.com (Aug 2026): one-shot chat, streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning replay on OpenAI-served models, function_call_output follow-up turns on OpenAI- and Fireworks-served models; store:false / prompt_cache_key / include:[reasoning.encrypted_content] / reasoning.summary accepted across backends; effort clamp confirmed to convert a would-be 400 (xhigh on o3) into a successful request via the disk mirror.
providers/
Registry and ABC for every inference provider Hermes knows about.
Each provider is declared once as a ProviderProfile. Every other layer —
auth resolution, transport kwargs, model listing, runtime routing — reads from
these profiles instead of maintaining its own parallel data.
Layout
providers/
├── base.py ProviderProfile dataclass + OMIT_TEMPERATURE sentinel
├── __init__.py Registry: register_provider(), get_provider_profile(), list_providers()
└── README.md This file
The profiles themselves live as plugins under
plugins/model-providers/<name>/ (bundled in this repo) and
$HERMES_HOME/plugins/model-providers/<name>/ (per-user overrides). The
registry in providers/__init__.py lazily discovers them the first time any
consumer calls get_provider_profile() or list_providers(). See
plugins/model-providers/README.md for the plugin contract and examples.
How it wires in
The registry is populated on first access. After that, every downstream layer reads from it:
hermes_cli/auth.pyextendsPROVIDER_REGISTRYwith every api-key profile it sees (skippingcopilot,kimi-coding,kimi-coding-cn,zai,openrouter,custom— those need bespoke token resolution).hermes_cli/models.pyextendsCANONICAL_PROVIDERSand callsprofile.fetch_models()insideprovider_model_ids().hermes_cli/doctor.pyadds a/modelshealth check for eachauth_type="api_key"profile.hermes_cli/config.pyinjects everyenv_varintoOPTIONAL_ENV_VARSso the setup wizard knows about it.hermes_cli/runtime_provider.pyreadsprofile.api_modeas a fallback when URL detection finds nothing.agent/model_metadata.pymaps hostname → provider viaprofile.get_hostname().agent/auxiliary_client.pyreadsprofile.default_aux_modelfirst before falling back to the legacy hardcoded dict.agent/transports/chat_completions.py::_build_kwargs_from_profile()invokesprofile.prepare_messages(),profile.build_extra_body(), andprofile.build_api_kwargs_extras()on every call.run_agent.pypassesprovider_profile=<ProviderProfile>so the transport takes the profile path instead of the legacy flag path.
Adding a provider
See plugins/model-providers/README.md — drop a new directory there (or
under $HERMES_HOME/plugins/model-providers/ for a private plugin).
Hooks you can override on ProviderProfile
| Hook | Purpose |
|---|---|
get_hostname() |
URL-based detection — default derives from base_url. |
prepare_messages(msgs) |
Provider-specific message preprocessing (Qwen normalises to list-of-parts, injects cache_control). |
build_extra_body(**ctx) |
Provider-specific extra_body (OpenRouter provider prefs, Gemini thinking_config). |
build_api_kwargs_extras(**ctx) |
(extra_body_additions, top_level_kwargs) — Kimi puts reasoning_effort top-level, Qwen splits enable_thinking/thinking_budget. |
supported_reasoning_efforts(model) |
Declared per-model reasoning-effort vocabulary for gateways that 400 on unknown levels (Ramp Router reads its live catalog). None = defer to transport defaults, () = model takes no reasoning params, tuple = clamp target. Must be cache-only — called on the request hot path. |
fetch_models(*, api_key) |
Live catalog fetch — default hits {models_url or base_url}/models with Bearer auth. Override for no-REST providers (Bedrock), OAuth catalogs (Anthropic), or public catalogs (OpenRouter). |
Configuration fields
Full reference in providers/base.py dataclass definition.