Both new call sites handed the loaded config positionally. The endpoint
doubles in tests/tui_gateway/test_local_model_session_identity.py (and any
keyword-only wrapper) reject a positional argument, the best-effort except
swallowed the TypeError, and llamacpp resolution fell through to "The local
model server is not running" / the anthropic env rung. Use config=... so the
call matches the keyword contract the rest of the tree already uses.
auxiliary.<task>.provider: openai was expanded to custom + the user's OpenAI
endpoint only by agent/auxiliary_client.py (compression, vision, title
generation). hermes_cli/runtime_provider.py::resolve_runtime_provider — the
path background_review, curator, MoA slots and delegation use — had no such
expansion, so "openai" hit auth.resolve_provider's registry lookup and raised
"Unknown provider 'openai'". The alias table now lives once in
runtime_provider_custom.py (the direct-alias/custom sibling) and both paths
call it; resolve_runtime_provider applies it before the ladder.
_host_gated_env_key_candidates also pairs OPENAI_API_KEY with a base_url that
is exactly OPENAI_BASE_URL: the alias lands on that proxy when no block
base_url is set, and the key was issued for it — the host gate otherwise sent
the "no-key-required" placeholder there while the aux-client path used the key.
Slim redo of #116083 (same direction: shared alias, applied in the runtime
resolver) without the extra key gate and effective_provider threading.
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.
- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
admits provider="custom" only when `codex_model_provider_id()` finds a
configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
`.modelProvider` (fields verified against the codex 0.147 app-server
schema). `_ensure_codex_session` fills them only for custom agents; codex
resolves base_url/env_key from its own `[model_providers.<name>]`, so the
Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
Desktop/TUI agent knows the provider id (it otherwise collapses to
"custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
bare-`custom` ineligibility.
Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).
Fixes#75186
`model.provider: custom` + `model.base_url` + `model.key_env` (the wizard's bare-custom
shape) resolved through `_resolve_openrouter_runtime`, whose candidate list only knew
`model.api_key` and the host-gated OPENAI/OPENROUTER env keys — `model.key_env` was never
read, so every request went out as `Bearer no-key-required` and the endpoint returned
401/403 in tens of milliseconds (#67453). The same rung is what the API server platform
resolves on each request, which is why the reporter saw it on every request.
- `_model_cfg_key_env_for()` supplies the declared variable's value on the bare-custom and
direct-alias rungs, only when the target base_url IS the configured `model.base_url`
(a CUSTOM_BASE_URL or alias endpoint elsewhere never receives that key).
- `_key_env_secret()` is the one key_env/api_key_env reader for custom blocks (model block
and custom_providers entries): a declared variable that resolves to nothing is now
WARNING-logged before the `no-key-required` substitution, so a misnamed/unexported var
points at the Hermes config instead of the provider's IAM. Blocks with no key_env stay
silent — that is the keyless local-server configuration.
- `is_output_cap_error()` recognises "max_completion_tokens is limited to N" (Scaleway), so
the budget step-down triggers instead of the compressor (second atom of #67453).
Supersedes #67554 (@JonthanaHanh), which added the same lookup on the explicit-base_url
branch only.
hermes_cli/runtime_provider._getenv was a 4-line copy of get_secret(name,
default) or default; it becomes agent.secret_scope.get_secret_str (returns
default only when the secret is genuinely unset, still raises
UnscopedSecretError — a child's unscoped read is a spawn-site bug). The
runtime_provider_backends/_custom siblings call it directly instead of via
the origin module.
tools/tts_tool, tools/transcription_tools and tools/xai_http each carried an
identical get_env_value re-export kept "so tests can patch" it; the seam is
hermes_cli.config.get_env_value, read lazily at call time. Callers
(tts_streaming, tts_tool_providers, transcription_cloud, voice_client_config,
tools_config) go there directly; resolve_provider_secret already defaults to
it so the env_getter kwarg is gone. Tests repointed at the canonical; the two
tests that only proved the shim forwarded are deleted.
Behavior change: none.
Report a routable provider in session.info instead of the resolved custom
billing class. The desktop carries that identity into new chats without
an endpoint, so losing llamacpp could send a local model to a cloud API.
Recover managed identity from the ownership-checked endpoint through the
existing custom-provider lookup. Bind metadata recovery to the session's
profile, preserving named endpoint precedence, pending selections and
remote compute metadata.
Cover new-chat and resume routing under a cloud default, conflicting
profile mappings, and negative endpoint-ownership cases.
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.