#109964 made same-model cache-parity forks (background review, /btw) inherit
the parent's resolved cache scope. That is the right call for
content-addressed caches (Anthropic, DeepSeek, Gemini) and for OpenAI's
routing-only prompt_cache_key. xAI is different: x-grok-conv-id /
prompt_cache_key select ONE server-side conversation slot, so a fork of
similar size that diverges early evicts the parent's slot and the parent's
next call reads cold.
Measured on grok-4.7 (xai-oauth Responses), ~160k context, fork of similar
size diverging right after the system prompt, parent's next call after the
fork:
shared scope (current): 1,152 / 162,239 read (2/2 runs)
derived scope (this PR): 162,176 / 162,239 read (2/2 runs)
no fork (control): 177,536 / 177,639 read (2/2 runs)
build_cache_parity_fork tags the fork (_prompt_cache_fork_tag). The
resolver derives "<scope>::<tag>" for tagged agents on slot-keyed routes
only (xai, xai-oauth, api.x.ai, OpenRouter x-ai/grok-*). The inherited
scope is kept everywhere else. The tag is applied outside the memo, so a
mid-run provider fallback re-evaluates. OpenRouter's Grok x-grok-conv-id
now honours a fork scope over the ambient affinity/conversation scope
(chat_completions threads cache_scope_id to the profile). session_id and
transcript identity are unchanged.
(cherry picked from commit 40fbb39d978648c76940cfb7835f77a42d3719b0)
Same race as the ACP client: _subprocess_died reads stderr_tail() the moment
is_alive() turns False, before the reader thread has the crash lines, so the user
saw 'exited unexpectedly' with no cause (CI on main: test_crash_mid_item...).
stderr_tail() now joins the reader briefly once the process has exited.
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Follow-up to the #123857 salvage. request_overrides is static config, so the
drop warning fired on every turn, retry and subagent call. It now fires once
per process. The Astra sanitizer docstring no longer suggests it handles the
key. The test is parametrized per route, and it pins the extra_body escape
hatch that the warning points to.
request_overrides are merged into the top-level Responses.create() kwargs,
where prompt_cache_options has no SDK parameter: the call fails with
TypeError before any request is sent, on every route. Drop the key after
the merge with a warning, matching the wire-boundary sanitizer precedent.
Wire-only fields still reach capable proxies via extra_body.
(cherry picked from commit c8e8227fdde75236246233565f9cb8ed74c6bd9c)
codex_runtime, codex_app_server and the runtime migration import this module
for HERMES_TOOLS_MCP_SERVER_NAME, so a module-level import of hermes_bootstrap
exported TMPDIR/TMP/TEMP/HERMES_SCRATCH_DIR into every library importer
(gateway.relay's read-only relay_fronted_platforms included; caught by
tests/gateway/relay/test_cold_opt_out.py).
Both run as their own `python -m` processes (the dashboard compute host, the
codex hermes-tools MCP server) and write os.environ while threads resolve
hosts, so they need the same never-free environ guard as the other entry points.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
Trim the salvaged fix to the shape main wants:
* Every gate asks ``agent.transports.registered_api_modes()`` directly. The three helper
spellings (``_has_registered_transport``, ``_registry_knows``, ``is_registered_api_mode``)
and the ``sys.modules`` peek are gone: no transport module imports ``providers`` or
``hermes_cli`` at module level, so a plain import cannot re-enter provider discovery.
* ``hermes_cli/auth.py`` late-registration pass dropped — main already re-syncs plugin
profiles into ``PROVIDER_REGISTRY`` on every registry miss
(``auth_plugin_providers.registry_lookup`` / ``sync_plugin_provider_registry``, #102123);
the probe shows a profile registered after the import-time mirror resolves and reaches
the wire on base.
* ``ProviderTransport.normalize_stream_delta`` and the streaming-assembler hook dropped —
legacy ``delta.function_call`` translation is a separate concern from api_mode
propagation and has no in-tree consumer.
* Tests: 15 gate-by-gate unit tests replaced by two invariants that install a REAL plugin
under a temp HERMES_HOME and walk profile → determine_api_mode → resolve_runtime_provider
→ agent ladder → delegation resolver (positive: red on origin/main; negative: an
unregistered mode still degrades to chat_completions).
A provider plugin ships a transport via `register_transport(api_mode, cls)` and
declares that same string as its profile's `api_mode`. Transports are selected by
that string, but every gate that validates an api_mode compared it against a
closed literal, so a plugin's mode was rejected at each one and rewritten to
`chat_completions`. The plugin's transport was then never selected: no error, no
tool call, the turn silently degraded to prose. `register_transport` was a public
seam with no way through.
Accept a mode when the transport registry knows it, via a new
`agent.transports.registered_api_modes()`, at each gate:
* `agent_init._resolve_api_mode` - the agent's mode ladder;
* `runtime_provider._parse_api_mode` - the config gate;
* `delegate_tool_config` - the delegation resolver;
* `providers.get_provider` - the reverse `TRANSPORT_TO_API_MODE` lookup recorded
an unknown mode as `openai_chat`, which made `determine_api_mode` report
`chat_completions` for a provider that has a dialect transport. This one is the
most deceptive: the other gates already pass, and the transport still is not used.
* `providers.determine_api_mode` - the same table lookup at the other end.
The registry read is deliberately lazy (`sys.modules.get("agent.transports")`,
never an import): this code is reached from `determine_api_mode`, which provider
discovery itself calls while the registry is being populated, and importing the
transport package there re-enters discovery.
Also add `ProviderTransport.normalize_stream_delta()`, the response-side twin of
`convert_messages()`: a provider that streams a tool call on the legacy OpenAI
`delta.function_call` pair instead of indexed `delta.tool_calls` had nowhere to
translate it, and the streaming assembler dropped the call. The default returns
the delta unchanged, so existing transports are untouched; the assembler now asks
the transport instead of hardcoding one provider's shape.
Finally, make plugin-provider registration repeatable. `hermes_cli.auth` registered
plugin profiles once, at import, from a list `hermes_cli.config` had already
partially discovered while importing itself. A profile that sorts LAST in discovery
was absent from that snapshot and never reached `PROVIDER_REGISTRY`, so every
consumer reported it unauthenticated and it silently vanished from the model
picker while working fine from the CLI. `ensure_plugin_providers_registered()` is
now called from `get_auth_status()` and `resolve_provider()`, so a late profile is
picked up instead of staying invisible.
Unregistered modes are still rejected everywhere, and the in-tree literal sets are
unchanged - they are simply no longer the only way in.
A codex thread that codex hands back via thread/resume already holds the conversation, but a
thread started fresh did not: a session that ran on another provider before /model switched to
openai-codex, a stored thread codex could not resume, or a thread retired mid-session (prompt
composition change, wedged client) answered the first turn blind. The prior user/assistant text,
tool names and tool-result previews (most recent 32K chars) now ride once on
thread/start.developerInstructions after the prompt composition; thread/resume never carries them,
and the recorded composition stays the bare prompt so the seed cannot make the next turn retire the
thread.
Direction from #26081 (first-turn seeding of the Hermes transcript); redone on the extracted
agent/codex_runtime.py path with the system prompt sent once (#115759) instead of duplicated.
Completes #26035 / #74712 (closed by #115759 for the prompt half; this is the history half).
Co-authored-by: LeonSGP43 <154585401+LeonSGP43@users.noreply.github.com>
Add ``resume_thread_id`` to CodexAppServerSession: when set, the first
ensure_started() issues ``thread/resume`` (with the same cwd / personality /
developerInstructions / model params thread/start sends) instead of starting
an empty thread, and verifies codex handed back the requested id. A refused
or mismatched resume raises the typed CodexThreadResumeError once; the next
ensure_started() falls through to ``thread/start`` on the same handshaken
client (initialize now runs once per client, not once per attempt).
Why: the thread id lived in memory only, so every new AIAgent for the same
Hermes session — a later API-server request or the first turn after a
restart — started a fresh codex thread and the model lost its own memory of
the conversation (#100531). The runtime decides the fail-closed policy; this
adapter only speaks the wire contract (verified against codex-cli 0.147.0:
thread/resume{threadId,...} -> result.thread.id; unknown id -> -32600 "no
rollout found"; killed writer -> -32600 "already has an active writer").
Salvaged from #103352 (thread/resume + id cross-fill + mismatch guard).
The codex_app_server runtime flattened every rich user turn into one text
item and replaced image parts with a literal "[image attached]" marker, so
screenshots and pasted images never reached the model. The app-server
`turn/start` protocol (schema v2/UserInput) accepts image inputs natively:
{type: text}, {type: image, url} for data:/http URLs and
{type: localImage, path} for local files.
_build_turn_input now maps Hermes content parts onto that list (text parts
stay text; image_url/{url} data or http refs become `image`, bare paths
become `localImage`) and run_turn sends the whole list. submitted_user_text
keeps the text portion only, which is what the wire echoes back, so the
echo-ownership dedup in codex_runtime is unchanged. Plain-string input and
the image-only default prompt behave as before.
Docs: trade-off table row for image attachments under the app-server runtime.
OpenAI Responses (api.openai.com and the ChatGPT Codex backend) now reserves the
``tool_search`` namespace for its native Tool Search. Progressive tool
disclosure advertises a client function literally named ``tool_search``, so
every request failed at validation with HTTP 400 "Function
'tool_search.tool_search' not allowed in reserved namespace 'tool_search'"
before any model output.
The xAI fix (#95003) already renames the bridge to ``hermes_tool_search`` on
the wire and maps it back in normalize_response; apply the same request-local
alias when the endpoint is the Codex backend or an OpenAI host. Other Responses
proxies keep the bare name.
Some OpenAI-compatible routers answer an upstream connect timeout with a 200
ChatCompletion whose only assistant text is "Connect timeout, please try again
later." and zero completion tokens. The cherry-picked predicate only guarded
ChatCompletionsTransport.validate_response(); the streaming assembler emitted
the text to live callbacks before that check, the iteration-limit summary
normalized the response directly, and the auxiliary _validate_llm_response()
only checked the message shape — so the router failure surfaced as the answer
(#68396, maintainer keep_open review on #68433).
Share the predicate (is_router_timeout_shim) and apply it where each path reads
the response:
- streaming: hold shim-prefix text in the existing pending-text seam used for
echoed SSE, and release it at assembly only when the finished response is not
a shim; the assembled object then fails validate_response and the loop retries
- iteration summary: a shim reads as an empty summary, taking the retry slot
- auxiliary: a shim raises like a malformed response so the fallback chain moves on
Live probe (fake OpenAI-compatible server, first reply shim, real AIAgent, temp
HERMES_HOME): before, non-stream / stream / summary all returned the shim text
(stream also emitted it live); after, all three retry and return the real answer;
control (no shim) still makes one call.
#75227 asked for an unsupported configuration to be reported rather than silently
falling back. A route whose vocabulary has no `none` (Astra, xAI) now logs one
warning per model per process from _resolve_reasoning; the request is unchanged.
Two silent wire mistakes on the Responses transport and its auxiliary
adapter shared one root cause: the transport had no notion of what a
model can be told about reasoning, so it always projected the default
`medium` and always omitted a disable.
- `reasoning_effort: none` was dropped from the request. On a reasoning
model that defaults to medium (gpt-5.6) the model kept thinking: 76
reasoning tokens against 0 with an explicit `reasoning.effort: none`.
The disable now goes on the wire as `{"effort": "none"}` wherever the
route's vocabulary has `none`; an unset effort is the only state that
omits the field. A route that rejects `none` already trips the
reasoning-mandatory recovery (warn, drop the disable, retry).
- Chat-era OpenAI models on api.openai.com (gpt-4o, gpt-4.1, -mini,
fine-tunes) 400 on any `reasoning` key, so every `openai-api` request
to them failed. `_codex_efforts_for_route` now returns `()` for them
(new `model_metadata.openai_model_rejects_reasoning`, a denylist so an
unknown future model keeps its dial) and both the main transport and
`_CodexCompletionsAdapter` send no reasoning field. Only the exact
OpenAI origin is judged: a relay serving those ids may translate.
Fixes#75227Fixes#76255
When codex dies and its stdout closes, the reader thread exited without
touching _pending, so an in-flight request() (initialize/thread/start/
turn/start) still rode out its per-call timeout; only close() failed
pending requests. Wrap the read loop in try/finally so EOF or a reader
failure fails them immediately with a transport error; close() afterwards
stays a no-op on the emptied map.
Also drop the dead 'except queue.Full: pass' in _fail_pending_requests:
_dispatch and _fail_pending_requests both pop the slot under _pending_lock
before putting, so each maxsize=1 queue sees at most one put.
When close() lands after turn/start is accepted but before _drive_turn
snapshots the client, the loop broke on the None snapshot and the post-loop
fallback labelled the result 'turn timed out after Ns'. Route the None
snapshot through _subprocess_died so it retires with the same 'session
closed while the turn was in flight' outcome as a mid-loop close().
run_turn()/compact_thread() re-read self._client on every poll iteration, so a
close() from another thread (gateway session-expiry watcher) nulled the client
mid-loop and the next call raised AttributeError out of the turn. The loop now
snapshots the client once, and _subprocess_died() treats a closed session like
subprocess death: TurnResult(interrupted=True, should_retire=True).
Transport write failures (CodexAppServerTransportError) now retire the session
at every boundary: turn/start and thread/compact/start via _request_for, the
approval/elicitation response via _drive_turn; steer and interrupt stay
non-fatal because they already catch CodexAppServerError.
Minimal hand-port of the snapshot idea from #87424.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Related: #87422, #83127
CodexAppServerClient._send raised a bare RuntimeError when stdin was closed or
the write hit BrokenPipe/EINVAL, so a codex child exiting between the liveness
check and the next JSON-RPC write escaped every session boundary (they catch
CodexAppServerError/TimeoutError only). The write failure is now a
CodexAppServerTransportError (a CodexAppServerError subclass) and request()
drops its pending slot on a failed send. close() drains pending requests with
the same transport error (tagged reply, not a fake server error) so callers
blocked in request() fail fast instead of riding out their own timeout.
Ported from #83129 (transport exception + pending cleanup) onto the close()
drain from #87434; fixes the drain helper to put on the Queue itself.
Related: #83127, #87433
CodexAppServerClient.close() never touched self._pending. A thread
blocked in request() waiting on a JSON-RPC reply had no way to know the
transport it depended on had just been torn down — it rode out its own
per-call timeout (up to 30s by default) even though the subprocess was
already dead.
This is the concrete mechanism behind "codex route hangs after session
expiry": gateway/run.py's _session_expiry_watcher can call
AIAgent.close() -> codex_session.close() while a turn is mid-flight
(e.g. blocked in turn/start), and the blocked caller thread would just
sit there for up to 30s instead of failing fast.
Fix: close() now pops everything out of self._pending and delivers a
synthetic JSON-RPC error to each queue, under the same lock discipline
_read_stdout already uses to dispatch real replies — a real reply
landing at the same instant still wins cleanly (queue.Full is caught)
instead of being dropped or corrupting the queue.
Includes a 300-iteration jittered-timing soak test proving every
blocked request fails within ~0ms of close(), not just on one lucky
timing window.
(cherry picked from commit 2e376cf315c9f628cfdf3afd4e331c4a3d97f5ae)
OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning
continuity; every other OpenAI-compatible route either ignores the field or, when
its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole
request with 400/422 once an earlier reasoning turn is in history — wedging the
session after an in-session model switch (#70233). Strip the field from the wire
copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url),
mirror it in the auxiliary wire boundary and the iteration-summary path; state.db
history keeps the field so switching back to OpenRouter/Nous replays it again.
The source-aware classifier replaced base's '401 unauthorized' needle with a
\b401\b regex on the primary error, so any unrelated primary message
carrying a bare 401 token (e.g. 'request body exceeded limit by 401 bytes')
produced the re-login hint and retired the session. Restore the base needle
alongside 'unauthorized' (which already covers the issue's case) and drop
the regex and its 're' import; pin the negative case.
`_classify_oauth_failure` joined the primary JSON-RPC error and the codex
stderr tail into one haystack and matched broad tokens ("unauthorized",
"401 unauthorized", "oauth"). codex writes independent ChatGPT plugin
prewarm failures ("HTTP 401 Unauthorized") to stderr while the core
JSON-RPC server keeps working, so any unrelated RPC error, timeout or
subprocess exit was rewritten into the `codex login` hint and the real
error plus stderr tail disappeared.
Classify by source: generic 401/unauthorized/oauth text is authoritative
only in the operation's own error; ambient stderr needs a strong
credential signal (invalid_grant, refresh/expired token, no auth profile).
Every call site (turn error, request timeout, dead subprocess, and the
compaction paths that share them) passes stderr by keyword.
Salvaged from #75182 by @cosin2077, hand-applied onto the refactored
session module (call sites collapsed into `_set_classified_error` /
`_request_for` / `_subprocess_died`).
Fixes#75167
_newest_reasoning_only pops older turns' codex_reasoning_items before the
converter runs, so those turns looked reasoning-free and replayed their
msg_* id with neither the reasoning item nor its rs_* id on the wire — the
exact orphan shape #97427 rejects. The trimmed row now carries a transient
codex_reasoning_trimmed marker and _replay_message_items honours it; the
single-use _turn_has_encrypted_reasoning wrapper is inlined into that
predicate. Test drives ResponsesApiTransport.build_kwargs on an Azure host.
Declares OpenAI's provider-executed Responses `web_search` built-in in place of
the client-side `web_search` function, mirroring the existing xAI native-search
path. Selected via `web.search_backend: openai-native`; search-only, so
`web_extract` keeps resolving to its own backend.
The Responses adapter already recognises built-in tool types
(`_RESPONSES_BUILTIN_TOOL_TYPES`) and preflight passes them through, so the only
missing piece was the swap itself plus a provider name for the config to point at.
Gating is deliberate: two-sided (Codex backend AND a selected openai-native
backend) and fail-closed, so a custom OpenAI-compatible endpoint or an
unresolved provider leaves the client tool untouched.
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.
- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
admits provider="custom" only when `codex_model_provider_id()` finds a
configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
`.modelProvider` (fields verified against the codex 0.147 app-server
schema). `_ensure_codex_session` fills them only for custom agents; codex
resolves base_url/env_key from its own `[model_providers.<name>]`, so the
Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
Desktop/TUI agent knows the provider id (it otherwise collapses to
"custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
bare-`custom` ineligibility.
Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).
Fixes#75186
New config key `agent.text_verbosity` ("" | low | medium | high, default ""
= not sent). When set, the Responses-family transport emits the top-level
`text: {"verbosity": ...}` field so GPT-5-family models can be asked for
terser or fuller final answers independently of reasoning effort (#20203).
Why this shape: the value is parsed once in agent_init._apply_agent_section
(unknown values warn and are ignored, so unset/"" can never flip the
provider default), passed to ResponsesApiTransport.build_kwargs like the
other per-request params, and never reaches chat_completions / Anthropic;
xAI's /responses is skipped the same way service_tier is. An explicit
request_overrides["text"] (e.g. structured-output format) still wins because
overrides merge after it. The Codex preflight whitelist entry that lets the
field through landed in the previous commit.
DEFAULT_CONFIG entry + docs row under Reasoning Effort.
Salvaged from PR #20258 (config key, docs and adapter direction); the
separate agent/text_verbosity.py module, constructor kwarg and the
cli/gateway/cron/tui_gateway plumbing were dropped in favour of the existing
agent-section config seam.
Hermes supplies its own agent identity and personality through the system
prompt. Sending personality: "none" on thread/start strips codex's built-in
"# Personality" section from the base instructions (verified against codex
0.147: the section disappears from the outgoing model request and thread/start
is accepted without error), so it can no longer compete with SOUL.md.
Hand-ported from #72106 (its test hunk is folded into the lane's own invariant
test commit).
Fixes#72104
CodexAppServerSession(developer_instructions=...) forwards Hermes' composed
system prompt as thread/start.developerInstructions when non-empty. Codex keeps
its own base instructions and inserts the text as the first developer message
of every model request, so SOUL.md / memory / channel overrides finally reach
the model on this runtime.
Hand-ported from #27998 (the developerInstructions half only; the SOUL-only
baseInstructions hunk is dropped because baseInstructions REPLACES codex's base
tool guidance, and #74712's own probe table shows the `instructions` spelling is
accepted but ignored).
Part of #74712#26035
Anthropic reports the reason for a stop_reason=refusal on the message's
stop_details (category + optional explanation), not in a content block.
The pinned SDK (0.87.0) exposes it only as an extra field and its stream
accumulator copies just stop_reason/stop_sequence from message_delta, so
the final snapshot loses it. Capture it while consuming the stream and
surface it as provider_data["stop_details"] so the refusal handler can
log and show it instead of "(no text)".
Cherry-picked from PR #108682 (the transport + adapter hunks only; the
compression half stays with that PR).
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes#114249 but has two
edges the transport must keep:
- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
non-None custom declaration short-circuited the host lookup: a
``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
Router catalog clamp and would send ``max`` to a gateway that 400s on it.
Resolve by endpoint host first, then by name — the host is the endpoint's
truth; the config-entry name is only a label (the existing Router test now
uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
codex_responses) is still OpenAI: its per-model ladder is known, so skip
profile declarations on the official origin and the Codex backend.
``_is_openai_api_origin`` is the shared exact-host check;
``_is_official_openai_responses_route`` reuses it.
Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.
Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
After a tool result, `CodexAppServerSession._run_started_turn` armed a
90s wire-silence watchdog that sent turn/interrupt and retired the
session. On large contexts codex legitimately emits no events for
minutes while it reasons after a big tool output, and the app-server
answers RPCs the whole time (the interrupt was acknowledged in ~25ms).
Silence is not evidence of a wedged process, so the watchdog killed
healthy work and surfaced as protocol_violation / crashed workers.
Keep `post_tool_quiet_timeout` as observability only: past the
threshold log one warning per tool result and keep waiting. Retirement
still happens on the two real signals the poll loop already checks
every iteration -- subprocess death (`_subprocess_died`) and the
overall `turn_timeout` deadline -- so a truly wedged codex stays
bounded.
The existing monotonic-clock watchdog test becomes the invariant for
the new behaviour: tool item, silence past the threshold, then
turn/completed -> no interrupt, no retirement, a warning logged.
Fixes#112928
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.
Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.
Refs #111707
Dispatcher-owned Kanban workers on the codex app-server runtime injected their
HERMES_KANBAN_* scope into `mcp_servers.hermes-mcp.env.*`, but the runtime
migration registers Hermes' MCP callback as `[mcp_servers.hermes-tools]`. The
override therefore materialised a second, env-only server entry that codex
rejects at bootstrap ("invalid transport in `mcp_servers.hermes-mcp`"), so no
worker could initialize. Point the overrides at the entry that actually exists.
Salvaged from #107337 (its 147-line parametrized test file is replaced by two
invariant tests in a follow-up commit). #111711 proposed the identical two lines.
Fixes#111707
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.
Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.
Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.