253 Commits

Author SHA1 Message Date
Kyzcreig
d159b6e0c3 fix(cache): cache-parity forks get a derived cache scope on xAI
#109964 made same-model cache-parity forks (background review, /btw) inherit
the parent's resolved cache scope. That is the right call for
content-addressed caches (Anthropic, DeepSeek, Gemini) and for OpenAI's
routing-only prompt_cache_key. xAI is different: x-grok-conv-id /
prompt_cache_key select ONE server-side conversation slot, so a fork of
similar size that diverges early evicts the parent's slot and the parent's
next call reads cold.

Measured on grok-4.7 (xai-oauth Responses), ~160k context, fork of similar
size diverging right after the system prompt, parent's next call after the
fork:
  shared scope (current):   1,152 / 162,239 read (2/2 runs)
  derived scope (this PR): 162,176 / 162,239 read (2/2 runs)
  no fork (control):       177,536 / 177,639 read (2/2 runs)

build_cache_parity_fork tags the fork (_prompt_cache_fork_tag). The
resolver derives "<scope>::<tag>" for tagged agents on slot-keyed routes
only (xai, xai-oauth, api.x.ai, OpenRouter x-ai/grok-*). The inherited
scope is kept everywhere else. The tag is applied outside the memo, so a
mid-run provider fallback re-evaluates. OpenRouter's Grok x-grok-conv-id
now honours a fork scope over the ambient affinity/conversation scope
(chat_completions threads cache_scope_id to the profile). session_id and
transcript identity are unchanged.

(cherry picked from commit 40fbb39d978648c76940cfb7835f77a42d3719b0)
2026-09-27 00:38:34 +05:30
teknium1
3ca79fd771 fix: codex app-server crash text reaches the user when stderr lags the exit (#121467)
Same race as the ACP client: _subprocess_died reads stderr_tail() the moment
is_alive() turns False, before the reader thread has the crash lines, so the user
saw 'exited unexpectedly' with no cause (CI on main: test_crash_mid_item...).
stderr_tail() now joins the reader briefly once the process has exited.
2026-09-26 11:22:06 -07:00
kshitijk4poor
d24aadfdd1 refactor(copilot): one GitHub effort clamp for the profile and main agent
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.

The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
2026-09-26 22:21:47 +05:30
Jash Lee
7a597324a1 fix(copilot): preserve Astra reasoning effort
(cherry picked from commit 108f20c553852246deb0ce86fd5474d5d7992dc1)
2026-09-26 22:21:47 +05:30
kshitijk4poor
c0cb1d7a34 fix(transports): warn once about a dropped prompt_cache_options
Follow-up to the #123857 salvage. request_overrides is static config, so the
drop warning fired on every turn, retry and subagent call. It now fires once
per process. The Astra sanitizer docstring no longer suggests it handles the
key. The test is parametrized per route, and it pins the extra_body escape
hatch that the warning points to.
2026-09-26 22:14:31 +05:30
kokhlo
07da054cf1 fix(transports): drop prompt_cache_options leaking from request_overrides
request_overrides are merged into the top-level Responses.create() kwargs,
where prompt_cache_options has no SDK parameter: the call fails with
TypeError before any request is sent, on every route. Drop the key after
the merge with a warning, matching the wire-boundary sanitizer precedent.
Wire-only fields still reach capable proxies via extra_body.

(cherry picked from commit c8e8227fdde75236246233565f9cb8ed74c6bd9c)
2026-09-26 22:14:31 +05:30
teknium1
18084881d5 fix: bootstrap the hermes-tools MCP server only when it runs as python -m
codex_runtime, codex_app_server and the runtime migration import this module
for HERMES_TOOLS_MCP_SERVER_NAME, so a module-level import of hermes_bootstrap
exported TMPDIR/TMP/TEMP/HERMES_SCRATCH_DIR into every library importer
(gateway.relay's read-only relay_fronted_platforms included; caught by
tests/gateway/relay/test_cold_opt_out.py).
2026-09-23 19:26:00 -07:00
teknium1
ae80cb7260 fix: import hermes_bootstrap first in the compute-host and hermes-tools MCP entry modules
Both run as their own `python -m` processes (the dashboard compute host, the
codex hermes-tools MCP server) and write os.environ while threads resolve
hosts, so they need the same never-free environ guard as the other entry points.
2026-09-23 19:26:00 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
0f3d32ec58 refactor(transports): one registry read per api_mode gate; drop the stream-delta hook
Trim the salvaged fix to the shape main wants:

* Every gate asks ``agent.transports.registered_api_modes()`` directly. The three helper
  spellings (``_has_registered_transport``, ``_registry_knows``, ``is_registered_api_mode``)
  and the ``sys.modules`` peek are gone: no transport module imports ``providers`` or
  ``hermes_cli`` at module level, so a plain import cannot re-enter provider discovery.
* ``hermes_cli/auth.py`` late-registration pass dropped — main already re-syncs plugin
  profiles into ``PROVIDER_REGISTRY`` on every registry miss
  (``auth_plugin_providers.registry_lookup`` / ``sync_plugin_provider_registry``, #102123);
  the probe shows a profile registered after the import-time mirror resolves and reaches
  the wire on base.
* ``ProviderTransport.normalize_stream_delta`` and the streaming-assembler hook dropped —
  legacy ``delta.function_call`` translation is a separate concern from api_mode
  propagation and has no in-tree consumer.
* Tests: 15 gate-by-gate unit tests replaced by two invariants that install a REAL plugin
  under a temp HERMES_HOME and walk profile → determine_api_mode → resolve_runtime_provider
  → agent ladder → delegation resolver (positive: red on origin/main; negative: an
  unregistered mode still degrades to chat_completions).
2026-09-20 10:11:40 -07:00
valerdoskin
ef8cdfc389 fix(transports): accept a provider plugin's own api_mode
A provider plugin ships a transport via `register_transport(api_mode, cls)` and
declares that same string as its profile's `api_mode`. Transports are selected by
that string, but every gate that validates an api_mode compared it against a
closed literal, so a plugin's mode was rejected at each one and rewritten to
`chat_completions`. The plugin's transport was then never selected: no error, no
tool call, the turn silently degraded to prose. `register_transport` was a public
seam with no way through.

Accept a mode when the transport registry knows it, via a new
`agent.transports.registered_api_modes()`, at each gate:

* `agent_init._resolve_api_mode` - the agent's mode ladder;
* `runtime_provider._parse_api_mode` - the config gate;
* `delegate_tool_config` - the delegation resolver;
* `providers.get_provider` - the reverse `TRANSPORT_TO_API_MODE` lookup recorded
  an unknown mode as `openai_chat`, which made `determine_api_mode` report
  `chat_completions` for a provider that has a dialect transport. This one is the
  most deceptive: the other gates already pass, and the transport still is not used.
* `providers.determine_api_mode` - the same table lookup at the other end.

The registry read is deliberately lazy (`sys.modules.get("agent.transports")`,
never an import): this code is reached from `determine_api_mode`, which provider
discovery itself calls while the registry is being populated, and importing the
transport package there re-enters discovery.

Also add `ProviderTransport.normalize_stream_delta()`, the response-side twin of
`convert_messages()`: a provider that streams a tool call on the legacy OpenAI
`delta.function_call` pair instead of indexed `delta.tool_calls` had nowhere to
translate it, and the streaming assembler dropped the call. The default returns
the delta unchanged, so existing transports are untouched; the assembler now asks
the transport instead of hardcoding one provider's shape.

Finally, make plugin-provider registration repeatable. `hermes_cli.auth` registered
plugin profiles once, at import, from a list `hermes_cli.config` had already
partially discovered while importing itself. A profile that sorts LAST in discovery
was absent from that snapshot and never reached `PROVIDER_REGISTRY`, so every
consumer reported it unauthenticated and it silently vanished from the model
picker while working fine from the CLI. `ensure_plugin_providers_registered()` is
now called from `get_auth_status()` and `resolve_provider()`, so a late profile is
picked up instead of staying invisible.

Unregistered modes are still rejected everywhere, and the in-tree literal sets are
unchanged - they are simply no longer the only way in.
2026-09-20 10:11:40 -07:00
teknium1
e84f0a1c5b fix(codex): a codex app-server thread started from scratch is seeded with the session's prior turns
A codex thread that codex hands back via thread/resume already holds the conversation, but a
thread started fresh did not: a session that ran on another provider before /model switched to
openai-codex, a stored thread codex could not resume, or a thread retired mid-session (prompt
composition change, wedged client) answered the first turn blind. The prior user/assistant text,
tool names and tool-result previews (most recent 32K chars) now ride once on
thread/start.developerInstructions after the prompt composition; thread/resume never carries them,
and the recorded composition stays the bare prompt so the seed cannot make the next turn retire the
thread.

Direction from #26081 (first-turn seeding of the Hermes transcript); redone on the extracted
agent/codex_runtime.py path with the system prompt sent once (#115759) instead of duplicated.

Completes #26035 / #74712 (closed by #115759 for the prompt half; this is the history half).
Co-authored-by: LeonSGP43 <154585401+LeonSGP43@users.noreply.github.com>
2026-09-19 20:44:24 -07:00
Ben Awad
ba586ed4b5 fix: let CodexAppServerSession resume a stored codex thread
Add ``resume_thread_id`` to CodexAppServerSession: when set, the first
ensure_started() issues ``thread/resume`` (with the same cwd / personality /
developerInstructions / model params thread/start sends) instead of starting
an empty thread, and verifies codex handed back the requested id. A refused
or mismatched resume raises the typed CodexThreadResumeError once; the next
ensure_started() falls through to ``thread/start`` on the same handshaken
client (initialize now runs once per client, not once per attempt).

Why: the thread id lived in memory only, so every new AIAgent for the same
Hermes session — a later API-server request or the first turn after a
restart — started a fresh codex thread and the model lost its own memory of
the conversation (#100531). The runtime decides the fail-closed policy; this
adapter only speaks the wire contract (verified against codex-cli 0.147.0:
thread/resume{threadId,...} -> result.thread.id; unknown id -> -32600 "no
rollout found"; killed writer -> -32600 "already has an active writer").

Salvaged from #103352 (thread/resume + id cross-fill + mismatch guard).
2026-09-19 12:27:01 -07:00
teknium1
049a62ab3d Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	website/docs/user-guide/features/codex-app-server-runtime.md
2026-09-19 11:23:47 -07:00
Teknium
271cf9e2d5 Merge pull request #115938 from NousResearch/fix/boa-res-R8-openai-native-search
feat(web): openai-native backend lets Codex Responses turns use the server-side web_search built-in (#19320, salvage #107377)
2026-09-19 11:15:31 -07:00
Teknium
1285de7bbc Merge pull request #115880 from NousResearch/fix/boa-res-R4-routing-catalog-azure-text-verbosity
feat(openai): agent.text_verbosity controls Responses answer length (#20203, salvage #20258)
2026-09-19 11:12:52 -07:00
teknium1
484ec8d07d chore: merge origin/main (resolve tests/agent/transports/test_codex_transport.py) 2026-09-19 10:52:11 -07:00
teknium1
349f67778f chore: merge origin/main (resolve agent/transports/codex.py) 2026-09-19 10:51:50 -07:00
teknium1
54c01bc19a chore: merge origin/main (resolve hermes_cli/runtime_provider.py, website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:51:04 -07:00
teknium1
bf6977fe16 chore: merge origin/main (resolve website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:47:48 -07:00
teknium1
b056e1f36e fix(codex): send image attachments natively in app-server turn/start (#51053)
The codex_app_server runtime flattened every rich user turn into one text
item and replaced image parts with a literal "[image attached]" marker, so
screenshots and pasted images never reached the model. The app-server
`turn/start` protocol (schema v2/UserInput) accepts image inputs natively:
{type: text}, {type: image, url} for data:/http URLs and
{type: localImage, path} for local files.

_build_turn_input now maps Hermes content parts onto that list (text parts
stay text; image_url/{url} data or http refs become `image`, bare paths
become `localImage`) and run_turn sends the whole list. submitted_user_text
keeps the text portion only, which is what the wire echoes back, so the
echo-ownership dedup in codex_runtime is unchanged. Plain-string input and
the image-only default prompt behave as before.

Docs: trade-off table row for image attachments under the app-server runtime.
2026-09-19 10:36:18 -07:00
teknium1
1f4fbd5145 fix(codex): alias the tool_search bridge on OpenAI Responses so Codex requests are not rejected
OpenAI Responses (api.openai.com and the ChatGPT Codex backend) now reserves the
``tool_search`` namespace for its native Tool Search. Progressive tool
disclosure advertises a client function literally named ``tool_search``, so
every request failed at validation with HTTP 400 "Function
'tool_search.tool_search' not allowed in reserved namespace 'tool_search'"
before any model output.

The xAI fix (#95003) already renames the bridge to ``hermes_tool_search`` on
the wire and maps it back in normalize_response; apply the same request-local
alias when the endpoint is the Codex backend or an OpenAI host. Other Responses
proxies keep the bare name.
2026-09-19 10:17:32 -07:00
teknium1
22753744fb fix: reject the HTTP-200 router timeout shim in every OpenAI-compatible consumer
Some OpenAI-compatible routers answer an upstream connect timeout with a 200
ChatCompletion whose only assistant text is "Connect timeout, please try again
later." and zero completion tokens. The cherry-picked predicate only guarded
ChatCompletionsTransport.validate_response(); the streaming assembler emitted
the text to live callbacks before that check, the iteration-limit summary
normalized the response directly, and the auxiliary _validate_llm_response()
only checked the message shape — so the router failure surfaced as the answer
(#68396, maintainer keep_open review on #68433).

Share the predicate (is_router_timeout_shim) and apply it where each path reads
the response:
- streaming: hold shim-prefix text in the existing pending-text seam used for
  echoed SSE, and release it at assembly only when the finished response is not
  a shim; the assembled object then fails validate_response and the loop retries
- iteration summary: a shim reads as an empty summary, taking the retry slot
- auxiliary: a shim raises like a malformed response so the fallback chain moves on

Live probe (fake OpenAI-compatible server, first reply shim, real AIAgent, temp
HERMES_HOME): before, non-stream / stream / summary all returned the shim text
(stream also emitted it live); after, all three retry and return the real answer;
control (no shim) still makes one call.
2026-09-19 10:03:43 -07:00
Jakub Wolniewicz
d4a496373d fix(agent): reject router timeout shim responses 2026-09-19 10:03:43 -07:00
teknium1
0def1fb1ea fix(codex): warn once when an explicit reasoning disable has no wire form on the route
#75227 asked for an unsupported configuration to be reported rather than silently
falling back. A route whose vocabulary has no `none` (Astra, xAI) now logs one
warning per model per process from _resolve_reasoning; the request is unchanged.
2026-09-19 09:39:50 -07:00
teknium1
3e74037cf9 fix(codex): send reasoning.effort none explicitly; no reasoning field for chat-era OpenAI models
Two silent wire mistakes on the Responses transport and its auxiliary
adapter shared one root cause: the transport had no notion of what a
model can be told about reasoning, so it always projected the default
`medium` and always omitted a disable.

- `reasoning_effort: none` was dropped from the request. On a reasoning
  model that defaults to medium (gpt-5.6) the model kept thinking: 76
  reasoning tokens against 0 with an explicit `reasoning.effort: none`.
  The disable now goes on the wire as `{"effort": "none"}` wherever the
  route's vocabulary has `none`; an unset effort is the only state that
  omits the field. A route that rejects `none` already trips the
  reasoning-mandatory recovery (warn, drop the disable, retry).
- Chat-era OpenAI models on api.openai.com (gpt-4o, gpt-4.1, -mini,
  fine-tunes) 400 on any `reasoning` key, so every `openai-api` request
  to them failed. `_codex_efforts_for_route` now returns `()` for them
  (new `model_metadata.openai_model_rejects_reasoning`, a denylist so an
  unknown future model keeps its dial) and both the main transport and
  `_CodexCompletionsAdapter` send no reasoning field. Only the exact
  OpenAI origin is judged: a relay serving those ids may translate.

Fixes #75227
Fixes #76255
2026-09-19 09:39:50 -07:00
teknium1
1db043b3dc fix(codex-app-server): fail pending requests when stdout hits EOF
When codex dies and its stdout closes, the reader thread exited without
touching _pending, so an in-flight request() (initialize/thread/start/
turn/start) still rode out its per-call timeout; only close() failed
pending requests. Wrap the read loop in try/finally so EOF or a reader
failure fails them immediately with a transport error; close() afterwards
stays a no-op on the emptied map.

Also drop the dead 'except queue.Full: pass' in _fail_pending_requests:
_dispatch and _fail_pending_requests both pop the slot under _pending_lock
before putting, so each maxsize=1 queue sees at most one put.
2026-09-19 09:36:13 -07:00
teknium1
8dbbc1333f fix(codex-app-server): report close() before the turn loop as 'session closed', not a timeout
When close() lands after turn/start is accepted but before _drive_turn
snapshots the client, the loop broke on the None snapshot and the post-loop
fallback labelled the result 'turn timed out after Ns'. Route the None
snapshot through _subprocess_died so it retires with the same 'session
closed while the turn was in flight' outcome as a mid-loop close().
2026-09-19 09:36:13 -07:00
teknium1
300201c977 fix(codex-app-server): end a turn cleanly when close() or a transport loss races the poll loop
run_turn()/compact_thread() re-read self._client on every poll iteration, so a
close() from another thread (gateway session-expiry watcher) nulled the client
mid-loop and the next call raised AttributeError out of the turn. The loop now
snapshots the client once, and _subprocess_died() treats a closed session like
subprocess death: TurnResult(interrupted=True, should_retire=True).

Transport write failures (CodexAppServerTransportError) now retire the session
at every boundary: turn/start and thread/compact/start via _request_for, the
approval/elicitation response via _drive_turn; steer and interrupt stay
non-fatal because they already catch CodexAppServerError.

Minimal hand-port of the snapshot idea from #87424.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Related: #87422, #83127
2026-09-19 09:36:13 -07:00
fangliquanflq
7c4a37626b fix(codex-app-server): raise CodexAppServerTransportError on write failure and closed-client drains
CodexAppServerClient._send raised a bare RuntimeError when stdin was closed or
the write hit BrokenPipe/EINVAL, so a codex child exiting between the liveness
check and the next JSON-RPC write escaped every session boundary (they catch
CodexAppServerError/TimeoutError only). The write failure is now a
CodexAppServerTransportError (a CodexAppServerError subclass) and request()
drops its pending slot on a failed send. close() drains pending requests with
the same transport error (tagged reply, not a fake server error) so callers
blocked in request() fail fast instead of riding out their own timeout.

Ported from #83129 (transport exception + pending cleanup) onto the close()
drain from #87434; fixes the drain helper to put on the Queue itself.

Related: #83127, #87433
2026-09-19 09:36:13 -07:00
joaomarcos
1b8245b885 fix(codex-app-server): cancel pending requests immediately on close()
CodexAppServerClient.close() never touched self._pending. A thread
blocked in request() waiting on a JSON-RPC reply had no way to know the
transport it depended on had just been torn down — it rode out its own
per-call timeout (up to 30s by default) even though the subprocess was
already dead.

This is the concrete mechanism behind "codex route hangs after session
expiry": gateway/run.py's _session_expiry_watcher can call
AIAgent.close() -> codex_session.close() while a turn is mid-flight
(e.g. blocked in turn/start), and the blocked caller thread would just
sit there for up to 30s instead of failing fast.

Fix: close() now pops everything out of self._pending and delivers a
synthetic JSON-RPC error to each queue, under the same lock discipline
_read_stdout already uses to dispatch real replies — a real reply
landing at the same instant still wins cleanly (queue.Full is caught)
instead of being dropped or corrupting the queue.

Includes a 300-iteration jittered-timing soak test proving every
blocked request fails within ~0ms of close(), not just on one lucky
timing window.

(cherry picked from commit 2e376cf315c9f628cfdf3afd4e331c4a3d97f5ae)
2026-09-19 09:36:13 -07:00
teknium1
c8ecc3db64 fix: drop replayed reasoning_details on every chat-completions route that does not read it
OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning
continuity; every other OpenAI-compatible route either ignores the field or, when
its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole
request with 400/422 once an earlier reasoning turn is in history — wedging the
session after an in-session model switch (#70233). Strip the field from the wire
copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url),
mirror it in the auxiliary wire boundary and the iteration-summary path; state.db
history keeps the field so switching back to OpenRouter/Nous replays it again.
2026-09-19 09:28:45 -07:00
teknium1
e132e69374 fix(codex): match the primary error on '401 unauthorized', not any bare 401 token
The source-aware classifier replaced base's '401 unauthorized' needle with a
\b401\b regex on the primary error, so any unrelated primary message
carrying a bare 401 token (e.g. 'request body exceeded limit by 401 bytes')
produced the re-login hint and retired the session. Restore the base needle
alongside 'unauthorized' (which already covers the issue's case) and drop
the regex and its 're' import; pin the negative case.
2026-09-19 09:27:40 -07:00
cosin2077
48d59f151f fix(codex): plugin 401 noise in app-server stderr no longer masks the real turn error
`_classify_oauth_failure` joined the primary JSON-RPC error and the codex
stderr tail into one haystack and matched broad tokens ("unauthorized",
"401 unauthorized", "oauth"). codex writes independent ChatGPT plugin
prewarm failures ("HTTP 401 Unauthorized") to stderr while the core
JSON-RPC server keeps working, so any unrelated RPC error, timeout or
subprocess exit was rewritten into the `codex login` hint and the real
error plus stderr tail disappeared.

Classify by source: generic 401/unauthorized/oauth text is authoritative
only in the operation's own error; ambient stderr needs a strong
credential signal (invalid_grant, refresh/expired token, no auth profile).
Every call site (turn error, request timeout, dead subprocess, and the
compaction paths that share them) passes stderr by keyword.

Salvaged from #75182 by @cosin2077, hand-applied onto the refactored
session module (call sites collapsed into `_set_classified_error` /
`_request_for` / `_subprocess_died`).

Fixes #75167
2026-09-19 09:27:40 -07:00
teknium1
4d1d3d05a3 fix(codex): drop the message id of Azure-trimmed reasoning turns too
_newest_reasoning_only pops older turns' codex_reasoning_items before the
converter runs, so those turns looked reasoning-free and replayed their
msg_* id with neither the reasoning item nor its rs_* id on the wire — the
exact orphan shape #97427 rejects. The trimmed row now carries a transient
codex_reasoning_trimmed marker and _replay_message_items honours it; the
single-use _turn_has_encrypted_reasoning wrapper is inlined into that
predicate. Test drives ResponsesApiTransport.build_kwargs on an Azure host.
2026-09-19 09:24:57 -07:00
teknium1
fddc8d1a0f chore: stack on #115759 to resolve agent/codex_runtime.py conflict
CodexAppServerSession takes developer_instructions and model/model_provider; thread/start
params = {cwd, personality:'none'} + developerInstructions + modelProvider/model.
_ensure_codex_session passes both; the named-custom-provider test expects personality:'none'.
2026-09-19 03:39:02 -07:00
lyswty
82c77ff9b9 feat(web): add openai-native backend for Codex server-side web_search
Declares OpenAI's provider-executed Responses `web_search` built-in in place of
the client-side `web_search` function, mirroring the existing xAI native-search
path. Selected via `web.search_backend: openai-native`; search-only, so
`web_extract` keeps resolving to its own backend.

The Responses adapter already recognises built-in tool types
(`_RESPONSES_BUILTIN_TOOL_TYPES`) and preflight passes them through, so the only
missing piece was the swap itself plus a provider name for the config to point at.

Gating is deliberate: two-sided (Codex backend AND a selected openai-native
backend) and fail-closed, so a custom OpenAI-compatible endpoint or an
unresolved provider leaves the client tool untouched.
2026-09-19 01:44:49 -07:00
Kevin
8b7caf226f feat(codex): named custom providers work with the codex_app_server runtime
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.

- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
  admits provider="custom" only when `codex_model_provider_id()` finds a
  configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
  and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
  for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
  `.modelProvider` (fields verified against the codex 0.147 app-server
  schema). `_ensure_codex_session` fills them only for custom agents; codex
  resolves base_url/env_key from its own `[model_providers.<name>]`, so the
  Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
  Desktop/TUI agent knows the provider id (it otherwise collapses to
  "custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
  bare-`custom` ineligibility.

Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).

Fixes #75186
2026-09-19 01:04:05 -07:00
Brandon Zarnitz
3910489f78 feat(openai): agent.text_verbosity controls Responses answer length
New config key `agent.text_verbosity` ("" | low | medium | high, default ""
= not sent). When set, the Responses-family transport emits the top-level
`text: {"verbosity": ...}` field so GPT-5-family models can be asked for
terser or fuller final answers independently of reasoning effort (#20203).

Why this shape: the value is parsed once in agent_init._apply_agent_section
(unknown values warn and are ignored, so unset/"" can never flip the
provider default), passed to ResponsesApiTransport.build_kwargs like the
other per-request params, and never reaches chat_completions / Anthropic;
xAI's /responses is skipped the same way service_tier is. An explicit
request_overrides["text"] (e.g. structured-output format) still wins because
overrides merge after it. The Codex preflight whitelist entry that lets the
field through landed in the previous commit.

DEFAULT_CONFIG entry + docs row under Reasoning Effort.

Salvaged from PR #20258 (config key, docs and adapter direction); the
separate agent/text_verbosity.py module, constructor kwarg and the
cli/gateway/cron/tui_gateway plumbing were dropped in favour of the existing
agent-section config seam.
2026-09-19 00:43:00 -07:00
luyifan
2a12555b3c fix(codex_app_server): disable codex's built-in personality on thread/start
Hermes supplies its own agent identity and personality through the system
prompt. Sending personality: "none" on thread/start strips codex's built-in
"# Personality" section from the base instructions (verified against codex
0.147: the section disappears from the outgoing model request and thread/start
is accepted without error), so it can no longer compete with SOUL.md.

Hand-ported from #72106 (its test hunk is folded into the lane's own invariant
test commit).

Fixes #72104
2026-09-18 23:36:33 -07:00
Momentum96
5cf6dcddcc feat(codex_app_server): accept developer_instructions and send them on thread/start
CodexAppServerSession(developer_instructions=...) forwards Hermes' composed
system prompt as thread/start.developerInstructions when non-empty. Codex keeps
its own base instructions and inserts the text as the first developer message
of every model request, so SOUL.md / memory / channel overrides finally reach
the model on this runtime.

Hand-ported from #27998 (the developerInstructions half only; the SOUL-only
baseInstructions hunk is dropped because baseInstructions REPLACES codex's base
tool guidance, and #74712's own probe table shows the `instructions` spelling is
accepted but ignored).

Part of #74712 #26035
2026-09-18 23:36:17 -07:00
Andrew Johnson
2cfeb505bd fix(agent): alias Perplexity-reserved Responses tools 2026-09-18 09:54:13 -07:00
Kyzcreig
92392218bf fix(anthropic): carry stop_details from the Messages response into provider_data
Anthropic reports the reason for a stop_reason=refusal on the message's
stop_details (category + optional explanation), not in a content block.
The pinned SDK (0.87.0) exposes it only as an extra field and its stream
accumulator copies just stop_reason/stop_sequence from message_delta, so
the final snapshot loses it. Capture it while consuming the stream and
surface it as provider_data["stop_details"] so the refusal handler can
log and show it instead of "(no text)".

Cherry-picked from PR #108682 (the transport + adapter hunks only; the
compression half stays with that PR).
2026-09-18 09:52:36 -07:00
teknium1
d7f2644141 fix(codex): profile declarations follow the host and never override OpenAI's own ladder
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes #114249 but has two
edges the transport must keep:

- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
  non-None custom declaration short-circuited the host lookup: a
  ``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
  Router catalog clamp and would send ``max`` to a gateway that 400s on it.
  Resolve by endpoint host first, then by name — the host is the endpoint's
  truth; the config-entry name is only a label (the existing Router test now
  uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
  codex_responses) is still OpenAI: its per-model ladder is known, so skip
  profile declarations on the official origin and the Codex backend.
  ``_is_openai_api_origin`` is the shared exact-host check;
  ``_is_official_openai_responses_route`` reuses it.

Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.

Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
2026-09-18 09:40:31 -07:00
funky-xamarin
2e6ff19ce2 fix: preserve Codex command exit status for replay 2026-09-18 09:34:10 -07:00
teknium1
2992ed4b3a fix(codex): healthy app-server sessions survive long post-tool reasoning silence
After a tool result, `CodexAppServerSession._run_started_turn` armed a
90s wire-silence watchdog that sent turn/interrupt and retired the
session. On large contexts codex legitimately emits no events for
minutes while it reasons after a big tool output, and the app-server
answers RPCs the whole time (the interrupt was acknowledged in ~25ms).
Silence is not evidence of a wedged process, so the watchdog killed
healthy work and surfaced as protocol_violation / crashed workers.

Keep `post_tool_quiet_timeout` as observability only: past the
threshold log one warning per tool result and keep waiting. Retirement
still happens on the two real signals the poll loop already checks
every iteration -- subprocess death (`_subprocess_died`) and the
overall `turn_timeout` deadline -- so a truly wedged codex stays
bounded.

The existing monotonic-clock watchdog test becomes the invariant for
the new behaviour: tool item, silence past the threshold, then
turn/completed -> no interrupt, no retirement, a warning logged.

Fixes #112928
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 16:53:20 -07:00
teknium1
204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
Shenrui Ma
6973f2622d fix(codex): target hermes-tools in Kanban worker overrides
Dispatcher-owned Kanban workers on the codex app-server runtime injected their
HERMES_KANBAN_* scope into `mcp_servers.hermes-mcp.env.*`, but the runtime
migration registers Hermes' MCP callback as `[mcp_servers.hermes-tools]`. The
override therefore materialised a second, env-only server entry that codex
rejects at bootstrap ("invalid transport in `mcp_servers.hermes-mcp`"), so no
worker could initialize. Point the overrides at the entry that actually exists.

Salvaged from #107337 (its 147-line parametrized test file is replaced by two
invariant tests in a follow-up commit). #111711 proposed the identical two lines.

Fixes #111707

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 19:05:29 -07:00
teknium1
9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K
af4a3eba0a fix(agent): skip corrupted Gemini thought signatures 2026-09-15 05:42:35 -07:00