Commit Graph

137 Commits

Author SHA1 Message Date
teknium1
e84f0a1c5b fix(codex): a codex app-server thread started from scratch is seeded with the session's prior turns
A codex thread that codex hands back via thread/resume already holds the conversation, but a
thread started fresh did not: a session that ran on another provider before /model switched to
openai-codex, a stored thread codex could not resume, or a thread retired mid-session (prompt
composition change, wedged client) answered the first turn blind. The prior user/assistant text,
tool names and tool-result previews (most recent 32K chars) now ride once on
thread/start.developerInstructions after the prompt composition; thread/resume never carries them,
and the recorded composition stays the bare prompt so the seed cannot make the next turn retire the
thread.

Direction from #26081 (first-turn seeding of the Hermes transcript); redone on the extracted
agent/codex_runtime.py path with the system prompt sent once (#115759) instead of duplicated.

Completes #26035 / #74712 (closed by #115759 for the prompt half; this is the history half).
Co-authored-by: LeonSGP43 <154585401+LeonSGP43@users.noreply.github.com>
2026-09-19 20:44:24 -07:00
Teknium
b5fcf635dc Merge pull request #116343 from NousResearch/boa-w3-codex-thread
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
2026-09-19 14:31:55 -07:00
teknium1
e77e24a6a6 fix: persist the codex thread id per session and thread/resume it across an API-server restart
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.

Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.

Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
2026-09-19 12:27:01 -07:00
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
teknium1
049a62ab3d Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	website/docs/user-guide/features/codex-app-server-runtime.md
2026-09-19 11:23:47 -07:00
Teknium
b1e7460aea Merge pull request #115900 from NousResearch/fix/boa-res-R2-app-server-supersede
fix(codex): superseded streams keep consuming so final text is not truncated (#69486, salvage #69502)
2026-09-19 11:13:49 -07:00
teknium1
54c01bc19a chore: merge origin/main (resolve hermes_cli/runtime_provider.py, website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:51:04 -07:00
teknium1
dfdeab9e54 chore: merge origin/main (resolve agent/codex_runtime.py, tests/agent/test_run_agent_codex_responses.py) 2026-09-19 10:50:57 -07:00
teknium1
bf6977fe16 chore: merge origin/main (resolve website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:47:48 -07:00
teknium1
68ef35deec fix: drop the redundant _request_is_current() from the first-event stamp guard
_on_event is only ever installed as on_event=_fenced(_on_event), and _fenced
already refuses every call once the request is retired, so the inner predicate
could never be false when reached. Comment now points at the fence.
2026-09-19 10:00:03 -07:00
Ahmett101
34ca29ad73 fix(codex): record first Responses stream event timing
Responses lifecycle events may arrive before answer text, but the streaming runtime never populated the existing post_api_request timing field. Record the first accepted event while preserving request-retirement fencing.\n\nCloses #105311.
2026-09-19 10:00:03 -07:00
teknium1
b3f6834ea4 fix: close the raw provider stream when the post-terminal drain times out under a Relay-managed wrapper
On the Relay-managed path ManagedLlmStream.close() cannot reach the provider
response while the loop is still running the drain (RuntimeError, swallowed by
_close_event_stream), so the httpx response and drain thread lingered until the
socket read timeout. Capture the raw stream in _codex_stream_created and close it
too on drain timeout; the unmanaged path is unchanged (idempotent double close).
2026-09-19 09:59:00 -07:00
teknium1
a8c317d483 fix(codex): make the post-terminal drain budget a config key (agent.stream_drain_timeout)
Replace the hardcoded 2.0s module constant from the salvaged commit with
``agent.stream_drain_timeout`` in DEFAULT_CONFIG (default 2.0, ``0`` skips the
drain), read through ``load_config_readonly`` at drain time. Non-secret knobs
belong in config.yaml, not a new HERMES_* env var. Documents the budget in the
timeouts table and adapts the salvaged test to the new seam.

Why: a relay that never closes the SSE socket after ``response.completed``
wedged the turn until the stale watchdog fired and discarded the already-billed
response (#103864). The drain is a courtesy to Relay's finalizer, so it must be
bounded; the bound needs a user-tunable knob for slow-but-honest relays.
2026-09-19 09:59:00 -07:00
fangliquanflq
39884274d6 fix(codex): bound post-terminal stream drain 2026-09-19 09:59:00 -07:00
teknium1
87de11b700 fix: spill oversized tool outputs before a Codex zero-event reconnect and log the size delta
#95429 acceptance criterion 3: retrying a zero-event oversized request must not
resend the same pathological payload unchanged. run_codex_stream's in-place
transport reconnect re-sent dict(api_kwargs) verbatim with no diagnostic.

When an attempt dies before any stream event was received and the serialized
input exceeds the per-turn budget, inline function_call_output strings over the
per-result threshold (results that escaped commit-time persistence) are spilled
through maybe_persist_tool_result -- bounded preview + recoverable
<persisted-output> reference -- and one warning records
serialized_input_bytes=before -> after with the attempt number. When nothing is
prunable the warning says the payload is being resent unchanged. Only the
retried wire payload changes; the caller's kwargs and history are untouched.
No new config or env vars.

Before: zero-event reconnect resent 760K-char tool output byte-identical, no log.
After: retried input < 1/10 the size, spill file holds the full text, one
"Codex zero-event retry (attempt 1/2): ... serialized_input_bytes=N -> M" line.
2026-09-19 09:47:00 -07:00
briandevans
1c362700b3 fix(codex): stop double-counting cached input in app-server usage
_record_codex_app_server_usage fed inputTokens — which the app-server protocol
reports INCLUSIVE of cachedInputTokens — into CanonicalUsage.input_tokens while
also recording cachedInputTokens as cache_read. CanonicalUsage.prompt_tokens sums
both buckets, so cached tokens were counted twice: inputTokens=80,
cachedInputTokens=20 produced prompt_tokens=100 instead of 80. The inflated value
reached session accounting, cost estimation and context_compressor.last_prompt_tokens
(premature compaction) on every cache-hit turn (#105412, #63654, #48801).

Subtract cache_read from the reported input to get the uncached remainder,
mirroring normalize_usage's codex_responses branch so both Codex transports yield
identical accounting. The transport exposes no cache-write, so only cache_read is
subtracted. totalTokens stays the provider passthrough. The existing integration
test asserted the wrong totals; it now pins prompt=80, uncached=60, cached=20.

Ported from PR #63654 onto the current adapter shape.

Fixes #105412
2026-09-19 09:42:18 -07:00
embwl0x
267ddabb2a feat(codex): configure app-server binary 2026-09-19 09:30:56 -07:00
teknium1
1484f82c26 fix(codex): reset the streamed-text buffer at each completed agentMessage
codex app-server emits commentary deltas + completed, then final deltas +
completed. The delta buffer accumulated "commentary + final", so the completed
final no longer prefix-matched its own deltas, was delivered with
already_streamed=False and the gateway posted a second copy next to the
streamed message (#74248 boundary 2). Drop the buffer once a completed
agentMessage has been compared against its own deltas.
2026-09-19 09:30:22 -07:00
teknium1
fddc8d1a0f chore: stack on #115759 to resolve agent/codex_runtime.py conflict
CodexAppServerSession takes developer_instructions and model/model_provider; thread/start
params = {cwd, personality:'none'} + developerInstructions + modelProvider/model.
_ensure_codex_session passes both; the named-custom-provider test expects personality:'none'.
2026-09-19 03:39:02 -07:00
teknium1
bd7b8bec86 Merge commit 'ca94646b7eb1687e78ff65dfddc6e7f955c5693a' into HEAD 2026-09-19 03:36:30 -07:00
teknium1
7c92677cbb Merge commit '6584242190579b914de9b42a02c9d4b9d1537cb1' into HEAD 2026-09-19 03:36:30 -07:00
teknium1
359a8f25c7 chore: stack on #115817 to resolve agent/codex_runtime.py conflict
writer_token now carries value/raw_stream/superseded_logged; the per-attempt reset clears all three
so the superseded raw stream is still closed (#115817) and the warn-once flag is kept (#115900).
2026-09-19 03:36:30 -07:00
teknium1
ee69433c57 fix(codex_app_server): retire the thread when the composed prompt changes mid-session
_ensure_codex_session records the developerInstructions composition it sent; when the
TUI/Desktop gateway mutates the live agent's prompt in place (/personality, prompt
mirror) the next turn closes the stale thread and starts one carrying the new prompt,
matching what the standard loop applies on its next API call. Sessions attached
without a recorded composition are kept. Docs sentence updated.
2026-09-19 01:23:16 -07:00
teknium1
6584242190 fix: drop the redundant _request_is_current() from the first-event stamp guard
_on_event is only ever installed as on_event=_fenced(_on_event), and _fenced
already refuses every call once the request is retired, so the inner predicate
could never be false when reached. Comment now points at the fence.
2026-09-19 01:17:19 -07:00
teknium1
69562a134e fix: close the raw provider stream when the post-terminal drain times out under a Relay-managed wrapper
On the Relay-managed path ManagedLlmStream.close() cannot reach the provider
response while the loop is still running the drain (RuntimeError, swallowed by
_close_event_stream), so the httpx response and drain thread lingered until the
socket read timeout. Capture the raw stream in _codex_stream_created and close it
too on drain timeout; the unmanaged path is unchanged (idempotent double close).
2026-09-19 01:15:24 -07:00
Kevin
8b7caf226f feat(codex): named custom providers work with the codex_app_server runtime
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.

- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
  admits provider="custom" only when `codex_model_provider_id()` finds a
  configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
  and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
  for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
  `.modelProvider` (fields verified against the codex 0.147 app-server
  schema). `_ensure_codex_session` fills them only for custom agents; codex
  resolves base_url/env_key from its own `[model_providers.<name>]`, so the
  Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
  Desktop/TUI agent knows the provider id (it otherwise collapses to
  "custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
  bare-`custom` ineligibility.

Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).

Fixes #75186
2026-09-19 01:04:05 -07:00
lbo728
09f307d390 fix(codex): superseded streams keep consuming so the final text is not truncated
When a newer attempt claimed the stream-writer slot, run_codex_stream's
accept_chunk hook returned False and Relay stopped consuming the Codex
Responses stream. The assembler then returned a "completed" response
holding only the text seen so far, and the gateway delivered a reply cut
mid-sentence (#69486) even with display streaming disabled.

Supersession now fences only the live callbacks (text and reasoning
deltas, commentary, first-delta) through the attempt's writer token and
consumption continues to the terminal frame, so the assembled final
response is complete. The single-writer invariant still holds: a
superseded attempt never writes into the live display. Request
retirement (watchdog kill) is unchanged and still stops the worker.

Ported from PR #69502 onto the Relay-backed stream path.

Fixes #69486
2026-09-19 00:57:47 -07:00
teknium1
ca94646b7e refactor: inline the pre-stream transport-cause check in run_codex_stream
The OpenAI SDK always raises ``APIConnectionError(...) from err`` with the
httpx transport error as the direct ``__cause__``, so the chain-walking
helper from the salvaged commit is more machinery than the signal needs.
Replace it with an ``isinstance(exc.__cause__, httpx.TransportError)``
check at the single call site and document WHY the raw
``transport_errors`` branch never sees a pre-stream failure (the SDK wraps
them), which is the root cause behind #103673.
2026-09-19 00:00:49 -07:00
Halldrix
37a2ab610b fix(agent): retry pre-stream Codex APIConnectionError caused by transport errors 2026-09-19 00:00:48 -07:00
Ahmett101
bef494b08a fix(codex): record first Responses stream event timing
Responses lifecycle events may arrive before answer text, but the streaming runtime never populated the existing post_api_request timing field. Record the first accepted event while preserving request-retirement fencing.\n\nCloses #105311.
2026-09-18 23:58:13 -07:00
teknium1
b949348c27 fix(codex): make the post-terminal drain budget a config key (agent.stream_drain_timeout)
Replace the hardcoded 2.0s module constant from the salvaged commit with
``agent.stream_drain_timeout`` in DEFAULT_CONFIG (default 2.0, ``0`` skips the
drain), read through ``load_config_readonly`` at drain time. Non-secret knobs
belong in config.yaml, not a new HERMES_* env var. Documents the budget in the
timeouts table and adapts the salvaged test to the new seam.

Why: a relay that never closes the SSE socket after ``response.completed``
wedged the turn until the stale watchdog fired and discarded the already-billed
response (#103864). The drain is a courtesy to Relay's finalizer, so it must be
bounded; the bound needs a user-tunable knob for slow-but-honest relays.
2026-09-18 23:57:54 -07:00
fangliquanflq
87b9570049 fix(codex): bound post-terminal stream drain 2026-09-18 23:54:28 -07:00
RelaxJonh
ab2861fb06 fix(codex_runtime): hand Hermes' composed system prompt to the codex thread
The codex_app_server early-return forked away from the standard loop before any
prompt assembly was used: the codex thread received only cwd and the raw user
text, so SOUL.md, MEMORY.md/USER.md and channel_overrides.system_prompt were
composed, persisted to sessions.system_prompt, and then silently discarded.

_ensure_codex_session now composes the prompt exactly as turn_context does
(agent._cached_system_prompt + "\n\n" + agent.ephemeral_system_prompt) and
passes it to CodexAppServerSession(developer_instructions=...) once per thread.
Sessions are created lazily per turn only when none exists, so the prompt is
never duplicated per turn and a retired session re-sends the current one.

Hand-ported from #74726: same composition, but the session kwarg is spelled
developer_instructions -> thread/start.developerInstructions because the
`instructions` field #74726 used is accepted and ignored by codex.

Fixes #74712
Fixes #26035
2026-09-18 23:36:59 -07:00
teknium1
59e427ff0e test(codex): trim exit-status replay coverage to two invariants
Replace the six parametrized subprocess-driven cases from #111616 with two
invariants: a successful marker-ending output round-trips SessionDB and
replay sanitize as data (envelope preserved), and a real interrupt or an
unknown exit status still recovers as UNKNOWN. The projector consumes the
event dict, so spawning real children added no signal.

Also re-word the live-card payload docstring in agent/codex_runtime.py: it
no longer mirrors the persisted tool-result content byte for byte.
2026-09-18 09:34:10 -07:00
teknium1
9cb1a65978 fix(codex): drop the stale 'watchdog tripped' retirement cause and re-scope the reset test
The post-tool quiet timer no longer retires the session (it only warns), so the
codex_runtime retirement comment listed a cause that cannot happen and the
retained test's name/docstring still promised a watchdog that cannot trip.
Comment and test wording now describe the warning semantics; assertions unchanged.
2026-09-16 16:53:20 -07:00
kshitijk4poor
a12b3c7aa3 refactor(agent): import the transform bypass from its defining module, no re-export shim
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
2026-09-15 19:23:53 -07:00
John Paul Soliva
1e39c93710 perf(agent): keep bulk chat-completions payloads out of the SDK request transform
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.

Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:

    101 msgs /  76 KB   13.6 ms -> 1.2 ms
    401 msgs / 190 KB   48.6 ms -> 2.0 ms
   1601 msgs / 650 KB  188.7 ms -> 5.8 ms

and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.

The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.

Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.

Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.

The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
2026-09-15 19:23:53 -07:00
teknium1
204f345816 refactor(codex): one shared constant for the hermes-tools MCP server name
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.

Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.

Refs #111707
2026-09-15 19:05:29 -07:00
teknium1
cfd752e6f7 fix(sessions): token-accounting guard stamps the agent's real source; trim salvage
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.

Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.

Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.

Refs #111999
2026-09-15 18:23:07 -07:00
teknium1
88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
Hermes Agent
e27f16365d fix(agent): preserve streamed refusals as text (port of anomalyco/opencode#43343)
A model that declines mid-stream delivers the explanation on the
structured refusal channel (chat_completions delta.refusal; Responses
response.refusal.delta / refusal content parts) and leaves content
empty. The streaming accumulators dropped that channel entirely, so a
streamed refusal assembled into an empty message and fell into the
empty/invalid-response retry loops - burning paid retries reproducing a
deterministic refusal - while the non-streaming path had already fixed
this class in #46013.

- chat_completions streaming: accumulate delta.refusal (incl.
  model_extra), expose message.refusal on the assembled mock response so
  ChatCompletionsTransport.normalize_response applies the existing
  sole-payload -> content_filter promotion; count refusal deltas in the
  zero-chunk guard; carry refusal in the Relay final-response dict.
- Codex Responses stream consumer: collect response.refusal.delta as
  answer text so a refusal-only stream no longer raises 'did not emit a
  terminal response' with zero usable content.
- Responses normalizer: read type=refusal content parts in
  _extract_responses_message_text (attr and dict shapes).

Sabotage-verified: each new test fails with its wiring line disabled.
E2E: refusal-only stream -> terminal content_filter with explanation;
refusal-alongside-content stays a normal usable turn; plain-text
streams unchanged.
2026-09-15 03:39:07 -07:00
kshitijk4poor
e625602a67 test(agent): trim stale-kill unwedge tests to the two invariants
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.

Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
2026-09-15 12:46:27 +05:30
Justin Bennington
8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Teknium
a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium
0f4587e336 refactor(compression): every compaction gate asks real usage first; rough estimates only decide whether to wait
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).

Now there is one authority:

- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
  last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
  the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
  over threshold waits ONE request for the provider's real count instead of compressing on a guess
  (first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
  (note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
  past the whole window, and provider-proven overflow all compress immediately; the post-compaction
  latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
  that scripted whole-history estimates now state the fact they relied on (provider omits usage).

Fixes #103391 (closes #103397 by construction — the baseline it repaired no longer exists).
2026-09-06 13:21:17 -07:00
686f6c61
c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
liuhao1024
a9e56c048a fix(compression): anchor codex app-server preflight on reported usage
The codex_app_server runtime bypasses the conversation loop, so the
usage-anchored context accounting captured there never ran: agent._usage_anchor
stayed None forever. Every preflight estimate therefore fell back to the rough
mirror-transcript heuristic, which is deliberately never compacted on this
runtime (_record_codex_app_server_compaction preserves the mirror), so it grows
monotonically while the real thread may be tiny or freshly compacted. With
compression.codex_app_server_auto=hermes that estimate alone tripped the
threshold and fired thread/compact/start on nearly every turn of a long-lived
session (#100381).

Mirror the main loop's post-response capture: _record_codex_app_server_usage
now snapshots the reported thread usage into agent._usage_anchor (base prompt
+ completion exactly as the provider counted them, with estimation confined to
messages appended since). A usage-less turn keeps the previous anchor, and the
post-compaction invalidation site is unchanged.
2026-09-06 13:21:17 -07:00
kshitijk4poor
3ac671dbba refactor(agent): drop the write-only agent-level codex watchdog timestamps
With a request-local watchdog state on every codex request, agent._codex_stream_last_event_ts / _last_progress_ts had no remaining reader; the state-None snapshot only runs for non-codex requests where no watchdog consults it.
2026-09-07 00:41:09 +05:30
kshitijk4poor
1ad8b5dd4a refactor(agent): trim the watchdog phase split to what the tests pin
- Explicit-vs-implicit idle timeout detection uses env_float with a sentinel
  instead of a hand-rolled parse (unset and unparseable both mean implicit).
- _on_event writes the agent-level timestamps only when there is no
  request-local watchdog state; with state present they were dual bookkeeping
  nobody read (the snapshot prefers the state).
- _abort_request keeps main's close-then-retire order: the reorder had no
  test teeth and is a separate concern from #90449.
2026-09-07 00:41:09 +05:30
PeaceMaker-best
463d4bcd15 fix(agent): defer codex idle watchdog until progress
Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
2026-09-07 00:41:09 +05:30