A codex thread that codex hands back via thread/resume already holds the conversation, but a
thread started fresh did not: a session that ran on another provider before /model switched to
openai-codex, a stored thread codex could not resume, or a thread retired mid-session (prompt
composition change, wedged client) answered the first turn blind. The prior user/assistant text,
tool names and tool-result previews (most recent 32K chars) now ride once on
thread/start.developerInstructions after the prompt composition; thread/resume never carries them,
and the recorded composition stays the bare prompt so the seed cannot make the next turn retire the
thread.
Direction from #26081 (first-turn seeding of the Hermes transcript); redone on the extracted
agent/codex_runtime.py path with the system prompt sent once (#115759) instead of duplicated.
Completes #26035 / #74712 (closed by #115759 for the prompt half; this is the history half).
Co-authored-by: LeonSGP43 <154585401+LeonSGP43@users.noreply.github.com>
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.
Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.
Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.
One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.
Part of #97548
_on_event is only ever installed as on_event=_fenced(_on_event), and _fenced
already refuses every call once the request is retired, so the inner predicate
could never be false when reached. Comment now points at the fence.
Responses lifecycle events may arrive before answer text, but the streaming runtime never populated the existing post_api_request timing field. Record the first accepted event while preserving request-retirement fencing.\n\nCloses #105311.
On the Relay-managed path ManagedLlmStream.close() cannot reach the provider
response while the loop is still running the drain (RuntimeError, swallowed by
_close_event_stream), so the httpx response and drain thread lingered until the
socket read timeout. Capture the raw stream in _codex_stream_created and close it
too on drain timeout; the unmanaged path is unchanged (idempotent double close).
Replace the hardcoded 2.0s module constant from the salvaged commit with
``agent.stream_drain_timeout`` in DEFAULT_CONFIG (default 2.0, ``0`` skips the
drain), read through ``load_config_readonly`` at drain time. Non-secret knobs
belong in config.yaml, not a new HERMES_* env var. Documents the budget in the
timeouts table and adapts the salvaged test to the new seam.
Why: a relay that never closes the SSE socket after ``response.completed``
wedged the turn until the stale watchdog fired and discarded the already-billed
response (#103864). The drain is a courtesy to Relay's finalizer, so it must be
bounded; the bound needs a user-tunable knob for slow-but-honest relays.
#95429 acceptance criterion 3: retrying a zero-event oversized request must not
resend the same pathological payload unchanged. run_codex_stream's in-place
transport reconnect re-sent dict(api_kwargs) verbatim with no diagnostic.
When an attempt dies before any stream event was received and the serialized
input exceeds the per-turn budget, inline function_call_output strings over the
per-result threshold (results that escaped commit-time persistence) are spilled
through maybe_persist_tool_result -- bounded preview + recoverable
<persisted-output> reference -- and one warning records
serialized_input_bytes=before -> after with the attempt number. When nothing is
prunable the warning says the payload is being resent unchanged. Only the
retried wire payload changes; the caller's kwargs and history are untouched.
No new config or env vars.
Before: zero-event reconnect resent 760K-char tool output byte-identical, no log.
After: retried input < 1/10 the size, spill file holds the full text, one
"Codex zero-event retry (attempt 1/2): ... serialized_input_bytes=N -> M" line.
_record_codex_app_server_usage fed inputTokens — which the app-server protocol
reports INCLUSIVE of cachedInputTokens — into CanonicalUsage.input_tokens while
also recording cachedInputTokens as cache_read. CanonicalUsage.prompt_tokens sums
both buckets, so cached tokens were counted twice: inputTokens=80,
cachedInputTokens=20 produced prompt_tokens=100 instead of 80. The inflated value
reached session accounting, cost estimation and context_compressor.last_prompt_tokens
(premature compaction) on every cache-hit turn (#105412, #63654, #48801).
Subtract cache_read from the reported input to get the uncached remainder,
mirroring normalize_usage's codex_responses branch so both Codex transports yield
identical accounting. The transport exposes no cache-write, so only cache_read is
subtracted. totalTokens stays the provider passthrough. The existing integration
test asserted the wrong totals; it now pins prompt=80, uncached=60, cached=20.
Ported from PR #63654 onto the current adapter shape.
Fixes#105412
codex app-server emits commentary deltas + completed, then final deltas +
completed. The delta buffer accumulated "commentary + final", so the completed
final no longer prefix-matched its own deltas, was delivered with
already_streamed=False and the gateway posted a second copy next to the
streamed message (#74248 boundary 2). Drop the buffer once a completed
agentMessage has been compared against its own deltas.
writer_token now carries value/raw_stream/superseded_logged; the per-attempt reset clears all three
so the superseded raw stream is still closed (#115817) and the warn-once flag is kept (#115900).
_ensure_codex_session records the developerInstructions composition it sent; when the
TUI/Desktop gateway mutates the live agent's prompt in place (/personality, prompt
mirror) the next turn closes the stale thread and starts one carrying the new prompt,
matching what the standard loop applies on its next API call. Sessions attached
without a recorded composition are kept. Docs sentence updated.
_on_event is only ever installed as on_event=_fenced(_on_event), and _fenced
already refuses every call once the request is retired, so the inner predicate
could never be false when reached. Comment now points at the fence.
On the Relay-managed path ManagedLlmStream.close() cannot reach the provider
response while the loop is still running the drain (RuntimeError, swallowed by
_close_event_stream), so the httpx response and drain thread lingered until the
socket read timeout. Capture the raw stream in _codex_stream_created and close it
too on drain timeout; the unmanaged path is unchanged (idempotent double close).
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.
- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
admits provider="custom" only when `codex_model_provider_id()` finds a
configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
`.modelProvider` (fields verified against the codex 0.147 app-server
schema). `_ensure_codex_session` fills them only for custom agents; codex
resolves base_url/env_key from its own `[model_providers.<name>]`, so the
Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
Desktop/TUI agent knows the provider id (it otherwise collapses to
"custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
bare-`custom` ineligibility.
Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).
Fixes#75186
When a newer attempt claimed the stream-writer slot, run_codex_stream's
accept_chunk hook returned False and Relay stopped consuming the Codex
Responses stream. The assembler then returned a "completed" response
holding only the text seen so far, and the gateway delivered a reply cut
mid-sentence (#69486) even with display streaming disabled.
Supersession now fences only the live callbacks (text and reasoning
deltas, commentary, first-delta) through the attempt's writer token and
consumption continues to the terminal frame, so the assembled final
response is complete. The single-writer invariant still holds: a
superseded attempt never writes into the live display. Request
retirement (watchdog kill) is unchanged and still stops the worker.
Ported from PR #69502 onto the Relay-backed stream path.
Fixes#69486
The OpenAI SDK always raises ``APIConnectionError(...) from err`` with the
httpx transport error as the direct ``__cause__``, so the chain-walking
helper from the salvaged commit is more machinery than the signal needs.
Replace it with an ``isinstance(exc.__cause__, httpx.TransportError)``
check at the single call site and document WHY the raw
``transport_errors`` branch never sees a pre-stream failure (the SDK wraps
them), which is the root cause behind #103673.
Responses lifecycle events may arrive before answer text, but the streaming runtime never populated the existing post_api_request timing field. Record the first accepted event while preserving request-retirement fencing.\n\nCloses #105311.
Replace the hardcoded 2.0s module constant from the salvaged commit with
``agent.stream_drain_timeout`` in DEFAULT_CONFIG (default 2.0, ``0`` skips the
drain), read through ``load_config_readonly`` at drain time. Non-secret knobs
belong in config.yaml, not a new HERMES_* env var. Documents the budget in the
timeouts table and adapts the salvaged test to the new seam.
Why: a relay that never closes the SSE socket after ``response.completed``
wedged the turn until the stale watchdog fired and discarded the already-billed
response (#103864). The drain is a courtesy to Relay's finalizer, so it must be
bounded; the bound needs a user-tunable knob for slow-but-honest relays.
The codex_app_server early-return forked away from the standard loop before any
prompt assembly was used: the codex thread received only cwd and the raw user
text, so SOUL.md, MEMORY.md/USER.md and channel_overrides.system_prompt were
composed, persisted to sessions.system_prompt, and then silently discarded.
_ensure_codex_session now composes the prompt exactly as turn_context does
(agent._cached_system_prompt + "\n\n" + agent.ephemeral_system_prompt) and
passes it to CodexAppServerSession(developer_instructions=...) once per thread.
Sessions are created lazily per turn only when none exists, so the prompt is
never duplicated per turn and a retired session re-sends the current one.
Hand-ported from #74726: same composition, but the session kwarg is spelled
developer_instructions -> thread/start.developerInstructions because the
`instructions` field #74726 used is accepted and ignored by codex.
Fixes#74712Fixes#26035
Replace the six parametrized subprocess-driven cases from #111616 with two
invariants: a successful marker-ending output round-trips SessionDB and
replay sanitize as data (envelope preserved), and a real interrupt or an
unknown exit status still recovers as UNKNOWN. The projector consumes the
event dict, so spawning real children added no signal.
Also re-word the live-card payload docstring in agent/codex_runtime.py: it
no longer mirrors the persisted tool-result content byte for byte.
The post-tool quiet timer no longer retires the session (it only warns), so the
codex_runtime retirement comment listed a cause that cannot happen and the
retained test's name/docstring still promised a watchdog that cannot trip.
Comment and test wording now describe the warning semantics; assertions unchanged.
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
`chat.completions.create` re-walks the whole request body against the
`CompletionCreateParams` union graph client-side, with the GIL held, before
any byte leaves the process. #93650 documented that class of walk wedging
for 12+ hours on a ~1.4 MB conversation: no in-process watchdog can fire
while the GIL is held, and no socket kill helps a pre-network hang.
through `extra_body`, which the SDK merges into the JSON body after the
transform — but scoped it to `responses.create`. The default chat path,
which every OpenRouter / Nous / xAI / DeepSeek / Kimi / llama.cpp /
Ollama / LM Studio / LiteLLM request takes, still pays the full walk.
Measured against a real `openai.OpenAI` over an `httpx.MockTransport`
(canned SSE, no network), with the request body captured from the
transport on both sides:
101 msgs / 76 KB 13.6 ms -> 1.2 ms
401 msgs / 190 KB 48.6 ms -> 2.0 ms
1601 msgs / 650 KB 188.7 ms -> 5.8 ms
and the bytes the server receives are IDENTICAL — literally equal, not
merely equivalent (194,894 == 194,894 at 401 messages). The cost is paid
per API call, so a tool-using turn multiplies it by its iteration count.
The three helpers move from agent/codex_runtime.py into a shared
agent/sdk_transform_bypass.py, re-exported under their original names so
agent/auxiliary_client.py and tests/run_agent/test_codex_sdk_transform_bypass.py
keep working untouched. The field tuple is now a parameter:
("input", "tools") for Responses, ("messages", "tools") for chat.
Two chat-specific details. `messages` is a @required_args parameter, so it
stays in the typed kwargs as an empty list and the extra_body copy
replaces it in the body — hence the new `required_empty` argument, which
Responses does not use. And the bypass is gated on the target actually
being the SDK's Completions: Hermes also drives chat-completions-shaped
facades that are NOT the SDK — the in-process MoA aggregator most
importantly — and those never merge extra_body, so handing them one would
silently send an empty message list. That guard is also why this needs no
edits to the 32 test files that assert on kwargs["messages"]: they mock
with stand-ins, not the SDK.
Every rail the merged PR was reviewed on is kept: the plain-JSON-only
guard so pydantic models and generators stay on the typed path, caller
`extra_body` precedence via setdefault (load-bearing here — the chat path
already populates extra_body from custom providers, reasoning config and
Nous Portal), and an env escape hatch, HERMES_CHAT_SDK_TRANSFORM=1,
mirroring HERMES_CODEX_SDK_TRANSFORM.
The summary/compression call sites at chat_completion_helpers.py:3449 and
:3514 carry the largest payloads in the process and are deliberately left
for a follow-up: they route through a lambda whose client is not in scope
at the call site, so they need a slightly different shape and a wider
test surface than this change.
The name of Hermes' MCP callback for the codex app-server runtime was spelled
as a string literal in five places (the server itself, the runtime migration
that writes `[mcp_servers.hermes-tools]`, the Kanban worker override launcher,
the elicitation auto-accept handler, the display-name stripper and the switch
report) and had already drifted once (#111707). Define it once in
agent/transports/hermes_tools_mcp_server.py — the module that IS the server and
whose module-level imports are stdlib only, so every higher layer (transports,
agent/codex_runtime, hermes_cli) can import it without a cycle — and read it
everywhere.
Two invariant tests in tests/agent/transports/: the worker's `-c
mcp_servers.<name>.env.*` overrides only ever target an entry the migration
really writes to config.toml (red on the pre-fix base: `{'hermes-mcp'}`), and
non-owned launches emit no override at all.
Refs #111707
When every row create of a turn loses to the SQLite lock, the queued token delta's
"ensure the row exists" guard becomes the session's first writer and minted the row as
source='unknown'. That placeholder was permanent on the real path even with the upsert
repair from #112045: the turn lease (turn_facade_lease.admit_durable_turn) treats an existing
row as proof the create already happened and sets _session_db_created, so the creator never
returns to repair it. Live probe: a platform="desktop" AIAgent whose create_session raised
"database is locked" for the whole first turn ended with a source='unknown' row on base AND
on the contributor head; with this change the row is minted 'desktop' by the guard itself.
Producer fix: update_token_counts gains an optional source= that the two agent call sites
(agent/turn_usage.py, agent/codex_runtime.py) fill from _session_source_for_agent(platform),
the same value _ensure_db_session would stamp. record_auxiliary_usage has no surface and
keeps the placeholder, which the creator's upsert now repairs.
Salvage trims: the contributor's SimpleNamespace dispatch test is replaced by a real-AIAgent
invariant test under tests/agent/ (the dispatch hunk in _run_prompt_submit is kept; the
INSERT-OR-IGNORE is idempotent under prompt.submit's own persist); narration comments cut
to the WHY; docs list 'unknown' among the startup-sweep sources.
Refs #111999
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
stream watchdog sees progress on a refusal-only stream instead of timing it
out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
parts; without it an aux refusal-only turn parsed to content=None and hit the
empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
alongside-content) so there is one test per surface; drop upstream product
references from docstrings (credit stays in the PR body); pass encoding= to
the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
content_filter result, not an empty response to retry.
A model that declines mid-stream delivers the explanation on the
structured refusal channel (chat_completions delta.refusal; Responses
response.refusal.delta / refusal content parts) and leaves content
empty. The streaming accumulators dropped that channel entirely, so a
streamed refusal assembled into an empty message and fell into the
empty/invalid-response retry loops - burning paid retries reproducing a
deterministic refusal - while the non-streaming path had already fixed
this class in #46013.
- chat_completions streaming: accumulate delta.refusal (incl.
model_extra), expose message.refusal on the assembled mock response so
ChatCompletionsTransport.normalize_response applies the existing
sole-payload -> content_filter promotion; count refusal deltas in the
zero-chunk guard; carry refusal in the Relay final-response dict.
- Codex Responses stream consumer: collect response.refusal.delta as
answer text so a refusal-only stream no longer raises 'did not emit a
terminal response' with zero usable content.
- Responses normalizer: read type=refusal content parts in
_extract_responses_message_text (attr and dict shapes).
Sabotage-verified: each new test fails with its wiring line disabled.
E2E: refusal-only stream -> terminal content_filter with explanation;
refusal-alongside-content stays a normal usable turn; plain-text
streams unchanged.
Keep the real httpx 0.28 wrapper-shape test (proves the shutdown reaches
the socket through BoundSyncStream/ResponseStream/PoolByteStream) and the
loopback E2E (a parked reader unwinds within its stale budget and the
retry lands). The other four were narrower restatements of the same paths.
Also treat httpx.ReadError as a transport error in codex_runtime: it is the
same abort-induced-read class the streaming retry loop now recovers from.
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.
Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.
Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698
Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).
Now there is one authority:
- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
over threshold waits ONE request for the provider's real count instead of compressing on a guess
(first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
(note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
past the whole window, and provider-proven overflow all compress immediately; the post-compaction
latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
that scripted whole-history estimates now state the fact they relied on (provider omits usage).
Fixes#103391 (closes#103397 by construction — the baseline it repaired no longer exists).
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).
- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
resumed turn while the durable transcript still matches, cleared with the row on
compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).
Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
The codex_app_server runtime bypasses the conversation loop, so the
usage-anchored context accounting captured there never ran: agent._usage_anchor
stayed None forever. Every preflight estimate therefore fell back to the rough
mirror-transcript heuristic, which is deliberately never compacted on this
runtime (_record_codex_app_server_compaction preserves the mirror), so it grows
monotonically while the real thread may be tiny or freshly compacted. With
compression.codex_app_server_auto=hermes that estimate alone tripped the
threshold and fired thread/compact/start on nearly every turn of a long-lived
session (#100381).
Mirror the main loop's post-response capture: _record_codex_app_server_usage
now snapshots the reported thread usage into agent._usage_anchor (base prompt
+ completion exactly as the provider counted them, with estimation confined to
messages appended since). A usage-less turn keeps the previous anchor, and the
post-compaction invalidation site is unchanged.
With a request-local watchdog state on every codex request, agent._codex_stream_last_event_ts / _last_progress_ts had no remaining reader; the state-None snapshot only runs for non-codex requests where no watchdog consults it.
- Explicit-vs-implicit idle timeout detection uses env_float with a sentinel
instead of a hand-rolled parse (unset and unparseable both mean implicit).
- _on_event writes the agent-level timestamps only when there is no
request-local watchdog state; with state present they were dual bookkeeping
nobody read (the snapshot prefers the state).
- _abort_request keeps main's close-then-retire order: the reorder had no
test teeth and is a separate concern from #90449.