Split _preflight_codex_input_items into per-item-type helpers over a
_PreflightCtx; streaming assembly moves into _CodexResponseAssembler with a
per-event dispatch table; run_codex_app_server_turn loses its session/usage/
interrupt regions to _ensure_codex_session/_finish_codex_turn/
_consume_user_interrupt/_queue_token_counts (counts built lazily so stub agents
without a session DB are never touched). Responses input/normalization and
stream callback order verified byte-identical against merge-base.
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.
Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.
Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.
Fix: publish a per-request retirement token so the worker can tell it has been
retired.
- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
`agent._active_codex_stream_request_token` before handing off to the worker
(codex_responses only) and clears it at all four kill sites plus the worker's
own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
which can raise — every other call site wraps it in try/except, and a leaked
token would let a later worker mistake itself for the owning attempt. The
request-local `_codex_request_retired` mirror also swallows the transport
error our own force-close causes, so the worker's local error cannot replace
the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
`TimeoutError` from `interrupt_check` when it no longer owns the request —
raising rather than breaking, because a break returns the partial `final`.
The four stream callbacks also drop post-retirement frames so an abandoned
attempt cannot stream tokens into the live turn's bubble (the gateway caches
AIAgent instances per session).
`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.
Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.
- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
delta estimate (stale reasoning excluded on all but the newest assistant
message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.
Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):
- codex app-server: in "native"/"off" auto-compaction mode (native is the
default) Hermes preflight is skipped and the codex agent compacts its
own thread inside run_turn() — the compress_context() rejection was
unreachable. init_agent now refuses checkpoint_required together with
api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
testable guard), and run_codex_app_server_turn() fails closed as
defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
native_compaction_context_management() now returns None while the gate
is armed, so context_management never goes on the wire and the
checkpoint-aware Hermes compressor stays authoritative. The suppression
is logged once per process, not silently applied.
Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.
Refs #93986
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.
Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.
Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
Review polish from the 3-angle pass on the final stack:
- The warning now says WHICH shape leaked (top-level, extra_body, or both).
Relay injects top-level while request_overrides typically inject via
extra_body, so the shape identifies the offending middleware when
debugging.
- Fold the 'always returns a fresh mapping' assertion into the parametrized
real-endpoint test (the caller mutates the result with stream=True, so the
copy contract is load-bearing on no-drop paths too) and drop the
SimpleNamespace stub test it strictly subsumes. The nested-preserve stub
stays: the parametrized test only exercises top-level retention.
The wire guard only removed the top-level prompt_cache_retention kwarg, but
the OpenAI SDK merges extra_body into the outgoing JSON body, so a nested
extra_body.prompt_cache_retention reaches chatgpt.com/backend-api/codex just
the same and still triggers the non-retryable HTTP 400. Both injection
vectors are real and probe-verified: the Relay overlay's 'key not in
baseline' arm admits an interceptor-added extra_body, and
request_overrides={'extra_body': {...}} lands verbatim in build_kwargs
output.
Close the gap in the same helper: strip the nested field too (copy-on-write,
never mutating the caller's mapping), drop extra_body entirely when it
empties, and log the same warning. Compatible endpoints keep nested
retention untouched.
Mutation-verified: removing the extra_body leg fails both new nested tests.
Reported by egilewski's review on #89969.
The salvaged compatibility test stubs `_is_codex_backend=lambda: False` on a
SimpleNamespace, so it proves the helper honors its own boolean but not that
the boolean is right for any real endpoint. A predicate change that widened
the drop onto retention-supporting hosts would keep it green.
Adds a parametrized test that builds a real AIAgent per base URL and asserts
the drop only fires for chatgpt.com/backend-api/codex, while api.meta.ai,
bedrock-mantle.*.api.aws, api.openai.com and a same-host/different-path
backend keep their supported 24h value. Also asserts prompt_cache_key
survives untouched on every endpoint, since retention and cache-key routing
are independent and the guard must not disturb caching.
Verified non-vacuous: relaxing the guard's condition to drop on every
endpoint fails 4 of the 6 cases (Meta, Bedrock, OpenAI, non-codex path).
Drive-by on the guard itself: drop the dead `None` default on the `pop` that
is already gated by an `in` check, and record why the predicate is resolved
via getattr -- run_codex_stream is driven with lightweight stand-in agents
that lack `_is_codex_backend`, so a bare call would raise AttributeError.
The PR added APIConnectionError handling to the main request and
iteration try blocks but missed the finalization drain loop (line ~1492).
That site catches httpx transport errors to preserve an already-completed,
already-billed response when the drain iterator fails. Without the
APIConnectionError handler, an SDK-wrapped transport error during drain
would propagate uncaught and discard the completed response.
Also strengthens the test's no-payload-leak assertion to check the full
request body and URL are absent from the log message, not just the
literal string 'payload'.
The native Responses stream does carry summary_index, so the part boundary is
structured data here rather than something to infer. Break on a change of
index, and leave streams that send no index (plain reasoning_text) untouched.
A transport error during the post-terminal SSE drain (used only to let the
Relay transport finalize the attempt) was sharing exception handling with
the pre-terminal assembly path, so it discarded an already-completed,
already-billed response and opened a brand-new physical request. Give the
drain step its own non-fatal error handling: log a finalization warning
and still return the terminal response.
Closes#74310
Every LLM call built a fresh openai.OpenAI wire client (new httpx pool,
TCP+TLS handshake, measured 19.2ms p50 / 35.5ms p95 per call at ~5 calls
per tool-loop turn) and closed it when the request finished. Cache ONE
reusable wire client on the agent, keyed by the effective client kwargs:
- _create_request_openai_client hands back the cached client when the
effective kwargs are identical; any change (credential rotation,
provider failover, vision default_headers) evicts and rebuilds.
- Only a request that produced a response reports a reuse close reason
(request_complete / stream_request_complete); error unwinds report
*_error_cleanup and really close, so a retry after a request error
always builds a fresh pool.
- Cross-thread aborts poison the slot: a pool whose sockets were
shutdown(SHUT_RDWR) from a stranger thread is never reused (#29507) —
the owner-thread close discards it and the next create rebuilds. The
holder read and the abort are atomic (under the holder lock) at all
three abort sites, so a late abort can never poison the NEXT request's
checked-out client.
- Worker-side interrupt breaks close the half-read SSE stream on the
owning thread before building the partial response; a failed close
poisons the slot (otherwise each interrupt leaked one checked-out
connection until the pool hit PoolTimeout). run_codex_stream gets the
same poison-on-close-failure handling.
- Single checked-out slot (in_use): a concurrent call gets an untracked
client with the old per-request lifecycle.
- release_clients() / close() really close the cached client when idle;
if a worker has it checked out they abort the sockets and detach the
slot, deferring the FD release to the worker's own close.
- MoA facade and Mock passthroughs never enter the cache; max_retries=0
is preserved on all request clients.
Every API call in the tool loop persisted its token/cost delta by
calling SessionDB.update_token_counts() synchronously on the turn
thread — a BEGIN IMMEDIATE sessions UPDATE plus a session_model_usage
upsert, measured in production at p50 3.3ms / p95 70.4ms per call and
up to 299ms against a cold multi-GB state.db. The tool loop stalls for
that long between calls, N times per multi-tool turn.
SessionDB gains queue_token_counts(): same signature and semantics as
update_token_counts(), but the critical path is a deque append plus a
condvar notify. A lazily started daemon thread applies deltas in
enqueue order through the existing update_token_counts ->
_execute_write path, so the established self._lock / BEGIN IMMEDIATE /
jitter-retry discipline is unchanged. When a backlog forms, adjacent
same-route incremental deltas coalesce into one UPDATE: token and
api-call fields sum, cost fields sum None-preservingly (an all-None
run stays None so COALESCE keeps the stored value), and absolute=True
deltas never merge and act as ordering barriers. Route equality is
required for a merge because those fields feed COALESCE backfill, the
last-non-None-wins status fields, and the per-model usage attribution
key — a merged apply is row-equivalent to sequential applies.
Correctness and durability:
- flush_token_counts() gives read-your-writes to token/cost readers
(get_session, list_sessions_rich, _get_session_rich_row,
list_gateway_sessions, InsightsEngine.generate) — a plain attribute
check when nothing is queued. The writer sets its busy flag before
popping the queue so the lock-free fast path can never miss an
in-flight batch.
- update_session_model / update_session_billing_route /
update_session_meta write the sessions row synchronously, bypassing
the queue, so they flush it first: a still-queued first-of-session
delta carries the pre-switch route, and applying it after the switch
UPDATE would trip the first_accounted_route branch (api_call_count
== 0 plus a route mismatch) and resurrect the old model/provider.
- AIAgent._persist_session flushes at turn finalize and every
error-exit persist point; close() stops and drains the writer before
the WAL checkpoint; an atexit hook (registered on first enqueue,
unregistered on close so closed instances are not pinned until
interpreter exit) drains at shutdown. Worst-case crash loss is the
in-flight call's delta — the same window as the old inline write.
- A flush trusts a live stop-flagged writer (its loop drains before
exiting) and only drains on the caller's thread when the writer is
dead or never started, claiming the same busy flag so concurrent
flushes wait instead of racing an in-flight batch.
- After close() has stopped the writer, queue_token_counts applies the
delta inline instead of parking it on a queue nothing will drain; a
closed-connection failure then raises at the call site, which
already guards for it, exactly like the old synchronous path.
- Writer apply failures are logged and never raise into a turn; the
writer thread survives and keeps applying.
Call sites switched to the queue: the per-call site in
agent/conversation_loop.py and both codex app-server sites in
agent/codex_runtime.py. In-memory per-turn counters
(agent.session_estimated_cost_usd etc.) stay synchronous, so live turn
displays never see the queue.
Tests: tests/agent/test_async_token_accounting.py (19 tests: enqueue
ordering, absolute-as-barrier, backlog coalescing with exact sums,
coalesced-vs-sequential row equivalence, merge unit rules, None-cost
preservation, read-your-writes, flush vs stop-flagged/concurrent
drains, inline apply after writer stop, close/atexit durability,
_persist_session drain, writer failure isolation);
tests/run_agent/test_token_persistence_non_cli.py updated to the
queue_token_counts contract.
- codex app-server sibling path: surface a WARNING (was silent debug) when
the projected-message flush fails — same bug class as the main fix, but
codex output has already streamed so fail-closed and agent_persisted=False
are both wrong here (#860/#42039 duplicate-write hazard); loud durability
gap logging instead.
- map session_persistence_failed in _format_turn_completion_explanation so
the user sees an actionable reason instead of 'The request failed:
unknown error' + explainer test.
- contributors/emails mapping for elco@thedaoist.gg (attribution CI).
The Codex app-server runtime bypasses the main conversation loop and drives
its own subprocess turn, so it needs first-class hooks rather than the
OpenAI-loop interrupt path.
- `AIAgent.interrupt()` now forwards a hard stop to
`CodexAppServerSession.request_interrupt()`, and `redirect()` uses Codex's
native `turn/steer` protocol instead of cancelling the subprocess.
- `run_turn()` no longer clears an interrupt that arrived during
`ensure_started()`: a stop landing mid-startup is honored before `turn/start`,
and the interrupt event is cleared on every exit path.
- `run_codex_app_server_turn()` mirrors the loop finalizer's interrupt handoff
(surface `interrupted` / `interrupt_message`, then `clear_interrupt()`) on
both the normal and exception early-return paths, so a hard stop can't leave
`_interrupt_requested` stale for the next turn.
Extends the app-server event bridge (make_codex_app_server_event_bridge)
to fire the authoritative stable-ID tool_start_callback /
tool_complete_callback alongside the existing tool_progress_callback,
and route item/reasoning/summaryDelta through the reasoning channel.
Surfaces that render structured tool cards (TUI, desktop) — not just
progress bubbles — now correlate live cards with the projected history
entry after a resume, because the call ids mirror CodexEventProjector's
_deterministic_call_id. Guarded per-callback so a broken display
consumer can't tear down the codex turn loop.
Grafted from PR #65412 by @HaiderSultanArc onto the merged bridge (the
PR's parallel _codex_live_event implementation was reconciled into the
bridge's existing _fire_tool_started/_fire_tool_completed helpers).
A cron job ("Daily Buzz Report") died with 'AIAgent' object has no
attribute '_claim_stream_writer'. The #65991 single-writer fence lives on
AIAgent (run_agent.py), but the streaming paths that use it live in other
modules — chat_completion_helpers (chat / anthropic / bedrock) and
codex_runtime (codex responses) — and called it directly as
agent._claim_stream_writer() / agent._stream_writer_is_current(). That makes
those modules hard-depend on the method being present on whatever object is
passed as agent.
The fence is an *additive* safety net that may only ever drop a provably
superseded stream, never the sole legitimate writer. But the direct calls
turned any agent that doesn't expose it — a version-skewed checkout (the
streaming helper module newer than run_agent), a hot-reloaded gateway mid
git-pull, a duck-typed agent, or a test double — into a fatal AttributeError
that aborts the whole turn (and, on cron, fails the job).
Route every cross-module claim/check through agent/stream_single_writer.py.
claim_stream_writer(agent) returns 0 when the fence is unavailable (or
raises), and stream_writer_is_current(agent, token) treats a 0 token or an
absent guard as "current" — so a guard-less agent degrades to "no fence"
instead of crashing, while a real AIAgent keeps the full single-writer
protection. Internal self.* uses inside run_agent are unchanged (self is
always a full AIAgent there).
Two more display gaps from #26541 grafted onto the merged bridge:
- webSearch: codex's built-in web search now produces a tool.started/
tool.completed bubble pair (query as preview + args). Previously the
item type wasn't in _CODEX_TOOL_ITEM_TYPES, so built-in searches
showed nothing.
- mcp.hermes-tools.* stripping: tools codex invokes through Hermes' own
hermes-tools MCP server display as their bare names (web_search,
browser_navigate) instead of mcp.hermes-tools.web_search. The inner
dispatch subprocess can't fire native progress events, so the
codex-level event is the display event — name it the way users know
the tool.
Credit: both behaviors designed and first implemented by @simpolism in
PR #26541 (May 15, earliest of the app-server display-bridge family).
Widen the #65991 single-writer fence to run_codex_stream: each codex
attempt claims the delta sink before consuming events, and the consume
loop's interrupt_check now also stops the instant a newer attempt
supersedes this one. Parity with the chat_completions / anthropic /
bedrock paths from the salvaged fix.
Two regression tests: superseded codex stream is fenced mid-stream;
sole-writer codex stream delivers unchanged.
Port from anomalyco/opencode#36130: the Responses spec carries streaming
error details at the top level of the error frame, but the official OpenAI
SDK and several OpenAI-compatible proxies wrap them in an HTTP-style nested
envelope ({"type": "error", "error": {code, message, param}}).
_raise_stream_error only read top-level fields, so nested-envelope frames
collapsed to the generic 'stream emitted error event' placeholder with
code=None — the error classifier never saw the provider's real failure
reason, misrouting rate-limit / context-overflow / entitlement errors into
the generic retry path.
Top-level fields keep precedence; the envelope is a fallback. Null-tolerant
for spec-compliant frames with explicit nulls.
Follow-ups on top of @xxxigm's salvaged bridge (#33294):
- Remove the now-dead narrow item/started-only mapper from #38835
(_codex_note_to_tool_progress) — the full bridge supersedes it and
keeps the same tool-name contract; its tests are repointed at the
bridge helpers.
- Preserve main's request_routing/approval-bypass wiring on the
CodexAppServerSession constructor (landed after the PR was filed).
- Gate agentMessage interim delivery on display.show_commentary so the
app-server runtime honors the same toggle as the codex_responses
commentary channel (tool progress is unaffected).
- Add json import (bridge helpers use json.dumps) and modernize the
wiring test's stub agent for main's usage-accounting attributes.
Pass ``on_event=make_codex_app_server_event_bridge(agent)`` when
spawning the per-session ``CodexAppServerSession``. The session has
always had a raw event hook but ``run_codex_app_server_turn`` never
supplied one, so Discord / Telegram / TUI users saw nothing while
codex was working — only the final answer landed.
Now each ``item/started`` for a tool-shaped item fires
``tool_progress_callback("tool.started", ...)``, ``item/completed``
fires the matching ``"tool.completed"`` with duration + result,
``item/agentMessage/delta`` flows through ``_fire_stream_delta`` and
each completed ``agentMessage`` surfaces through
``_emit_interim_assistant_message`` so the gateway's
``already_streamed`` dedupe keeps interim commentary in the channel
without duplicating text the stream already showed.
Adds ``make_codex_app_server_event_bridge(agent)`` plus four small
mapping helpers (``_codex_item_to_tool_name`` / ``_codex_item_to_args``
/ ``_codex_item_to_preview`` / ``_codex_item_completion_payload``)
that translate codex JSON-RPC ``item/*`` notifications into the
exact shape Hermes' gateway UI callbacks expect — tool names match
``CodexEventProjector`` so the progress bubbles and the projected
``tool_calls`` entries use the same identifiers.
No behaviour change yet: the next commit wires the bridge into
``run_codex_app_server_turn`` (#33200).
Commentary delivery is on by default; users who find the extra mid-turn
narration noisy can set display.show_commentary: false to restore the
previous behavior (commentary routed to the reasoning channel, visible
only with show_reasoning).
- hermes_cli/config.py: display.show_commentary default true
- agent/agent_init.py: wire config -> agent.show_commentary
- run_agent.py: gate structured commentary extraction on the flag
- agent/codex_runtime.py: gate live-stream commentary callback (falls
back to legacy reasoning-channel routing when off)
- docs + 2 tests (interim path off, live stream fallback)
Also adds AUTHOR_MAP entries for davidrobertson and 100yenadmin.
Preserve deletability, route identity, stored costs, aggregate reconciliation,
and zero-usage Codex route accounting on top of the salvaged per-model usage
work.