No ASCII code point is a Unicode format control (Cf), so ASCII text returns on the
O(1) isascii() flag and other text categorises set(text) minus ASCII instead of every
character. Output is unchanged (same unicodedata predicate).
Python 3.12, per call: CJK 9.2k chars 3.60 -> 1.57 ms, mixed 3.49 -> 0.30 ms,
ASCII 3.16 -> ~0 ms (main -> this commit). A per-match finditer over non-ASCII
characters was slower than main on CJK (11.0 ms) and is not used.
_neutralize_harmony_tokens walked every character of each string leaf that
contains both '<' and '|' through unicodedata.category() to look for format
controls. No ASCII code point is Cf and str.isascii() is an O(1) flag check,
so gate the scan on it. 160 KB ASCII code text: 85.7 ms -> 0.18 ms per call;
output unchanged.
Partial salvage of #98340: kept the isascii fast path; dropped the compiled
per-Unicode-version Cf pattern table, the benchmark script and the
table-drift/differential-matrix tests (source-reading change-detectors).
(cherry picked from commit 84e6561ce710b861ec1861e92558d2ae433dccab)
A tool call whose function.name violates the provider pattern
^[A-Za-z0-9_-]{1,64}$ — OpenAI's synthetic `multi_tool_use.parallel`, or a
whole shell command a weak fallback model put into `name` (372 chars) — is
persisted once and then 400s every later request on a strict endpoint, so
the session silently pins itself to the lenient fallback model (#51944).
`agent/message_sanitization.py::coerce_tool_name` is now the single owner of
the coercion (valid → identity, invalid runs → `_`, cut at 64, empty →
fallback); the Codex Responses adapter uses it in place of its private copy,
and the pre-call sanitizer's nameless-call repair becomes
`_repair_invalid_tool_call_names`, so both outbound builders that already
call `sanitize_api_messages` — the main loop (turn_request_assembly) and the
iteration-limit summary (chat_completion_helpers) — send valid names. Dict
tool calls are rewritten copy-on-write, so the summary path's shallow message
copy never edits persisted history; tool results follow through
`_realign_tool_result_names`. Deterministic, so identical stored bytes always
render identical wire bytes (prompt-cache prefix stays stable).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Review follow-ups on the inline-image guard:
- `image/jpg` is the JPEG alias every other image site accepts
(vision_message_prep, conversation_compression, image_gen_provider,
mcp_tool_content); the Responses guard downgraded it to a text
placeholder, losing a valid image. It now counts as JPEG.
- The Anthropic converter forwarded data:image/svg+xml (bmp, tiff) verbatim
as media_type, which 400s every turn once the part is in history. It now
applies the same rule: SVG is rasterized to PNG when a rasterizer exists,
any other unsupported inline subtype becomes a text placeholder, and
image/jpg is normalized to image/jpeg.
Both wire paths share one helper next to the existing supported-set
constant in tools/vision_tools_image_prep (import-time deps: hermes_constants
only, no cycle).
A/B: new Anthropic test and the jpg assertion red on the previous head,
green now.
The send-layer guard from c9f8cb72a6a downgrades every inline
data:image/svg+xml part to a text placeholder, so the model never sees a
drawing the agent (or a user attachment) wanted it to look at. The
reporter's ask was to rasterize instead of drop when that is possible.
Reuse the vision_analyze rasterizer chain (cairosvg / svglib+reportlab /
rsvg-convert / inkscape, all soft deps) through a small
rasterize_svg_data_url helper in tools/vision_tools_image_prep: the SVG
payload is decoded to a temp file, rasterized, re-encoded as a PNG data URL
and forwarded as input_image; the temp files are removed. Without a
rasterizer the existing placeholder still applies, so the request can never
carry SVG source. Import is lazy inside _input_image_part (no cycle; the prep
module only depends on hermes_constants at import time).
Live (real converter on a temp HERMES_HOME): base with cairosvg installed ->
input_text placeholder; fixed with cairosvg -> input_image
data:image/png (PNG magic verified); fixed without any rasterizer ->
placeholder.
A `data:image/svg+xml` (or BMP/TIFF/...) part reaching the Codex Responses
converter was forwarded verbatim as `input_image`; the backend rejects the
WHOLE request with 400 "The image data you provided does not represent a
valid image", and because the part is baked into history (a persisted
vision_analyze tool result, a user turn) every later continuation re-trips
the same 400 until the user drops the turn.
The inputs that create SVG parts are already guarded (image_routing skips
SVG for native attach; vision_analyze rasterizes it), so this is the
converging send-layer seam: `_input_image_part` now downgrades any inline
image whose subtype is outside jpeg/png/gif/webp to an `input_text`
placeholder. It is the single seam for user/tool message content,
`function_call_output.output`, and the preflight validator, so all three
carriers are covered; remote http(s) URLs pass through untouched (the
provider owns their validation). Valid images in the same message still go
as `input_image`.
The 400-recovery matcher already recognises the Codex wording since
0241619068, so only the send side was missing.
Docs: vision.md now documents `agent.image_input_mode` (auto|native|text)
and the explicit `auxiliary.vision` override — the way to keep an
openai-codex main model while routing image analysis to another vision
provider when the backend answers image requests with server_error.
Salvages #47299 (@hanzckernel) — same fix direction, redone against the
current converter (the original diffs predate the `_input_image_part` seam).
'Next, I'll create the script.' / 'First, let me check the directory.' /
'Okay — running the tests.' followed by a closing {"cmd": ...} object were not
classified as leaked tool calls because the lead-in pattern anchored the action
verb at line start. Allow an optional Next/First/Then/Okay/OK/Alright marker
(with comma/dash) before the existing prefixes; the 'closes the message' and
'cmd' key constraints are unchanged, so bare/explained JSON answers still pass
through (#56920).
gpt-5.x on the Codex Responses backend sometimes serializes the tool call it
meant to make as Codex-CLI shell JSON closing the assistant text
('Creating the script now.\n{"cmd": "mkdir -p ..."}') instead of a structured
function_call item. _normalize_codex_response only recognised the Harmony
`to=functions.<name>` leak, so this shape was normalised as plain content with
finish_reason=stop: the JSON was printed, nothing ran, and the turn ended.
Classify it through the same leak seam: a trailing `{"cmd": ...}` object (scalar
siblings only) whose previous line is an action lead-in is a failed tool call, so
the turn is incomplete and the existing Codex continuation re-elicits a real
function_call. A bare or explained `{"cmd": ...}` payload, or one that does not
close the message, stays a legitimate answer (false-positive guard).
Leaked text (both leak shapes) also no longer populates codex_message_items, so
the continuation cannot replay it as a completed assistant message.
Salvaged from #56958 (@YanzhongSu, commit by @minhngoc25a): compact lead-in
regex, one detector for both leak shapes, tests trimmed to two invariants.
Fixes#56920
_newest_reasoning_only pops older turns' codex_reasoning_items before the
converter runs, so those turns looked reasoning-free and replayed their
msg_* id with neither the reasoning item nor its rs_* id on the wire — the
exact orphan shape #97427 rejects. The trimmed row now carries a transient
codex_reasoning_trimmed marker and _replay_message_items honours it; the
single-use _turn_has_encrypted_reasoning wrapper is inlined into that
predicate. Test drives ResponsesApiTransport.build_kwargs on an Azure host.
Stateless Responses replay (store=False) strips every reasoning item's rs_*
id, but the assistant message minted in the same response kept its msg_* id.
GPT-5.6-family endpoints validate that link and reject the continuation with
HTTP 400 "Item 'msg_…' of type 'message' was provided without its required
'reasoning' item: 'rs_…'" on every post-tool turn, deterministically.
The converter now drops the message id whenever the stored turn carried
encrypted reasoning (replayed, suppressed by recovery, or dropped as a
foreign-issuer blob), keeping content/status/phase. Reasoning-free turns keep
their id for prefix-cache affinity. Done in _replay_message_items so the main
transport, preflight and the auxiliary Codex adapter all emit the same shape.
Co-authored-by: salehelsayed <saleh.fekry@gmail.com>
The second failure mode in #51512: with no reasoning replay at all, a single
``{"role": "user", "content": "<text>"}`` item still 400s on the ChatGPT Codex
backend with ``{"detail": "Unsupported content type"}``. The classifier mapping
from the previous commit cannot recover that turn (turn_recovery's strip needs
cached codex_reasoning_items), so the wire shape has to be right up front.
``_chat_messages_to_responses_input`` now wraps string user/assistant text as
``input_text`` / ``output_text`` parts when the issuer is ``codex_backend``; every
other Responses route keeps the string shorthand it has always received. The
preflight already validates typed parts, so the real call path (build_kwargs ->
preflight_kwargs) needs no separate change.
Test drives ResponsesApiTransport.build_kwargs + preflight_kwargs, red on the
old head for the codex case. Three test_native_compaction asserts pinned the
assistant string incidentally (they check history survives, not its shape) and
now expect the typed part on the codex route.
_PREFLIGHT_ALLOWED_KEYS is derived from _PREFLIGHT_OPTIONAL_FIELDS and `text`
was not on that table, so any Responses request carrying `text.verbosity` (or
the structured-output `text.format` block) died inside Hermes with
"Codex Responses request has unsupported field(s): text." before reaching the
provider. One table entry both allows the key and passes it through
normalization like service_tier / context_management; empty or non-dict
values are dropped rather than rejected, matching the other optional fields.
Transport-level prerequisite for #20203. Salvaged from PR #103329 (test hunk
trimmed; invariant test lands in a separate commit).
Stored tool call ids are minted per turn (terminal:0, terminal:1…), so the
same id recurs on later turns of one session. The Responses adapter replayed
each pair verbatim, putting the same call_id on the wire N times; strict
validators (opencode Console, OpenAI) reject the request with 400
"Duplicate function_call_output for call_id" and every retry fails the same
way (#102629, #111231 bug 1).
Every occurrence past the first now gets a `_dup<n>` wire id and the matching
tool output pops the id its function_call was given, in call order, so pairs
stay intact and stored history is untouched.
Re-applied onto the facade split (helpers moved into _replay_tool_call_items /
_tool_output_items) from PR #102634.
A model that declines mid-stream delivers the explanation on the
structured refusal channel (chat_completions delta.refusal; Responses
response.refusal.delta / refusal content parts) and leaves content
empty. The streaming accumulators dropped that channel entirely, so a
streamed refusal assembled into an empty message and fell into the
empty/invalid-response retry loops - burning paid retries reproducing a
deterministic refusal - while the non-streaming path had already fixed
this class in #46013.
- chat_completions streaming: accumulate delta.refusal (incl.
model_extra), expose message.refusal on the assembled mock response so
ChatCompletionsTransport.normalize_response applies the existing
sole-payload -> content_filter promotion; count refusal deltas in the
zero-chunk guard; carry refusal in the Relay final-response dict.
- Codex Responses stream consumer: collect response.refusal.delta as
answer text so a refusal-only stream no longer raises 'did not emit a
terminal response' with zero usable content.
- Responses normalizer: read type=refusal content parts in
_extract_responses_message_text (attr and dict shapes).
Sabotage-verified: each new test fails with its wiring line disabled.
E2E: refusal-only stream -> terminal content_filter with explanation;
refusal-alongside-content stays a normal usable turn; plain-text
streams unchanged.
`_classify_responses_issuer` reimplemented endpoint canonicalisation with
its own urlsplit/urlunsplit pass. The repo already owns that logic in
`hermes_cli/route_identity.py::normalize_route_base_url` (stdlib-only,
used by agent/backend_identity.py), so delegate to it.
Reasoning items persisted before canonicalisation were stamped with the
raw `other:<agent.base_url>` (trailing slash, host case). Comparing them
verbatim against the now-canonical `current_issuer_kind` marked them
foreign and dropped them on the very endpoint that minted them. Run the
persisted stamp through the same canonicaliser (`_canonical_issuer_kind`,
non-`other:` kinds untouched) before comparing.
Also fixes the `_chat_messages_to_responses_input` docstring, which still
stated the pre-stack rule that legacy endpoint-stamped items drop when the
current model is known; the stack replays them on a matching issuer.
The openai SDK appends a trailing slash to `client.base_url`, so the aux
adapter stamped `other:https://h/v1/` while the main transport stamped
`other:https://h/v1`. On custom Responses endpoints every aux call
(compression, flush_memories) therefore dropped all main-minted reasoning
items as "foreign".
`_classify_responses_issuer` now strips whitespace and trailing slashes and
lowercases scheme+netloc before stamping. The aux adapter also derives its
route flags from `classify_responses_route` — the single owner of the
codex/xai/github predicates — instead of an inline chatgpt.com host check,
and reuses the same flags for the effort clamp.
Native compaction checkpoints and reasoning items persisted before model
stamping existed carry only `_issuer_kind`. Treating a missing `_issuer_model`
as foreign dropped every such item once the current model was known, which
wiped existing sessions' native-compaction context on upgrade (four consumer
tests in test_native_compaction / test_native_preflight_estimate /
test_413_compression went red on the stack).
Trust the endpoint stamp when no model stamp is present, as main does today.
Items minted after this change carry the model stamp and still drop on a
same-endpoint model switch; a wrong guess on a legacy item is caught by the
invalid_encrypted_content 400 classifier and the replay kill switch.
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.
Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.
Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.
Closes#95834
Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.
OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.
Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.
Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.
Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.
Fixes#100611
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Split _preflight_codex_input_items into per-item-type helpers over a
_PreflightCtx; streaming assembly moves into _CodexResponseAssembler with a
per-event dispatch table; run_codex_app_server_turn loses its session/usage/
interrupt regions to _ensure_codex_session/_finish_codex_turn/
_consume_user_interrupt/_queue_token_counts (counts built lazily so stub agents
without a session DB are never touched). Responses input/normalization and
stream callback order verified byte-identical against merge-base.
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.