The native generateContent adapter never runs uncapped: when
model.max_tokens is unset it sends maxOutputTokens=65,535
(GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) because Gemini treats an omitted cap
as a low internal default. The context compressor's trigger is
pct×(window − max_tokens), and constructing it with max_tokens=None
reserved 0 — so on a 128K Gemma window the trigger landed at 98,304
while the real safe input budget was 65,537, and the provider 400'd
before compaction fired.
Live repro (real imports, temp HERMES_HOME, native Gemini base_url,
window=131072, max_tokens unset):
before: compressor.max_tokens=None, threshold_tokens=98304,
wire maxOutputTokens=65535 → trigger ABOVE the safe budget
after: compressor.max_tokens=65535, threshold_tokens=64000 → below it
Scoped to the native Gemini wiring (provider names + native base_url via
is_native_gemini_base_url; the /openai compat endpoint is excluded). The
generic provider-default reservation gap remains tracked in #63839.
Reported by @Artemonim in #57275 (residual claim 4).
model.ollama_num_ctx is resolved AFTER the context compressor is
constructed, so a config that sets only ollama_num_ctx (without
model.context_length) ran every request at the smaller served num_ctx
while the compressor still targeted the probed GGUF window (e.g. 256K
Gemma metadata). The compaction trigger then sat several times above the
window the server actually serves and never fired — reproducing the
original #57275 'blows past the limit' symptom on current main.
Live repro (real imports, temp HERMES_HOME, config = {model:
{ollama_num_ctx: 65536}}, probed window 262144):
before: _ollama_num_ctx=65536, compressor.context_length=262144,
threshold_tokens=196608 (300% of the served window)
after: compressor.context_length=65536, threshold below the window
The clamp is one-directional (a num_ctx larger than the resolved window
never inflates the compressor) and reuses update_model() so every
threshold-derived budget recalibrates. Overlaps #60103 (silent-clamp
dead zone) — this is the init-order half.
Reported by @Artemonim in #57275 (residual claim 3).
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.
Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.
Reported by @Artemonim in #57275 (residual claim 5).
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only
degraded to the 85% trigger when it met or exceeded the effective
window exactly (#14690). Near-minimum windows slipped through: at
context_length=65536 the threshold passed through at 64,000 — 97.7%
of the window, ~1.5K tokens of output room — so pre-API compaction
effectively could not fire.
Providers that silently truncate over-window prompts instead of
rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never
deliver the reactive context-overflow backstop either. Observed live
on a 65,536-token local model: the session rode into the window
ceiling and each length-continuation retry re-sent a window-filling
prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room)
until the turn died with "Response remained truncated after 4
continuation attempts" — every retry paying a full multi-minute
prefill.
Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the
effective input budget whenever the floor is the binding term. An
explicit threshold_percent above 85% is user intent and stays
uncapped; windows where the floor lands at/below the cap are
unchanged.
The isinstance(primary, Path) branch in _candidate_sentinel_paths was dead
weight: the surrounding except Exception already covers non-Path test
doubles, and .resolve() failing on them falls through to the plain
inequality comparison. Verified the pre-existing fail-safe stat fixture
(test_is_engaged_fails_safe_on_stat_error) still passes without it.
Module docstring still claimed 'a single os.stat'; the fleet-root check
makes it one or two stats. Updated.
Profile processes launch with HERMES_HOME=~/.hermes/profiles/<name>, so
`hermes pause` at the fleet root did not bind fleet-analyst dispatch
(t_7b65ff88). Check/resume both the process home and the fleet root.
The _CFG_SECRET_WORD_RE pre-gate only skips secret-FREE text. A compaction
payload containing one real secret assignment plus a long opaque dotted run
still reaches _CFG_DOTTED_RE's backtrackable '*' prefix, which re.sub retries
from every byte of the run — quadratic while holding the GIL (same class as
the _ENV_ASSIGN_LOWER_RE fix in this branch, #99255).
Anchor each attempt to the start of a key run with a negative lookbehind.
Match set is unchanged: any match starting mid-run implies a leftmost match
at the run start, verified 20/20 identical over a dotted-config corpus.
30k-char adversarial run: 102s -> 0.015s.
The Codex auxiliary Responses adapter enforced a single absolute
deadline (300s floor for compression). A dead stream held the entire
budget before fallback ran, and repeated compression attempts stacked
those waits into 20+ minute 'Summarizing thread' stalls (masoria debug
bundle, Aug 31 2026). Meanwhile a healthy-but-slow reasoning summary
was killed at the same absolute deadline even while producing tokens.
Replace the absolute kill with progress-aware deadlines:
- 60s no-progress window for the first substantive payload AND between
payloads; keepalive/lifecycle frames do not re-arm (mirrors the
commit-fence gating, #96707)
- a live stream re-arms per token and is bounded only by
_aux_stream_total_ceiling() (max(600s, 4x configured timeout)), the
same backstop the streamed chat.completions path already uses
- the compression critical-path retry gate now distinguishes failure
cost: a cheap first-token no-progress failure retries the same
provider once; mid-stream stalls and ceiling hits still skip straight
to provider fallback (#54465 semantics preserved)
Live A/B (real OpenAI SDK against a local SSE server, real adapter):
dead keepalive-only stream: main waits the full budget; fixed fails
over at the window. Slow-but-alive stream (tokens past the configured
timeout): main kills it mid-generation; fixed completes.
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.
Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
main-model fallback (a larger output budget may finish the summary), and
on terminal failure ABORTS compression preserving the session unchanged
(new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
existing retry/backoff loop.
_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.
Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.
Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:
- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
alias map THIS request emitted; the transport stashes it
(_last_wire_aliases) and normalize_response reverses ONLY those aliases.
A real user/plugin/MCP tool named hermes_tool_search is never silently
dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
a prior request can't leak into the next response's dispatch.
Refs #95003
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
xAI reserves the function name `tool_search` for Grok's native
server-side Tool Search and rejects the client declaration outright:
HTTP 400 {"code":"invalid-argument","error":"The function name
tool_search is reserved for the tool_search tool"}
Hermes's progressive-disclosure bridge registers exactly that literal
(`TOOL_SEARCH_NAME` in tools/tool_search.py) and assembly is not
provider gated, so with the default `tools.tool_search.enabled: auto`
every grok turn fails the moment the catalog crosses the threshold —
mid-session, which reads to the user as a session reset.
Same treatment as the two collisions already handled on this
transport (xAI `web_search` #48108, OpenCode reserved names #85589):
alias to `hermes_tool_search` on the wire in build_kwargs, map back in
normalize_response so Hermes dispatch and the bridge contract are
untouched. `tool_describe` / `tool_call` are not reserved by xAI and
are left alone.
Folds the per-provider rename helpers into one `_alias_reserved_tools`
owner parameterized by the reserved-name tuple, and extends the
existing `_RESERVED_ALIAS_TO_NAME` reverse map so the dispatch-side
un-aliasing needs no new branch.
Scope note: this covers the Responses transport, which is where every
api.x.ai route lands by default (`_fallback_api_mode` maps api.x.ai →
codex_responses, and the xai provider profile declares it). An xAI
model forced onto `api_mode: chat_completions` would still hit the
400; that path has no provider-specific tool rewriting today and would
need the symmetric hook in agent/transports/chat_completions.py. Happy
to add it here if you'd rather have both in one change.
Tests: new TestXaiReservedToolSearchAlias covering the wire alias,
non-xAI backends keeping the canonical name, composition with the
native web_search swap, and the normalize_response round trip.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012vLaAmnsdii3Gm9jMDs5gw
The curator LLM fork was steered by its own prompt to re-home skill
support files with terminal `mkdir -p ... && mv ...`. A terminal move
writes the same bytes with NO ledger entry, so the archive that follows
snapshots an already-stripped package (files: 1) and `hermes curator
rollback` restores a hollow skill — SKILL.md back, references/ gone.
Remove the capability rather than guard it: the fork's enabled_toolsets
drops "terminal", so terminal and process disappear together and there
is no shell to parse, no process stdin to feed, no remote-backend
divergence — a heuristic command guard over a Turing-complete input
space can guarantee none of that. Every mutation the pass needs has a
ledgered skill_manage action (write_file / remove_file / delete), and
the prompt now steers exactly those. Reading works through skill_view.
Tests pin both halves: the call-site kwarg (["skills"] only), the
resolved surface (no execution/write tools), and the prompt steering
(no mkdir -p / mv shapes).
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.
Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.
Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
(_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
message);
- micro-compaction splices [prefix, marker, suffix] - everything except
the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
keeping the historical archive-everything behavior.
Fixes#86366
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).
Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.
TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
session_reset (deliberate boundary) => still aborts BEFORE the #47202
flush, preserving #88411's no-growth contract where an abort remains
correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
MemoryManager.on_pre_compress() detects checkpoint API v2 providers,
selects the normalized evidence list for them, and re-raises their
failures under require_checkpoint — but it never tells the provider
that a checkpoint is required: the call passes only the messages.
A v2 provider therefore runs in its default best-effort mode, swallows
durable-write failures, and returns normally; the host then treats the
checkpoint as succeeded and lossy compression proceeds. With
compression.checkpoint_required: true this silently defeats the
guarantee the option exists to provide.
Forward require_checkpoint only to providers advertising the requested
checkpoint API version. Legacy providers keep the strict one-argument
on_pre_compress(self, messages) contract, so bundled v1 providers
(honcho, mem0, supermemory, ...) are unaffected.
Regression tests cover required and best-effort signaling, legacy
signature compatibility, and required-mode failure propagation.
Idle and preflight compaction arrived as lifecycle status without the
"Compacting context" marker, so TUI never entered a compacting state.
Re-tag those lines and freeze the busy FaceTicker on "compacting" for
the whole pause instead of restoring "running…" after 4s.
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
comment pinning WHY it is the timeout ladder's first rung (transient
lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
test_compression_rotation_state.py / test_context_compressor.py; kept the
force-bypass test (only site pinning that behavior for split failures) and
the E2E test (now asserting the named constant).
The cooldown is recorded inside the split-failure except handler; a raw
call on a stub/partial compressor would replace the real split error with
an AttributeError. Mirror the try/except-debug convention of the adjacent
record_rejected_compaction call.
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):
1. publish_compression_child gains require_lease_refresh: the lease is
extended inside the same transaction as the expiry check (same conn, no
TOCTOU), giving a worker whose refresher thread died from transient DB
failures one final chance to keep its completed work.
2. A failed compression split now records a 60s failure cooldown, so the
next turn cannot immediately re-trigger the same compression.
Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
A ceiling/idle-timeout host now joins its fence-cancelled worker for a
bounded grace before returning. A cooperative worker (which polls the
poison fence between provider phases) is reaped, proving quiescence, so
the durable lease releases normally. An uninterruptible worker is
orphaned behind the poison fence: its late result is discarded, and on
the total-ceiling path the holder-qualified lease stays retained until
it exits so no new attempt can overlap the unchanged session.
Supersession: a late candidate from an attempt whose compressor
generation was claimed by a newer attempt is discarded before the
commit boundary (failure_class=attempt_superseded), never committed
over newer state — covering fenceless callers the fence poison cannot
see.
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.
Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.
Refs #90976 and #91477.
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).