Superseding a stale micro marker joins the now-adjacent user turns into one
model-facing row, while the originals stay in display history as compacted
rows, so resumed display history painted every merged input twice (and one
more time per later pass). Flag the join display_metadata.model_only and skip
it in every display projection: resume dedupe, the indexed and legacy
get_messages pages, and the prompt timeline. The model payload is unchanged.
`min(threshold_tokens_cap, context_length)` behind the same `cap is not None and cap > 0`
guard was written in both `_derive_trigger` and `_apply_threshold_tokens_cap`, contradicting
`_derive_trigger`'s "one place the trigger math lives". Both now call
`_effective_threshold_cap(context_length)`. The `getattr(self, "_config_threshold_percent",
...)` fallback was dead: `__init__` sets the attribute unconditionally and no bare-`__new__`
test instance reaches `_derive_trigger`. Behaviour-preserving; the cap tests still fail when
the helper is mutated to return None.
`_finish` (agent/context_compressor.py:3586-3588) only wrapped
`_coverage` and re-derived `shown` from the final `selected`; it was
called once, at the return. With `_merged` hoisted above the slice loop
the classmethod now has two closures (`_merged`, `_render`) plus
`_coverage` instead of four. Same expressions, same evaluation order
(`_render` first, then the coverage counters), so the output and
coverage dict are byte-identical — verified on the 47-shape capture and
probe_C2 (0 violations, base identity x4 True).
`_bound_oversized_record` (agent/context_compressor.py:3485-3488) split
`remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded
the tail slice with `if tail_len else ""`. `remaining = limit -
marker_reserve` is >= 1 because the line above already returned when
`limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1`
and `tail_len = remaining - head_len >= 1` always: the guard can never
fire. Use `remaining // 2` and drop the dead guard.
Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative
int (checked exhaustively for r in 0..10**6), and the guard was dead.
Capture of 800 bounded-record cases (4 record shapes x remaining 1..200)
plus the 47-shape sampling capture is identical before/after.
The `if cursor < len(records)` block at the end of `_render`
(agent/context_compressor.py:3580-3583) duplicated the gap-marker
arithmetic of the in-loop branch but had no path to it.
Invariant: `_render` is only reached with non-empty `records`
(empty returns early), and n >= 1, so the slice-build loop always runs
its last iteration, which anchors `end = len(records)` and
`start = end - 1 >= 0`, hence `end > start` and the final slice ends at
`len(records)`. `_merged` keeps `max(end)` for the last interval, and the
extension pass only grows the newest slice backward (`(s - 1, e)`) or
older slices forward but never past the next slice's start, so the last
slice's end stays `len(records)`. Therefore `cursor == len(records)`
after the loop for every input.
Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge
shapes incl. single-record and empty); probe_C2 -> 0 violations;
under-cap base identity x4 True.
The slice-build loop in _sample_summary_records hand-rolled the same
"overlapping/touching slice -> extend previous" fold that the `_merged`
closure 25 lines below implements (agent/context_compressor.py:3563-3568
vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end)
pairs, and merge once before the extension pass; the first `_render`
no longer wraps `selected` in a redundant `_merged` since it is already
merged. One interval-merge implementation instead of two.
Output is byte-identical: the merge is applied once to the same ordered
slice list, so the extension pass starts from the same `selected`.
Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes)
diffed before/after: identical sampled text and coverage; probe_C2
40 shapes -> 0 violations; under-cap base identity x4 True.
The comment claimed "marker widths shrink as records leave a gap" — but
shrinkage can only make the render smaller, which the pre-check already
tolerates. The exact re-render exists for the opposite case: a gap's
first index can gain a digit or thousands separator (999 -> 1,000, +2
chars) when the added record moves a marker, and that growth can push a
render sitting at cap over it. Comment only; no behaviour change.
Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619).
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.
Before -> after (records per slice, fill unchanged):
10000x100 [1,1,1,1,1,1,1,8] -> [1,2,2,2,2,2,2,2] fill 0.944
8000x100 [2,2,2,2,2,2,2,5] -> [2,2,2,2,2,3,3,3] fill 0.957
4000x300 [4,4,4,4,4,4,4,11] -> [4,5,5,5,5,5,5,5] fill 0.987
2000x600 [9,9,9,9,9,9,9,15] -> [9,9,10,10,10,10,10,10] fill 0.995
500x2000 [37,...,37,39] -> [37,...,37,38,38] fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.
Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).
Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
The coverage docstring (agent/context_compressor.py:3510-3511) said the
counters count "record content only", but `sampled_chars` sums the
*display* records (post `_bound_oversized_record` truncation) while
`input_chars` sums the raw records. Spell that out so telemetry
consumers do not compute `omitted` two different ways (gate 2c
suggestion, L3512-3513/3524-3525).
`_begin_compression_telemetry` (agent/context_compressor.py:1947-1960)
did not seed the six `summary_input_*` counters that
`_record_summary_input_coverage` adds on the lean path, so the emitted
`compression_attempt` line (conversation_compression.py:1451-1470) had a
different key set for lean and legacy attempts. Seed them as None so the
schema is stable; probe: lean and legacy attempts now emit identical key
sets (gate 2c suggestion, L1947-1960).
The whole-record greedy fill in `_sample_summary_records`
(agent/context_compressor.py:3541,3555-3563) stops a slice as soon as the
next record would cross `target` (~20K), so each slice can be short by up
to one record. With mid-sized records that wastes a large share of the
160K cap: probe worst case (10K records) filled 90,660/160,000 (57%),
8K records 129,021 (81%); main filled 159,988 by splitting records.
Add a bounded extension pass after `selected` is built: newest slice
grows backward, older slices grow forward, one whole record at a time,
only while the exact rendered length stays <= cap and the slice does not
run into its neighbour (adjacent slices are merged so no separator is
lost). Records are never split and the cap is never exceeded.
After: worst random shape 150,581 (94%, remaining headroom < one 10K
record), 8K records 153,117 (96%), 1551-record region 159,996 with
968 whole records, 0 partial, 8 buckets within +-10%, telemetry
input_chars == sampled_chars + omitted_chars (probe: $D/fold/probe_C_fold.py).
The existing keeps_record_boundaries fixture (1.2K records) fills 99.1%
without this pass, so no fill assert was added there — it would be toothless.
Delete the post-render overflow-trim loop (agent/context_compressor.py
:3591-3607) and state the invariant in its place.
Why it is unreachable: every display record is pre-bounded to `target`
via _bound_oversized_record; each slice's greedy fill stops before
`size + sep + next > target`, so a slice renders to <= target chars and
the n slices to <= n*target (a merge of overlapping/adjacent slices is
their union: overlap only removes chars, adjacency adds one 2-char
separator but drops one marker of >= 2 chars). There are at most n-1
markers, each <= marker_len because marker_len is formatted with
first=last=len(records) and elided=total_len, the maxima. Since
target = (cap - (n-1)*marker_len) // n, the rendered length is
<= n*target + (n-1)*marker_len <= cap. Gate mutation M5 (loop disabled)
left all 7 sampler tests green; 30 random shapes never entered it.
Delete the `if len(res) > limit` re-trim at the end of
`_bound_oversized_record` (agent/context_compressor.py:3488-3489).
Why it can never fire: `marker_reserve` is the marker width formatted
with `elided=len(record)`, the largest value `elided` can take, so the
real marker is <= marker_reserve; `head` is `record[:head_len].rstrip`
(<= head_len) and `tail` is `record[-tail_len:].lstrip` (<= tail_len);
head_len + tail_len == limit - marker_reserve. Hence
len(head + marker + tail) <= limit always. The branch was untested and
unreachable (gate 2c W2).
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
The record split added .rstrip("\n") to every serialized record. That
changes the bytes _serialize_for_summary returns for legacy tail_mode
and for agent/micro_compaction.py's _micro_summarize_one, which never
sample and never needed it: a message with trailing newlines was no
longer reproduced verbatim in the summarizer prompt. Sampling receives
records structurally, so an empty pseudo-record can no longer become the
tail anchor and the strip has no remaining purpose.
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.
- agent/context_compressor.py:
* Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
* Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
* Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
* Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
* Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
* Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
* Test newest message ending in blank lines does not produce empty tail anchor.
* Test exact omission marker separator accounting.
* Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.
(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)
The compound refusal check `_response_refusal_text(response) or
_is_summary_refusal(content)` was spelled out verbatim at both call sites
(agent/context_compressor.py:3639 in `_call_summary_llm` and
agent/micro_compaction.py:157 in `_micro_summarize_one`, the latter through
two separate `_cc()` facade lookups). `_classify_summary_failure` already
describes it as one concept ("refusal content (prose or provider `refusal`
field)"), and the ordering invariant — an explicit provider `refusal` field
wins over summary-shaped content — lived in two callers instead of one place.
Add `_is_refusal_response(response, content)` next to the detectors and call
it from both sites (micro: one `_cc()` lookup). `_is_summary_refusal` and
`_response_refusal_text` stay as-is for the unit tests. Behaviour unchanged.
Proof: guard + micro test files green (62 passed); mutating the new predicate
to `return False` fails exactly the batch refusal-body test, the provider
refusal-field test and the micro refusal test (3 failed, 59 passed).
`_response_refusal_text` (agent/context_compressor.py:199-201) returned ""
early when `_coerce_llm_message` yielded None or a str. That branch is dead:
`_message_field` (agent/auxiliary_client.py:8014) is
`msg.get(name) if isinstance(msg, dict) else getattr(msg, name, None)`, and
`getattr(None, "refusal", None)` / `getattr("x", "refusal", None)` are both
None, which the `isinstance(refusal, str)` tail already maps to "". So the
guard could never change the result — invariant: a non-dict, non-object
message has no `refusal` attribute, so the field lookup is None.
Probe (None, "x", bare dict, dict-with-choices, choices=None, dict refusal,
str message, refusal=None) returns ['', '', 'no', 'nope', '', 'r', '', '']
both before and after this change; the guard test file stays green.
Repo convention: no defensive guards around code that cannot fail.
agent/context_compressor.py:_classify_summary_failure (L697 marker list)
maps the "refusal content" RuntimeError onto empty_content without saying
so in the docstring, which makes the "LLM returned empty content" fallback
log line look wrong when the real cause was a refusal. Document the
intentional reuse (cooldown + main-model fallback + abort) in one place.
Gate finding: B.2ab (misleading fallback log for refusals).
agent/context_compressor.py:_is_summary_refusal (L177-184) matched on the
opener only, so a genuine templated summary that starts with a hedge such as
"I cannot see the earlier turns, but here is the summary:\n## Goal ..."
was classified as a refusal and compression aborted (gate probe: the
[preamble] compress() run aborted with "refusal content"). A refusal-only body
never carries the summary template's "## " section headings, while a real
summary always does, so exempt any body containing a "## " heading after the
prefix match.
Test: TestSummaryRefusalGuard::test_narrowing_and_heading_exemption. Red when
the heading exemption is removed (L91 assert True is False) and red when the
summary-term narrowing is dropped via an early `return True` (L85 assert True
is False, pins gate mutation M3); green at head (25 passed). Live probe:
[preamble] now compresses 25->7 with prev_summary_set=True.
Gate finding: B.2ab (false positive on prose-preambled summary).
agent/context_compressor.py:_SUMMARY_REFUSAL_PREFIX_RE (L170) only matched
"cannot"/"can't", so the two-word "I can not provide a summary" slipped
through, and _is_summary_refusal (L184) looked for the exact terms
summary/summarize/checkpoint, so British "I can't summarise this" was
accepted as a checkpoint. Allow optional whitespace in can(?:\s*not|['’]t)
and match the stems ("summar", "checkpoint") instead.
Test: two new params on TestSummaryRefusalGuard::test_rejects_refusal_body.
Red with the previous prod file (2 failed: assert False is True), green at
head (24 passed).
Gate finding: B.2c (detector misses "can not" / "summarise").
agent/context_compressor.py:_response_refusal_text (L187-201) re-implemented
the choices[0].message traversal that agent/auxiliary_client.py already
exposes as _coerce_llm_message (L7997) + _message_field (L8014) and that the
sibling extract_content_or_reasoning uses on the very same response. Use the
shared helpers so the two extractors cannot drift on response shape
(dict-shaped proxies vs ChatCompletion objects vs bare messages).
Behaviour kept: None / SDK object / dict-with-choices / missing attr /
refusal=None -> "" ; the dict-unwrap (message/reason/text) + isinstance(str)
tail is unchanged. Gate live probe ($D/gate/B_live_probe.py) output is
byte-identical before and after; the guard test file is green.
Gate finding: B.2c (duplicate response traversal).
OpenAI-style structured-output refusals arrive in
`choices[0].message.refusal` with `content` empty or filler, so the
prose detector never sees the refusal and the filler would be committed
as the compaction checkpoint. Read that field (str, or the dict shapes
some proxies emit) in both the batch (`_call_summary_llm`) and micro
(`_micro_summarize_one`) paths and route it through the same
"refusal content" failure so it can never become `_previous_summary`.
Also widen the prose opener with #118406's extra shapes — `we` as
subject, `refuse to`, `could not` / `couldn't`, `apologise`, and a bare
`I'm unable` / `I am not able` opener. #118406's `could(?:not|['’]t)`
missed "couldn't" (it only matched "couldnot"/"could't"); spelled out
here as `could\s*not|couldn['’]t`.
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.
TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).
Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
ctx 8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows; window [4,10) -> [4,42)
ctx 16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
ctx 32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
ctx 128K: identical before/after
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.
Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.
Part of #116472 (request 4).
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.
Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.
Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.
(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
Two corrections to 0b9a9a0f5c.
The `if allow_split_turn else -1` gate did not skip the scan it claimed to:
`_ensure_last_user_message_in_tail` runs the identical lookup as its first
statement, so the micro-compaction pass paid the same scan and the batch path
paid it twice. Reverted to the unconditional call, which also removes the `-1`
sentinel every reader had to reason about.
The dedupe stopped at two of five spellings of the same predicate. Three more
sites inline it: the in-flight replay's "a real request follows the summary"
check, the handoff-candidate admission test (as its negation), and the merge
pre-check. All now call `_is_real_user_turn`. The one remaining inline pair is
deliberately different — it also requires `_is_real_user_message`, which rejects
metadata-flagged scaffolding this predicate cannot see.
Equivalence: all three predicates are pure, so the negated and reordered forms
are the same test; 221 tests pass across the compressor/anchor/micro-compaction
files, including the source-shape anchor-order test that the call shape here
leaves untouched.
Follow-up to the oversized-turn exception merged in #116181.
- `_find_last_user_message_idx` and `_real_user_indices_desc` each spelled out
the same actionable-and-not-synthetic predicate; both now call one
`_is_real_user_turn`. No behaviour change — same two classmethods, same rows.
- The newest-user index is only read by the split exception, which rolling
micro-compaction disables, so the scan is skipped on that pass instead of
running and being discarded.
`_is_actionable_user_turn` / `_is_synthetic_compression_user_turn` are pure, so
the dedupe is equivalence by construction; the row set each scan returns is
unchanged.
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).
Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Review of #115833: every new test called `_prune_stale_reasoning_replay`
directly, so deleting the production call in `compress()` left the file
green. Replace the redundant first test with one that drives
`ContextCompressor.compress()` (summary patched, 12-carrier transcript) and
asserts the returned transcript holds exactly the checkpoints
`prune_pre_checkpoint_items` would keep; the class is down to two tests.
The docstring claimed the shadowed copies were "charged in full by
_ALWAYS_REPLAYED_BUDGET_KEYS": the estimator already prices the sidecar
through `strip_opaque_replay_items` (562e6e4824), so the real effect is the
compacted transcript, child sessions and persisted rows. Factor the newest-
carrier rule into `drop_shadowed_checkpoints` so the durable prune can reuse
it instead of duplicating it.
`_prune_stale_reasoning_replay` exempted every `type: "compaction"` item from
pruning ("They must survive on every retained message"). The wire builder
disagrees: `native_compaction.prune_pre_checkpoint_items` rebuilds each request
around the NEWEST checkpoint run and discards every earlier one, and
`codex_responses_adapter` drops checkpoints wholesale once native compaction is
no longer eligible. In both gate states a checkpoint shadowed by a newer
carrier has no reader on any wire.
It was still charged in full against the protected-tail budget
(`codex_reasoning_items` is in `_ALWAYS_REPLAYED_BUDGET_KEYS`, ~120 KB of
ciphertext each), copied into compacted and forked transcripts, and persisted
to `messages.codex_reasoning_items` for the life of the DB. Field forensics on
a 3232 MB `state.db`: 2473 MB sat in that one column, ~120 KB per assistant
row, 400 rows carrying only 118 distinct blobs.
Keep the newest carrier — the same item the wire builder would have picked —
and drop the shadowed ones. Single-checkpoint transcripts are untouched, so
every pre-existing test in the suite passes unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lqi8r25VMJv5TgGXYzHba3
(cherry picked from commit 00b0b7b672a2821a3baa62df70f052bc09f2a562)
Review findings folded into the exception added by the previous commit:
- The anchored-region check now reuses `_walk_tail_budget` instead of a
second token sum with the default thought-charge rule. The walk charges
thinking only on the newest assistant turn unless the route replays stale
thinking (#73624/#84371), so the second rule could fire the exception on a
region the walk itself considers inside the ceiling.
- Two conjuncts (`last_user_idx >= head_end`, `last_user_idx < cut_idx`) are
implied by `user_anchored_cut < cut_idx`, which only changes when the
anchor found a real user turn inside the compressible region; the check
that depended on them is documented where it is read.
- The latest-assistant anchor is no longer computed and discarded under the
exception (it also logged an anchor it never applied).
- `_find_last_user_message_idx` early-exits instead of materialising every
actionable user index on the per-attempt boundary path.
- The exception logs at debug, like the sibling anchor decisions, rather
than at info from both `_compress_window` and `has_content_to_compress`.
Adds the missing guard test for the new ceiling check: a transcript that
fits the tail budget must still anchor the active request rather than
triggering the exception. Red when the ceiling conjunct is removed.
The oversized-turn exception (#80449) relaxed three tail anchors at once,
which voided two guarantees that hold on main:
- `compression.min_tail_user_messages` was skipped whenever the exception
fired, so an N-user tail could come back with no user turn at all. The
N-user anchor now always runs: the setting is a user-facing promise and
outranks the budget.
- The exception fired even when the oversized weight was the active turn's
own newest tool group. The pre-anchor cut retains that group anyway, so
the split bought no reclaim while taking the active request out of the
tail and losing the #10896 anchor. The exception now requires real turn
body (a tool-call group) between the opening request and the cut.
Both failures are bound by tests, each red when its conjunct is removed:
`TestTailTokenBudgetCeiling::test_message_floor_does_not_unboundedly_override_soft_ceiling`
and `TestMinTailUserMessages::test_n_guarantee_wins_over_tail_token_budget_and_floor`
(each passes on main), plus `test_n_user_tail_guarantee_outranks_the_split`
for the N-user guarantee on a genuinely oversized turn.
Second review round on the previous commit, both findings reproduced:
- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
about: a replayed leaf of head + marker followed by new content was never
shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
the marker to close the leaf, so that shape shrinks to head + marker with
true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
adds separator whitespace, so compact args with many keys could come back
LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
The helper now returns the caller's string unless the whole rewrite is a net
reduction.
Tests cover both shapes plus a many-key compact payload.
Review findings on the previous commit, all reproduced:
- The "already marked" guard was a substring test, so a leaf that merely
contains the marker — including one a model imitated into a new call, the
#83714 failure mode itself — was exempt from shrinking forever. The marker
is always written at `head_chars`, so the guard now tests that position: a
1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
JSON came back with inserted spaces and callers read it as a change
(rewriting replayed history and counting a pressure hit for zero reclaim).
Nothing replaced now returns the original string: a 547-char compact blob
is byte-identical.
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:
- A leaf just over `head_chars` came back LONGER: a 628-char args blob
shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
compaction, so a leaf's marker was rewritten from the true count
("2,800 of 3,000 chars omitted") to a self-referential one
("223 of 423"), destroying the per-instance count property the marker
relies on and churning those bytes on each pass.
A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
Root cause for #83714 (write_file/patch_tool writing literal
"...[truncated]" into files, PR #83752's guard is the safety net, not
the fix): _truncate_tool_call_args_json() in the compression pass
shrinks long string values inside a PAST assistant message's
tool_calls[].function.arguments — the exact field that represents the
model's own prior generated output, replayed back to it verbatim on
every subsequent turn. The old marker, a bare "...[truncated]" suffix,
is indistinguishable from something the model itself could have
written (it's exactly the kind of terse ellipsis abbreviation models
already produce). A model conditioned on seeing itself "get away with"
that pattern in its own history imitates it in a new tool call,
writing the literal marker instead of real content.
This is the second bug from the same root text. The first (#11762,
MiniMax 400s from unterminated JSON) was fixed by shrinking inside the
parsed structure so the JSON stays valid, but kept the same visible
marker text — fixing the syntax problem while leaving the imitation
problem untouched.
Fix: replace the marker with one deliberately NOT shaped like prose a
model would write — distinctive non-ASCII delimiters, an explicit "not
part of the original tool call" disclaimer, and a per-instance
char-count that won't match the next omission point even if copied
verbatim. The shrunk value stays a plain string (not a nested object)
so the #11762 valid-JSON/matching-shape contract is unchanged — only
the marker text changed.
Checked context_compressor.py's other "...[truncated]" call sites
(_serialize_for_summary, _compact_fallback_turn, the user-message-only
one near _ACTIVE_TASK_MAX_CHARS) — none of them write into a value
that gets replayed as the main model's own assistant/tool_calls
history, so they don't share this priming risk and were left as-is.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
`preview_threshold_tokens` restated the resolve -> floor -> compute -> cap chain that `update_model`
runs; two copies of the trigger math drift the next time a step is added — the exact bug class
#83450 fixes (the guard quoting a number the compressor will not install). `_derive_trigger` is the
single pure derivation; the auxiliary-summariser ceiling stays in `_apply_threshold_tokens_cap`
because it is per-runtime, not per-model.
The startup banner names the cap only when it set the trigger; on windows where the ratio already
sits below it "(capped at 256,000)" was noise. Comments no longer repeat the default literal.
delegation.compression_threshold_tokens and _apply_child_compression_cap both described a 1M-window
child compacting at 500K "like its parent"; with compression.threshold_tokens defaulting to 256K the
parent (and so the child, via the min of both caps) compacts at 256K. preview_threshold_tokens now
says it ignores the auxiliary-summariser ceiling, which only the post-switch probe can know.
The preflight-compression warning computed `context_length × threshold_percent` itself, so with the
absolute `threshold_tokens` cap (and provider-scoped `model_thresholds`, the small-window floor) it
quoted a trigger the compressor would never use — "auto-compress at ~500,000" on a 1M switch that
really compacts at 256,000. `ContextCompressor.preview_threshold_tokens()` now computes the post-switch
trigger with the same machinery `update_model` uses, without mutating state; the guard asks for it and
keeps the plain ratio only for duck-typed engines that lack the method.
Salvaged from #83523 (commits bb210c09a7 + a6918d8b1b squashed; the ContextEngine ABC default was
dropped — the guard's fallback already covers engines without a preview).
The two walks differ only in the break predicate's `(n - i) >= min_tail` conjunct, so when the
ceiling can hold the floor's wire overhead the floorless walk's cut is exactly what the fix chose
on every branch (identical break row when the floor fits; the bounded cut when it does not; the
#40803 re-cut already returns a cut at or below the floor). Verified by a 40,000-transcript
randomized differential probe against the two-walk form: 0 mismatches.
Also drops the second O(n) estimate pass from the micro-compaction per-turn path.