Commit Graph

464 Commits

Author SHA1 Message Date
brooklyn!
1674499d00 fix(compression): archive the rows the compressor held, not every row up to the newest
A gap below the newest held id, and turns appended above an unpersisted current
turn, were summarized away without being read. Name those held ids and clone
the rest.
2026-09-24 11:57:56 -05:00
kshitijk4poor
b7f91258b9 fix(agent): tools-available note on dropped-tools continuation; keep legacy stub recognized (#74990)
The dropped-tools partial-stream continuation now also states the cut was a
transport interruption and tools remain available. The pre-rewording
network-stub text stays in the compressor's synthetic-turn set so
crash-persisted nudges from older sessions are not mistaken for user turns.
2026-09-24 22:27:18 +05:30
kshitijk4poor
2492193c44 fix(compression): a failed watermark read takes prune's logged no-op path
_archive_watermark_for sat outside prune's try, so a raising
get_active_message_watermark/get_message_role escaped prune_tool_results_only and
surfaced only at debug level in turn_preflight, skipping the "keeping the original
transcript" warning. Read it inside the existing try, with StaleHeldHistory caught
ahead of the generic handler. Also add the missing E302 blank line.
2026-09-24 18:01:21 +05:30
kshitijk4poor
40051c2a13 fix(compression): a stale prune/micro-compaction generation aborts instead of publishing beside the winner
Prune and micro-compaction hold no compression lease. When the newest exact row of the history they
hold is no longer active, another compaction (a /compress on this or another surface, or an earlier
pass) has already committed. `held_archive_watermark` then falls back to the start watermark, which is
right for the in-place commit (its lease rules out overlap) but for a lease-less caller archives the
winner's rows and clones them back as a "concurrent tail": two summary generations live in one session.

`_archive_watermark_for` now asks `held_archive_watermark` to raise `StaleHeldHistory` on that branch.
The prune commit returns its input unchanged (`prune:stale_generation`), micro-compaction checks before
the slow summary call and skips the pass (telemetry `stale_generation`), and its commit re-checks so a
compaction that lands during the summary call is not overwritten either. The in-place commit keeps its
fallback.

Raised on #120634 by @ehz0ah (agent/conversation_compression.py:3623 thread) and independently by
@JoaoMarcos44 in #120821.

Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 18:01:21 +05:30
John Paul Soliva
3f7e0fd072 fix(compression): proactive prune and micro-compaction keep turns another surface appended
Both commits call archive_and_compact with no watermark, which archives
every active row. A turn another surface appended to the same session
after this process loaded it (a Desktop session continued from Telegram)
was archived with the rest as compacted (active=0, compacted=1): marked
summarized away though no summary holds it. The display, REST with
include_compacted and session_search still show it, but it is gone from
the model's history, so the agent forgets a turn the user can still see.
Micro-compaction also runs a slow aux summary call before its commit, so
a turn that arrived during it went the same way.

Both now cap the archive at the newest row the process held, the rule the
in-place compaction commit applies, so those rows take the concurrent-
append path and are cloned after the new set. Micro-compaction captures
the held history and the store's watermark before the summary call and
caps at commit. The cap moves into held_archive_watermark(session_db,
session_id, ...), which the compressor paths can call; _held_watermark
stays as the in-place commit's wrapper.

A store without get_active_message_watermark or get_message_role keeps
today's archive-everything commit, as a store without archive_and_compact
already skips the prune. Micro-compaction's DB sync runs under a broad
except, so an unguarded call there would have failed silently.

Held lists without row ids (the gateway's replay dicts) keep today's
behaviour, as in the in-place commit. Both features are opt-in
(compression.proactive_prune_tokens > 0, compression.micro_compact).

(cherry picked from commit 8d67bee760db0fc65d4330df15198ddfd5f24480)
2026-09-24 18:01:21 +05:30
kshitijk4poor
159f716313 refactor(agent): derive the pruned-marker matcher from the compressor template
The dispatch-boundary detector re-typed the marker's wording next to the
producer's template and hid the import behind functools.lru_cache, citing a
context_compressor -> prompt_builder -> tool_dispatch_helpers cycle that does
not exist (prompt_builder only mentions this module in a comment; both import
orders succeed). Two copies of one string means a template edit silently
disables the guard.

Move the prefix/template into a dependency-free leaf, agent/compression_marker,
and build the regex from the template's first sentence with the count
placeholders swapped for \d[\d,]*. context_compressor re-exports the names, so
its callers and tests are unchanged. The matcher keeps the intended behaviour:
prefix-only mentions do not match; a minted marker, flat or nested, does. The
old \s+ tolerance for multiple spaces is dropped because the producer never
emits one, so only a template-shaped copy counts.

tool_dispatch_helpers stays light: importing it alone still does not load
context_compressor (auxiliary_client, context_engine, ...).
2026-09-24 16:29:19 +05:30
teknium1
ab2a206e5e fix(micro-compaction): show each input once after a superseded marker
Superseding a stale micro marker joins the now-adjacent user turns into one
model-facing row, while the originals stay in display history as compacted
rows, so resumed display history painted every merged input twice (and one
more time per later pass). Flag the join display_metadata.model_only and skip
it in every display projection: resume dedupe, the indexed and legacy
get_messages pages, and the prompt timeline. The model payload is unchanged.
2026-09-23 10:44:29 -07:00
kshitijk4poor
94e8a255df refactor(compression): one effective-cap helper; drop dead _config_threshold_percent fallback
`min(threshold_tokens_cap, context_length)` behind the same `cap is not None and cap > 0`
guard was written in both `_derive_trigger` and `_apply_threshold_tokens_cap`, contradicting
`_derive_trigger`'s "one place the trigger math lives". Both now call
`_effective_threshold_cap(context_length)`. The `getattr(self, "_config_threshold_percent",
...)` fallback was dead: `__init__` sets the attribute unconditionally and no bare-`__new__`
test instance reaches `_derive_trigger`. Behaviour-preserving; the cap tests still fail when
the helper is mutated to return None.
2026-09-23 00:10:27 +05:30
kshitijk4poor
e4f76b8277 refactor(agent): inline the single-use _finish closure in _sample_summary_records
`_finish` (agent/context_compressor.py:3586-3588) only wrapped
`_coverage` and re-derived `shown` from the final `selected`; it was
called once, at the return. With `_merged` hoisted above the slice loop
the classmethod now has two closures (`_merged`, `_render`) plus
`_coverage` instead of four. Same expressions, same evaluation order
(`_render` first, then the coverage counters), so the output and
coverage dict are byte-identical — verified on the 47-shape capture and
probe_C2 (0 violations, base identity x4 True).
2026-09-22 13:40:13 +05:30
kshitijk4poor
8e12cd92d7 refactor(agent): integer head/tail split in _bound_oversized_record
`_bound_oversized_record` (agent/context_compressor.py:3485-3488) split
`remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded
the tail slice with `if tail_len else ""`. `remaining = limit -
marker_reserve` is >= 1 because the line above already returned when
`limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1`
and `tail_len = remaining - head_len >= 1` always: the guard can never
fire. Use `remaining // 2` and drop the dead guard.

Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative
int (checked exhaustively for r in 0..10**6), and the guard was dead.
Capture of 800 bounded-record cases (4 record shapes x remaining 1..200)
plus the 47-shape sampling capture is identical before/after.
2026-09-22 13:40:13 +05:30
kshitijk4poor
e17d49174d refactor(agent): drop unreachable trailing-gap branch from lean-sampling _render
The `if cursor < len(records)` block at the end of `_render`
(agent/context_compressor.py:3580-3583) duplicated the gap-marker
arithmetic of the in-loop branch but had no path to it.

Invariant: `_render` is only reached with non-empty `records`
(empty returns early), and n >= 1, so the slice-build loop always runs
its last iteration, which anchors `end = len(records)` and
`start = end - 1 >= 0`, hence `end > start` and the final slice ends at
`len(records)`. `_merged` keeps `max(end)` for the last interval, and the
extension pass only grows the newest slice backward (`(s - 1, e)`) or
older slices forward but never past the next slice's start, so the last
slice's end stays `len(records)`. Therefore `cursor == len(records)`
after the loop for every input.

Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge
shapes incl. single-record and empty); probe_C2 -> 0 violations;
under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
56b63ecf82 refactor(agent): merge lean-sampling slices once via _merged instead of inline
The slice-build loop in _sample_summary_records hand-rolled the same
"overlapping/touching slice -> extend previous" fold that the `_merged`
closure 25 lines below implements (agent/context_compressor.py:3563-3568
vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end)
pairs, and merge once before the extension pass; the first `_render`
no longer wraps `selected` in a redundant `_merged` since it is already
merged. One interval-merge implementation instead of two.

Output is byte-identical: the merge is applied once to the same ordered
slice list, so the extension pass starts from the same `selected`.
Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes)
diffed before/after: identical sampled text and coverage; probe_C2
40 shapes -> 0 violations; under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
f3642a27d8 docs(agent): say why the extension pass re-renders after the cap pre-check
The comment claimed "marker widths shrink as records leave a gap" — but
shrinkage can only make the render smaller, which the pre-check already
tolerates. The exact re-render exists for the opposite case: a gap's
first index can gain a digit or thousands separator (999 -> 1,000, +2
chars) when the added record moves a marker, and that growth can push a
render sitting at cap over it. Comment only; no behaviour change.

Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619).
2026-09-22 13:40:13 +05:30
kshitijk4poor
be63d5bcd3 fix(agent): hand out lean-sampling headroom round-robin across slices
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.

Before -> after (records per slice, fill unchanged):
  10000x100  [1,1,1,1,1,1,1,8]      -> [1,2,2,2,2,2,2,2]      fill 0.944
   8000x100  [2,2,2,2,2,2,2,5]      -> [2,2,2,2,2,3,3,3]      fill 0.957
   4000x300  [4,4,4,4,4,4,4,11]     -> [4,5,5,5,5,5,5,5]      fill 0.987
   2000x600  [9,9,9,9,9,9,9,15]     -> [9,9,10,10,10,10,10,10] fill 0.995
    500x2000 [37,...,37,39]         -> [37,...,37,38,38]      fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.

Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).

Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
2026-09-22 13:40:13 +05:30
kshitijk4poor
c94fe7a259 docs(agent): say sampled_chars counts display chars in _sample_summary_records
The coverage docstring (agent/context_compressor.py:3510-3511) said the
counters count "record content only", but `sampled_chars` sums the
*display* records (post `_bound_oversized_record` truncation) while
`input_chars` sums the raw records. Spell that out so telemetry
consumers do not compute `omitted` two different ways (gate 2c
suggestion, L3512-3513/3524-3525).
2026-09-22 13:40:13 +05:30
kshitijk4poor
eb05bde6df refactor(agent): pre-declare summary_input_* telemetry keys
`_begin_compression_telemetry` (agent/context_compressor.py:1947-1960)
did not seed the six `summary_input_*` counters that
`_record_summary_input_coverage` adds on the lean path, so the emitted
`compression_attempt` line (conversation_compression.py:1451-1470) had a
different key set for lean and legacy attempts. Seed them as None so the
schema is stable; probe: lean and legacy attempts now emit identical key
sets (gate 2c suggestion, L1947-1960).
2026-09-22 13:40:13 +05:30
kshitijk4poor
949d1707f9 fix(agent): spend leftover lean-sampling budget on whole neighbouring records
The whole-record greedy fill in `_sample_summary_records`
(agent/context_compressor.py:3541,3555-3563) stops a slice as soon as the
next record would cross `target` (~20K), so each slice can be short by up
to one record. With mid-sized records that wastes a large share of the
160K cap: probe worst case (10K records) filled 90,660/160,000 (57%),
8K records 129,021 (81%); main filled 159,988 by splitting records.

Add a bounded extension pass after `selected` is built: newest slice
grows backward, older slices grow forward, one whole record at a time,
only while the exact rendered length stays <= cap and the slice does not
run into its neighbour (adjacent slices are merged so no separator is
lost). Records are never split and the cap is never exceeded.

After: worst random shape 150,581 (94%, remaining headroom < one 10K
record), 8K records 153,117 (96%), 1551-record region 159,996 with
968 whole records, 0 partial, 8 buckets within +-10%, telemetry
input_chars == sampled_chars + omitted_chars (probe: $D/fold/probe_C_fold.py).
The existing keeps_record_boundaries fixture (1.2K records) fills 99.1%
without this pass, so no fill assert was added there — it would be toothless.
2026-09-22 13:40:13 +05:30
kshitijk4poor
f94e296ae6 refactor(agent): remove unreachable overflow-trim loop in _sample_summary_records
Delete the post-render overflow-trim loop (agent/context_compressor.py
:3591-3607) and state the invariant in its place.

Why it is unreachable: every display record is pre-bounded to `target`
via _bound_oversized_record; each slice's greedy fill stops before
`size + sep + next > target`, so a slice renders to <= target chars and
the n slices to <= n*target (a merge of overlapping/adjacent slices is
their union: overlap only removes chars, adjacency adds one 2-char
separator but drops one marker of >= 2 chars). There are at most n-1
markers, each <= marker_len because marker_len is formatted with
first=last=len(records) and elided=total_len, the maxima. Since
target = (cap - (n-1)*marker_len) // n, the rendered length is
<= n*target + (n-1)*marker_len <= cap. Gate mutation M5 (loop disabled)
left all 7 sampler tests green; 30 random shapes never entered it.
2026-09-22 13:40:13 +05:30
kshitijk4poor
5aeef2889f refactor(agent): drop unreachable re-trim in _bound_oversized_record
Delete the `if len(res) > limit` re-trim at the end of
`_bound_oversized_record` (agent/context_compressor.py:3488-3489).

Why it can never fire: `marker_reserve` is the marker width formatted
with `elided=len(record)`, the largest value `elided` can take, so the
real marker is <= marker_reserve; `head` is `record[:head_len].rstrip`
(<= head_len) and `tail` is `record[-tail_len:].lstrip` (<= tail_len);
head_len + tail_len == limit - marker_reserve. Hence
len(head + marker + tail) <= limit always. The branch was untested and
unreachable (gate 2c W2).
2026-09-22 13:40:13 +05:30
kshitijk4poor
33a1d46dfc feat(agent): record summary-input coverage telemetry (from #118390)
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-22 13:40:13 +05:30
kshitijk4poor
a3a03d48d4 refactor(agent): sampler takes serialized records only
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
2026-09-22 13:40:13 +05:30
kshitijk4poor
e28157be98 fix(agent): keep _serialize_for_summary byte-identical to main
The record split added .rstrip("\n") to every serialized record. That
changes the bytes _serialize_for_summary returns for legacy tail_mode
and for agent/micro_compaction.py's _micro_summarize_one, which never
sample and never needed it: a message with trailing newlines was no
longer reproduced verbatim in the summarizer prompt. Sampling receives
records structurally, so an empty pseudo-record can no longer become the
tail anchor and the strip has no remaining purpose.
2026-09-22 13:40:13 +05:30
joaomarcos
04fe735c57 fix: preserve structural record framing in lean summary sampling
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.

- agent/context_compressor.py:
  * Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
  * Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
  * Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
  * Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
  * Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
  * Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
  * Test newest message ending in blank lines does not produce empty tail anchor.
  * Test exact omission marker separator accounting.
  * Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.

(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)
2026-09-22 13:40:13 +05:30
joaomarcos
723fc9ffb5 fix: bound oversized records and protect newest slice in lean summary sampling
(cherry picked from commit 1d0d938cc3572c992f3b6540486b98dbcc777429)
2026-09-22 13:40:13 +05:30
joaomarcos
3bdc07efef fix: sample complete summary records
(cherry picked from commit b557fc4534a4557f95540e29cb77ac8acfffa473)
2026-09-22 13:40:13 +05:30
kshitijk4poor
bc2373110c refactor(agent): share one _is_refusal_response predicate across both summarizers
The compound refusal check `_response_refusal_text(response) or
_is_summary_refusal(content)` was spelled out verbatim at both call sites
(agent/context_compressor.py:3639 in `_call_summary_llm` and
agent/micro_compaction.py:157 in `_micro_summarize_one`, the latter through
two separate `_cc()` facade lookups). `_classify_summary_failure` already
describes it as one concept ("refusal content (prose or provider `refusal`
field)"), and the ordering invariant — an explicit provider `refusal` field
wins over summary-shaped content — lived in two callers instead of one place.

Add `_is_refusal_response(response, content)` next to the detectors and call
it from both sites (micro: one `_cc()` lookup). `_is_summary_refusal` and
`_response_refusal_text` stay as-is for the unit tests. Behaviour unchanged.

Proof: guard + micro test files green (62 passed); mutating the new predicate
to `return False` fails exactly the batch refusal-body test, the provider
refusal-field test and the micro refusal test (3 failed, 59 passed).
2026-09-22 13:38:25 +05:30
kshitijk4poor
787e89ae69 refactor(agent): drop the unreachable None/str guard in _response_refusal_text
`_response_refusal_text` (agent/context_compressor.py:199-201) returned ""
early when `_coerce_llm_message` yielded None or a str. That branch is dead:
`_message_field` (agent/auxiliary_client.py:8014) is
`msg.get(name) if isinstance(msg, dict) else getattr(msg, name, None)`, and
`getattr(None, "refusal", None)` / `getattr("x", "refusal", None)` are both
None, which the `isinstance(refusal, str)` tail already maps to "". So the
guard could never change the result — invariant: a non-dict, non-object
message has no `refusal` attribute, so the field lookup is None.

Probe (None, "x", bare dict, dict-with-choices, choices=None, dict refusal,
str message, refusal=None) returns ['', '', 'no', 'nope', '', 'r', '', '']
both before and after this change; the guard test file stays green.
Repo convention: no defensive guards around code that cannot fail.
2026-09-22 13:38:25 +05:30
kshitijk4poor
6cdcf0f9fb docs(agent): note that refusal content rides the empty_content failure class
agent/context_compressor.py:_classify_summary_failure (L697 marker list)
maps the "refusal content" RuntimeError onto empty_content without saying
so in the docstring, which makes the "LLM returned empty content" fallback
log line look wrong when the real cause was a refusal. Document the
intentional reuse (cooldown + main-model fallback + abort) in one place.

Gate finding: B.2ab (misleading fallback log for refusals).
2026-09-22 13:38:25 +05:30
kshitijk4poor
94fbc155ec fix(agent): do not reject a preambled real summary as a refusal
agent/context_compressor.py:_is_summary_refusal (L177-184) matched on the
opener only, so a genuine templated summary that starts with a hedge such as
"I cannot see the earlier turns, but here is the summary:\n## Goal ..."
was classified as a refusal and compression aborted (gate probe: the
[preamble] compress() run aborted with "refusal content"). A refusal-only body
never carries the summary template's "## " section headings, while a real
summary always does, so exempt any body containing a "## " heading after the
prefix match.

Test: TestSummaryRefusalGuard::test_narrowing_and_heading_exemption. Red when
the heading exemption is removed (L91 assert True is False) and red when the
summary-term narrowing is dropped via an early `return True` (L85 assert True
is False, pins gate mutation M3); green at head (25 passed). Live probe:
[preamble] now compresses 25->7 with prev_summary_set=True.

Gate finding: B.2ab (false positive on prose-preambled summary).
2026-09-22 13:38:25 +05:30
kshitijk4poor
d5159aa31d fix(agent): catch "can not" and "summarise" in the summary-refusal detector
agent/context_compressor.py:_SUMMARY_REFUSAL_PREFIX_RE (L170) only matched
"cannot"/"can't", so the two-word "I can not provide a summary" slipped
through, and _is_summary_refusal (L184) looked for the exact terms
summary/summarize/checkpoint, so British "I can't summarise this" was
accepted as a checkpoint. Allow optional whitespace in can(?:\s*not|['’]t)
and match the stems ("summar", "checkpoint") instead.

Test: two new params on TestSummaryRefusalGuard::test_rejects_refusal_body.
Red with the previous prod file (2 failed: assert False is True), green at
head (24 passed).

Gate finding: B.2c (detector misses "can not" / "summarise").
2026-09-22 13:38:25 +05:30
kshitijk4poor
a8f87c397a refactor(agent): reuse auxiliary_client message helpers in _response_refusal_text
agent/context_compressor.py:_response_refusal_text (L187-201) re-implemented
the choices[0].message traversal that agent/auxiliary_client.py already
exposes as _coerce_llm_message (L7997) + _message_field (L8014) and that the
sibling extract_content_or_reasoning uses on the very same response. Use the
shared helpers so the two extractors cannot drift on response shape
(dict-shaped proxies vs ChatCompletion objects vs bare messages).

Behaviour kept: None / SDK object / dict-with-choices / missing attr /
refusal=None -> "" ; the dict-unwrap (message/reason/text) + isinstance(str)
tail is unchanged. Gate live probe ($D/gate/B_live_probe.py) output is
byte-identical before and after; the guard test file is green.

Gate finding: B.2c (duplicate response traversal).
2026-09-22 13:38:25 +05:30
kshitijk4poor
460fef87ca fix(agent): honour an explicit provider refusal field (from #118406)
OpenAI-style structured-output refusals arrive in
`choices[0].message.refusal` with `content` empty or filler, so the
prose detector never sees the refusal and the filler would be committed
as the compaction checkpoint. Read that field (str, or the dict shapes
some proxies emit) in both the batch (`_call_summary_llm`) and micro
(`_micro_summarize_one`) paths and route it through the same
"refusal content" failure so it can never become `_previous_summary`.

Also widen the prose opener with #118406's extra shapes — `we` as
subject, `refuse to`, `could not` / `couldn't`, `apologise`, and a bare
`I'm unable` / `I am not able` opener. #118406's `could(?:not|['’]t)`
missed "couldn't" (it only matched "couldnot"/"could't"); spelled out
here as `could\s*not|couldn['’]t`.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-22 13:38:25 +05:30
KoNit-K
42545ab0b0 fix(agent): guard micro-compaction refusals
(cherry picked from commit 1dc4bac1292c90b8d9e0b497d961bd2c383b50e2)
2026-09-22 13:38:25 +05:30
KoNit-K
f6f7b13f7c fix(agent): reject refusal-only context summaries
(cherry picked from commit 30eaf2bee34c2976115cdabf63afa2147cdd0c1f)
2026-09-22 13:38:25 +05:30
teknium1
b7803a1763 fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.

TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).

Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
  ctx    8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows;  window [4,10) -> [4,42)
  ctx   16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
  ctx   32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
  ctx  128K: identical before/after
2026-09-21 01:04:16 -07:00
Chukuwebuka-2003
7d3c0b2f94 fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.

Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.

Part of #116472 (request 4).
2026-09-20 18:53:10 -07:00
teknium1
0765099ff4 fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
2026-09-20 15:50:49 -07:00
beardthelion
0e33dc9ebc fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.

Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.

Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.

(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
2026-09-20 15:50:49 -07:00
kshitij
f88c6fc46e refactor(compression): finish the row-test dedupe and drop a gate that saved nothing
Two corrections to 0b9a9a0f5c.

The `if allow_split_turn else -1` gate did not skip the scan it claimed to:
`_ensure_last_user_message_in_tail` runs the identical lookup as its first
statement, so the micro-compaction pass paid the same scan and the batch path
paid it twice. Reverted to the unconditional call, which also removes the `-1`
sentinel every reader had to reason about.

The dedupe stopped at two of five spellings of the same predicate. Three more
sites inline it: the in-flight replay's "a real request follows the summary"
check, the handoff-candidate admission test (as its negation), and the merge
pre-check. All now call `_is_real_user_turn`. The one remaining inline pair is
deliberately different — it also requires `_is_real_user_message`, which rejects
metadata-flagged scaffolding this predicate cannot see.

Equivalence: all three predicates are pure, so the negated and reordered forms
are the same test; 221 tests pass across the compressor/anchor/micro-compaction
files, including the source-shape anchor-order test that the call shape here
leaves untouched.
2026-09-20 02:23:46 +05:30
kshitij
0b9a9a0f5c refactor(compression): dedupe the actionable-user row test, skip an unread scan
Follow-up to the oversized-turn exception merged in #116181.

- `_find_last_user_message_idx` and `_real_user_indices_desc` each spelled out
  the same actionable-and-not-synthetic predicate; both now call one
  `_is_real_user_turn`. No behaviour change — same two classmethods, same rows.
- The newest-user index is only read by the split exception, which rolling
  micro-compaction disables, so the scan is skipped on that pass instead of
  running and being discarded.

`_is_actionable_user_turn` / `_is_synthetic_compression_user_turn` are pure, so
the dedupe is equivalence by construction; the row set each scan returns is
unchanged.
2026-09-20 02:11:36 +05:30
fangliquanflq
af93d57ca7 fix(agent): preserve context when summary provider overloads 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
e1f46b8518 fix(compression): pin the shadowed-checkpoint prune through compress(); drop the tail-budget claim
Review of #115833: every new test called `_prune_stale_reasoning_replay`
directly, so deleting the production call in `compress()` left the file
green. Replace the redundant first test with one that drives
`ContextCompressor.compress()` (summary patched, 12-carrier transcript) and
asserts the returned transcript holds exactly the checkpoints
`prune_pre_checkpoint_items` would keep; the class is down to two tests.

The docstring claimed the shadowed copies were "charged in full by
_ALWAYS_REPLAYED_BUDGET_KEYS": the estimator already prices the sidecar
through `strip_opaque_replay_items` (562e6e4824), so the real effect is the
compacted transcript, child sessions and persisted rows. Factor the newest-
carrier rule into `drop_shadowed_checkpoints` so the durable prune can reuse
it instead of duplicating it.
2026-09-19 10:08:14 -07:00
joaomarcos
addb689a7b fix(compression): prune native-compaction checkpoints a newer carrier shadows
`_prune_stale_reasoning_replay` exempted every `type: "compaction"` item from
pruning ("They must survive on every retained message"). The wire builder
disagrees: `native_compaction.prune_pre_checkpoint_items` rebuilds each request
around the NEWEST checkpoint run and discards every earlier one, and
`codex_responses_adapter` drops checkpoints wholesale once native compaction is
no longer eligible. In both gate states a checkpoint shadowed by a newer
carrier has no reader on any wire.

It was still charged in full against the protected-tail budget
(`codex_reasoning_items` is in `_ALWAYS_REPLAYED_BUDGET_KEYS`, ~120 KB of
ciphertext each), copied into compacted and forked transcripts, and persisted
to `messages.codex_reasoning_items` for the life of the DB. Field forensics on
a 3232 MB `state.db`: 2473 MB sat in that one column, ~120 KB per assistant
row, 400 rows carrying only 118 distinct blobs.

Keep the newest carrier — the same item the wire builder would have picked —
and drop the shadowed ones. Single-checkpoint transcripts are untouched, so
every pre-existing test in the suite passes unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lqi8r25VMJv5TgGXYzHba3
(cherry picked from commit 00b0b7b672a2821a3baa62df70f052bc09f2a562)
2026-09-19 10:08:14 -07:00
kshitij
469296c748 refactor(compression): tighten the oversized-turn exception after review
Review findings folded into the exception added by the previous commit:

- The anchored-region check now reuses `_walk_tail_budget` instead of a
  second token sum with the default thought-charge rule. The walk charges
  thinking only on the newest assistant turn unless the route replays stale
  thinking (#73624/#84371), so the second rule could fire the exception on a
  region the walk itself considers inside the ceiling.
- Two conjuncts (`last_user_idx >= head_end`, `last_user_idx < cut_idx`) are
  implied by `user_anchored_cut < cut_idx`, which only changes when the
  anchor found a real user turn inside the compressible region; the check
  that depended on them is documented where it is read.
- The latest-assistant anchor is no longer computed and discarded under the
  exception (it also logged an anchor it never applied).
- `_find_last_user_message_idx` early-exits instead of materialising every
  actionable user index on the per-attempt boundary path.
- The exception logs at debug, like the sibling anchor decisions, rather
  than at info from both `_compress_window` and `has_content_to_compress`.

Adds the missing guard test for the new ceiling check: a transcript that
fits the tail budget must still anchor the active request rather than
triggering the exception. Red when the ceiling conjunct is removed.
2026-09-19 21:12:46 +05:30
kshitij
d32f162cad fix(compression): keep the tail anchors the split exception was bypassing
The oversized-turn exception (#80449) relaxed three tail anchors at once,
which voided two guarantees that hold on main:

- `compression.min_tail_user_messages` was skipped whenever the exception
  fired, so an N-user tail could come back with no user turn at all. The
  N-user anchor now always runs: the setting is a user-facing promise and
  outranks the budget.
- The exception fired even when the oversized weight was the active turn's
  own newest tool group. The pre-anchor cut retains that group anyway, so
  the split bought no reclaim while taking the active request out of the
  tail and losing the #10896 anchor. The exception now requires real turn
  body (a tool-call group) between the opening request and the cut.

Both failures are bound by tests, each red when its conjunct is removed:
`TestTailTokenBudgetCeiling::test_message_floor_does_not_unboundedly_override_soft_ceiling`
and `TestMinTailUserMessages::test_n_guarantee_wins_over_tail_token_budget_and_floor`
(each passes on main), plus `test_n_user_tail_guarantee_outranks_the_split`
for the N-user guarantee on a genuinely oversized turn.
2026-09-19 21:12:46 +05:30
embwl0x
cc4bb888a4 fix(compression): split oversized active turns 2026-09-19 21:12:46 +05:30
kshitijk4poor
39abfdf4ec fix(context-compressor): the marked-leaf guard must match the whole tail
Second review round on the previous commit, both findings reproduced:

- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
  about: a replayed leaf of head + marker followed by new content was never
  shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
  the marker to close the leaf, so that shape shrinks to head + marker with
  true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
  adds separator whitespace, so compact args with many keys could come back
  LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
  The helper now returns the caller's string unless the whole rewrite is a net
  reduction.

Tests cover both shapes plus a many-key compact payload.
2026-09-19 21:06:12 +05:30
kshitijk4poor
a5381a7dca fix(context-compressor): match the marker by position, and leave args byte-identical
Review findings on the previous commit, all reproduced:

- The "already marked" guard was a substring test, so a leaf that merely
  contains the marker — including one a model imitated into a new call, the
  #83714 failure mode itself — was exempt from shrinking forever. The marker
  is always written at `head_chars`, so the guard now tests that position: a
  1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
  JSON came back with inserted spaces and callers read it as a change
  (rewriting replayed history and counting a pressure hit for zero reclaim).
  Nothing replaced now returns the original string: a 547-char compact blob
  is byte-identical.
2026-09-19 21:06:12 +05:30
kshitijk4poor
1e216d1330 fix(context-compressor): only shrink args when it reclaims, never re-shrink
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:

- A leaf just over `head_chars` came back LONGER: a 628-char args blob
  shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
  compaction, so a leaf's marker was rewritten from the true count
  ("2,800 of 3,000 chars omitted") to a self-referential one
  ("223 of 423"), destroying the per-instance count property the marker
  relies on and churning those bytes on each pass.

A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
2026-09-19 21:06:12 +05:30