Commit Graph

429 Commits

Author SHA1 Message Date
Chukuwebuka-2003
7d3c0b2f94 fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.

Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.

Part of #116472 (request 4).
2026-09-20 18:53:10 -07:00
teknium1
0765099ff4 fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
2026-09-20 15:50:49 -07:00
beardthelion
0e33dc9ebc fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.

Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.

Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.

(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
2026-09-20 15:50:49 -07:00
kshitij
f88c6fc46e refactor(compression): finish the row-test dedupe and drop a gate that saved nothing
Two corrections to 0b9a9a0f5c.

The `if allow_split_turn else -1` gate did not skip the scan it claimed to:
`_ensure_last_user_message_in_tail` runs the identical lookup as its first
statement, so the micro-compaction pass paid the same scan and the batch path
paid it twice. Reverted to the unconditional call, which also removes the `-1`
sentinel every reader had to reason about.

The dedupe stopped at two of five spellings of the same predicate. Three more
sites inline it: the in-flight replay's "a real request follows the summary"
check, the handoff-candidate admission test (as its negation), and the merge
pre-check. All now call `_is_real_user_turn`. The one remaining inline pair is
deliberately different — it also requires `_is_real_user_message`, which rejects
metadata-flagged scaffolding this predicate cannot see.

Equivalence: all three predicates are pure, so the negated and reordered forms
are the same test; 221 tests pass across the compressor/anchor/micro-compaction
files, including the source-shape anchor-order test that the call shape here
leaves untouched.
2026-09-20 02:23:46 +05:30
kshitij
0b9a9a0f5c refactor(compression): dedupe the actionable-user row test, skip an unread scan
Follow-up to the oversized-turn exception merged in #116181.

- `_find_last_user_message_idx` and `_real_user_indices_desc` each spelled out
  the same actionable-and-not-synthetic predicate; both now call one
  `_is_real_user_turn`. No behaviour change — same two classmethods, same rows.
- The newest-user index is only read by the split exception, which rolling
  micro-compaction disables, so the scan is skipped on that pass instead of
  running and being discarded.

`_is_actionable_user_turn` / `_is_synthetic_compression_user_turn` are pure, so
the dedupe is equivalence by construction; the row set each scan returns is
unchanged.
2026-09-20 02:11:36 +05:30
fangliquanflq
af93d57ca7 fix(agent): preserve context when summary provider overloads 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
e1f46b8518 fix(compression): pin the shadowed-checkpoint prune through compress(); drop the tail-budget claim
Review of #115833: every new test called `_prune_stale_reasoning_replay`
directly, so deleting the production call in `compress()` left the file
green. Replace the redundant first test with one that drives
`ContextCompressor.compress()` (summary patched, 12-carrier transcript) and
asserts the returned transcript holds exactly the checkpoints
`prune_pre_checkpoint_items` would keep; the class is down to two tests.

The docstring claimed the shadowed copies were "charged in full by
_ALWAYS_REPLAYED_BUDGET_KEYS": the estimator already prices the sidecar
through `strip_opaque_replay_items` (562e6e4824), so the real effect is the
compacted transcript, child sessions and persisted rows. Factor the newest-
carrier rule into `drop_shadowed_checkpoints` so the durable prune can reuse
it instead of duplicating it.
2026-09-19 10:08:14 -07:00
joaomarcos
addb689a7b fix(compression): prune native-compaction checkpoints a newer carrier shadows
`_prune_stale_reasoning_replay` exempted every `type: "compaction"` item from
pruning ("They must survive on every retained message"). The wire builder
disagrees: `native_compaction.prune_pre_checkpoint_items` rebuilds each request
around the NEWEST checkpoint run and discards every earlier one, and
`codex_responses_adapter` drops checkpoints wholesale once native compaction is
no longer eligible. In both gate states a checkpoint shadowed by a newer
carrier has no reader on any wire.

It was still charged in full against the protected-tail budget
(`codex_reasoning_items` is in `_ALWAYS_REPLAYED_BUDGET_KEYS`, ~120 KB of
ciphertext each), copied into compacted and forked transcripts, and persisted
to `messages.codex_reasoning_items` for the life of the DB. Field forensics on
a 3232 MB `state.db`: 2473 MB sat in that one column, ~120 KB per assistant
row, 400 rows carrying only 118 distinct blobs.

Keep the newest carrier — the same item the wire builder would have picked —
and drop the shadowed ones. Single-checkpoint transcripts are untouched, so
every pre-existing test in the suite passes unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Lqi8r25VMJv5TgGXYzHba3
(cherry picked from commit 00b0b7b672a2821a3baa62df70f052bc09f2a562)
2026-09-19 10:08:14 -07:00
kshitij
469296c748 refactor(compression): tighten the oversized-turn exception after review
Review findings folded into the exception added by the previous commit:

- The anchored-region check now reuses `_walk_tail_budget` instead of a
  second token sum with the default thought-charge rule. The walk charges
  thinking only on the newest assistant turn unless the route replays stale
  thinking (#73624/#84371), so the second rule could fire the exception on a
  region the walk itself considers inside the ceiling.
- Two conjuncts (`last_user_idx >= head_end`, `last_user_idx < cut_idx`) are
  implied by `user_anchored_cut < cut_idx`, which only changes when the
  anchor found a real user turn inside the compressible region; the check
  that depended on them is documented where it is read.
- The latest-assistant anchor is no longer computed and discarded under the
  exception (it also logged an anchor it never applied).
- `_find_last_user_message_idx` early-exits instead of materialising every
  actionable user index on the per-attempt boundary path.
- The exception logs at debug, like the sibling anchor decisions, rather
  than at info from both `_compress_window` and `has_content_to_compress`.

Adds the missing guard test for the new ceiling check: a transcript that
fits the tail budget must still anchor the active request rather than
triggering the exception. Red when the ceiling conjunct is removed.
2026-09-19 21:12:46 +05:30
kshitij
d32f162cad fix(compression): keep the tail anchors the split exception was bypassing
The oversized-turn exception (#80449) relaxed three tail anchors at once,
which voided two guarantees that hold on main:

- `compression.min_tail_user_messages` was skipped whenever the exception
  fired, so an N-user tail could come back with no user turn at all. The
  N-user anchor now always runs: the setting is a user-facing promise and
  outranks the budget.
- The exception fired even when the oversized weight was the active turn's
  own newest tool group. The pre-anchor cut retains that group anyway, so
  the split bought no reclaim while taking the active request out of the
  tail and losing the #10896 anchor. The exception now requires real turn
  body (a tool-call group) between the opening request and the cut.

Both failures are bound by tests, each red when its conjunct is removed:
`TestTailTokenBudgetCeiling::test_message_floor_does_not_unboundedly_override_soft_ceiling`
and `TestMinTailUserMessages::test_n_guarantee_wins_over_tail_token_budget_and_floor`
(each passes on main), plus `test_n_user_tail_guarantee_outranks_the_split`
for the N-user guarantee on a genuinely oversized turn.
2026-09-19 21:12:46 +05:30
embwl0x
cc4bb888a4 fix(compression): split oversized active turns 2026-09-19 21:12:46 +05:30
kshitijk4poor
39abfdf4ec fix(context-compressor): the marked-leaf guard must match the whole tail
Second review round on the previous commit, both findings reproduced:

- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
  about: a replayed leaf of head + marker followed by new content was never
  shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
  the marker to close the leaf, so that shape shrinks to head + marker with
  true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
  adds separator whitespace, so compact args with many keys could come back
  LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
  The helper now returns the caller's string unless the whole rewrite is a net
  reduction.

Tests cover both shapes plus a many-key compact payload.
2026-09-19 21:06:12 +05:30
kshitijk4poor
a5381a7dca fix(context-compressor): match the marker by position, and leave args byte-identical
Review findings on the previous commit, all reproduced:

- The "already marked" guard was a substring test, so a leaf that merely
  contains the marker — including one a model imitated into a new call, the
  #83714 failure mode itself — was exempt from shrinking forever. The marker
  is always written at `head_chars`, so the guard now tests that position: a
  1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
  JSON came back with inserted spaces and callers read it as a change
  (rewriting replayed history and counting a pressure hit for zero reclaim).
  Nothing replaced now returns the original string: a 547-char compact blob
  is byte-identical.
2026-09-19 21:06:12 +05:30
kshitijk4poor
1e216d1330 fix(context-compressor): only shrink args when it reclaims, never re-shrink
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:

- A leaf just over `head_chars` came back LONGER: a 628-char args blob
  shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
  compaction, so a leaf's marker was rewritten from the true count
  ("2,800 of 3,000 chars omitted") to a self-referential one
  ("223 of 423"), destroying the per-instance count property the marker
  relies on and churning those bytes on each pass.

A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
2026-09-19 21:06:12 +05:30
Daniel JB Clark
262a6436fa fix(context-compressor): stop injecting an imitable truncation marker into replayed tool_calls
Root cause for #83714 (write_file/patch_tool writing literal
"...[truncated]" into files, PR #83752's guard is the safety net, not
the fix): _truncate_tool_call_args_json() in the compression pass
shrinks long string values inside a PAST assistant message's
tool_calls[].function.arguments — the exact field that represents the
model's own prior generated output, replayed back to it verbatim on
every subsequent turn. The old marker, a bare "...[truncated]" suffix,
is indistinguishable from something the model itself could have
written (it's exactly the kind of terse ellipsis abbreviation models
already produce). A model conditioned on seeing itself "get away with"
that pattern in its own history imitates it in a new tool call,
writing the literal marker instead of real content.

This is the second bug from the same root text. The first (#11762,
MiniMax 400s from unterminated JSON) was fixed by shrinking inside the
parsed structure so the JSON stays valid, but kept the same visible
marker text — fixing the syntax problem while leaving the imitation
problem untouched.

Fix: replace the marker with one deliberately NOT shaped like prose a
model would write — distinctive non-ASCII delimiters, an explicit "not
part of the original tool call" disclaimer, and a per-instance
char-count that won't match the next omission point even if copied
verbatim. The shrunk value stays a plain string (not a nested object)
so the #11762 valid-JSON/matching-shape contract is unchanged — only
the marker text changed.

Checked context_compressor.py's other "...[truncated]" call sites
(_serialize_for_summary, _compact_fallback_turn, the user-message-only
one near _ACTIVE_TASK_MAX_CHARS) — none of them write into a value
that gets replayed as the main model's own assistant/tool_calls
history, so they don't share this priming risk and were left as-is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:06:12 +05:30
kshitijk4poor
af0be164c1 refactor(compression): one trigger derivation for update_model and the switch-guard preview
`preview_threshold_tokens` restated the resolve -> floor -> compute -> cap chain that `update_model`
runs; two copies of the trigger math drift the next time a step is added — the exact bug class
#83450 fixes (the guard quoting a number the compressor will not install). `_derive_trigger` is the
single pure derivation; the auxiliary-summariser ceiling stays in `_apply_threshold_tokens_cap`
because it is per-runtime, not per-model.

The startup banner names the cap only when it set the trigger; on windows where the ratio already
sits below it "(capped at 256,000)" was noise. Comments no longer repeat the default literal.
2026-09-19 16:28:57 +05:30
kshitijk4poor
9a18c8de43 docs(compression): comments that quoted the uncapped 500K trigger follow the new default
delegation.compression_threshold_tokens and _apply_child_compression_cap both described a 1M-window
child compacting at 500K "like its parent"; with compression.threshold_tokens defaulting to 256K the
parent (and so the child, via the min of both caps) compacts at 256K. preview_threshold_tokens now
says it ignores the auxiliary-summariser ceiling, which only the post-switch probe can know.
2026-09-19 16:28:57 +05:30
fangliquanflq
b8a77fa92c fix(cli): model-switch warning quotes the capped compression trigger
The preflight-compression warning computed `context_length × threshold_percent` itself, so with the
absolute `threshold_tokens` cap (and provider-scoped `model_thresholds`, the small-window floor) it
quoted a trigger the compressor would never use — "auto-compress at ~500,000" on a 1M switch that
really compacts at 256,000. `ContextCompressor.preview_threshold_tokens()` now computes the post-switch
trigger with the same machinery `update_model` uses, without mutating state; the guard asks for it and
keeps the plain ratio only for duck-typed engines that lack the method.

Salvaged from #83523 (commits bb210c09a7 + a6918d8b1b squashed; the ContextEngine ABC default was
dropped — the guard's fallback already covers engines without a preview).
2026-09-19 16:28:57 +05:30
kshitijk4poor
31447bda5e refactor(compression): pick the tail walk's floor once instead of walking twice
The two walks differ only in the break predicate's `(n - i) >= min_tail` conjunct, so when the
ceiling can hold the floor's wire overhead the floorless walk's cut is exactly what the fix chose
on every branch (identical break row when the floor fits; the bounded cut when it does not; the
#40803 re-cut already returns a cut at or below the floor). Verified by a 40,000-transcript
randomized differential probe against the two-walk form: 0 mismatches.

Also drops the second O(n) estimate pass from the micro-compaction per-turn path.
2026-09-19 16:21:37 +05:30
KoNit-K
fdbcdef914 fix(agent): bound tail message floor by token budget 2026-09-19 16:21:37 +05:30
kshitijk4poor
80f0d1b52f fix(agent): re-prompt once when a turn that did tool work ends on a collapsed fragment
A Responses-wire collapse (#103483): the tool work runs correctly, then the
final text stop is a fragment — a stray wrong-script word ("пар"), a token
starting mid-punctuation ("?warming up") — and the loop accepted it as the
answer, so the turn reported completed and an unattended job abandoned the
task.

The guard rides the existing ack-continuation path in finalize_turn: same
scope knob (agent.intent_ack_continuation, default auto = Responses
transports), the same bounded per-turn counter, the same durable interim +
nudge rows, and the same synthetic-user recognition in the compressor. It
fires only when the turn has tool results after the last user row and the
answer matches a deliberately narrow predicate: <= 24 chars, no sentence
terminal, and either leading punctuation no answer begins with or letters
with no ASCII letter/digit at all. "42", "SQLite", "report.csv", "€12.50",
"你好。", "Done." never match. Because shape cannot prove a collapse, the
nudge asks for the same answer again when it WAS complete, so a false
positive costs one call, never the answer. The nudge row closes the
tool-work window, so a second fragment ends the turn as before.

Rebuilt from PR #111472 by dankkush against current main; the mid-task
"stall note" arm and the ephemeral-row popping are not taken (the former
false-positives on declarative answers, the latter buries flagged rows once
a recovery response calls a tool).

Co-authored-by: dankkush <brandan@wrengineers.com>
2026-09-19 12:12:04 +05:30
teknium1
dbf19a6cfa fix(compression): collapse duplicate task sections during snapshot grounding
Follow-up to the salvaged #114480 (@jonameijers) and #114553 (@JoaoMarcos44).

A small summarizer can emit both the canonical "## Historical Task Snapshot"
and the legacy "## Active Task" heading in one summary (the 4B model in the
issue "duplicates the section set"). With count=1 the grounding pass replaced
only the first match and left the second as a live-looking, undisclaimed
task section. Replace the first task section with the deterministic snapshot
and drop every later one.

Tests: fold the two contributor test files into one file with two invariants
(prompt names the emitted heading; grounding collapses alias + duplicate
sections). The original #114480 test asserted the pre-fix prepend behaviour,
which #114553 makes false.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 19:18:23 -07:00
joaomarcos
d295ef02ee fix(compression): replace alias task headings during snapshot grounding
Align the iterative update instruction with HISTORICAL_TASK_HEADING and
treat leftover ## Active Task sections as the same snapshot so grounding
replaces them instead of prepending a second live-looking task heading.

Fixes #114479
2026-09-18 19:18:23 -07:00
openclaw
4ab7dec756 fix(compaction): name the emitted task heading in the update instruction
The iterative-update instruction told the summarizer to update
"## Active Task", but the template it is given emits HISTORICAL_TASK_HEADING
("## Historical Task Snapshot"). Leftover from #44454, which renamed the
heading in the template and SUMMARY_PREFIX but not in this literal.

A summarizer that follows the literal emits a section
_HISTORICAL_TASK_SECTION_RE cannot match, so
_ground_historical_task_snapshot prepends instead of replacing and the
summary carries two task sections. Only the grounded one is disclaimed by
SUMMARY_PREFIX, leaving the stale one readable as live work - the hijack
class #44454 closed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 19:18:23 -07:00
liuhao1024
a0562d17d8 fix(auth): missing-credential hints name the real env var or the OAuth login (#114405, #78996)
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.

One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.

Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.

Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
2026-09-18 11:11:44 -07:00
teknium1
303bcd804a fix(compression): a timed-out preflight compaction sends a fitting request and prune-commits an over-window one
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.

- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
  exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
  stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
  compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
  prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".

Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
2026-09-18 10:36:11 -07:00
teknium1
9f5b7ea02e fix(compression): keep the aux ceiling and re-probe feasibility on every main-runtime change
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).

- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
  _apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
  only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
  every runtime change: switch_model (outside the rollback guard), fallback activation
  and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
  fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
  rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:49:14 -07:00
teknium1
07c92d675a fix(compression): repeated summary stall escalates to the deterministic fallback summary
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).

WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.

Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.

Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
2026-09-17 08:59:50 -07:00
kshitijk4poor
fa5da7cf1a perf(agent): size outbound image payloads by data length, not by re-serializing them
`_image_payload` ran `json.dumps` over every image part on every request (pass runs on
each API call, before the HTTP client serializes the same bytes again). Measured per
request on the send-path clone, 20 images:

  per image   json.dumps   len(data)
  256 KB         50 ms       0.02 ms
  1 MB          226 ms       0.01 ms
  5 MB         ~1.0 s        0.03 ms

origin/main's keep-3 walk was 0.4 ms. The image payload is ASCII base64, so the data-URL
(or Anthropic `source.data`) length is the byte count within ~17 bytes of JSON framing
per part — noise against a 24 MB budget with 8 MB of headroom.
2026-09-17 13:52:36 +05:30
kshitijk4poor
5dcfae44b4 refactor(agent): one image-eviction policy for both send-path passes
The compressor pass and the Anthropic wire pass each carried their own copy of the
limit/batch/floor constants (under different names) and of the retire-count loop.
Retuning one alone would make the wire pass re-evict on a second frontier after the
compressor pass — the exact prefix-rewrite class #113517 fixes.

- agent/image_eviction_policy.py: stdlib-only leaf holding the constants and
  outbound_image_retire_count(); the wire converter stays a leaf (no compaction import).
- evict_stale_outbound_tool_images loses `keep_newest`: no caller passed it, and its
  default was the COMPACTION window constant the comments say the send path must not use.
- _retire_stale_tool_result_images (compaction) back to its main body; behaviour was
  unchanged, so the rewrite was churn.
- One `_tool_result_parts` unwrap instead of three copies; one `_image_payload` walk
  per message instead of two.
- Wire pass counts blocks via `_block_type` (non-dict inner blocks no longer raise).

Tests import the constants from the policy module; the two "stops at the floor" tests
assert the invariant (newest floor survives, request under the limit) instead of the
exact frontier. test_user_uploads_are_not_evicted now builds a request that actually
trips eviction — at 6 blocks it never fired and proved nothing.
2026-09-17 13:52:36 +05:30
teknium1
09d68f8a65 fix(agent): a batch retire that would blind the model stops at the keep-newest floor
Both eviction passes retire the oldest image-bearing tool results in whole
batches (8) and let the floor shelter a breach only when reserved uploads alone
exceed the ceiling. When uploads fill most of the ceiling without exceeding it,
the first batch step overshoots past the newest frames: fifteen uploads plus six
one-frame tool results retired all six (the model blind on the frame it had just
asked for) although keeping the newest three already clears the 20-block limit.

The batch step is now cut back to `total - floor` when it would cross the floor
and keeping the floor fits; a floor that does not fit (one 25-block tool_result,
byte pressure) still retires past it, and uploads alone over the ceiling still
shelter the floor. Pure-screenshot sessions are unchanged: same batch-aligned
frontier at 21/29/37 images.

Review-fix on #113519 (found while checking @ehz0ah's mixed-session boundary on
the sibling PR #113619).
2026-09-17 13:52:36 +05:30
Meowrasaki
b150dfa87c fix(agent): make the keep-newest floor satisfiability-aware
Three residual witnesses from review on this branch and #113619.

W1 (@ehz0ah): reserved user uploads were counted against the BLOCK ceiling but
not against the byte budget, so five 5 MB uploads plus three 3 MB tool images
serialized to 34.0 MB against Anthropic's 32 MB Messages limit and eviction
returned pruned=0. The API answers 413 request_too_large.

W2/W3 (@andrexibiza, @teknium1): the floor applied unconditionally, so it
sheltered breaches the tool content itself caused. One tool_result carrying 25
image blocks left retire=0 because only one carrier existed against a floor of
three; four 10-block results retired one and kept 30 blocks.

Reserve upload bytes alongside blocks, and make the floor conditional on
satisfiability: it shelters a violation only when retiring every removable tool
image still cannot bring the request inside the ceiling, and never when bytes
are the breach, since the request-size limit is hard. Uploads are still never
rewritten, and the unsatisfiable case still keeps the newest frames.

  W1  34.0 MB -> 25.0 MB, pruned=3
  W2  25 blocks -> 0, pruned=1
  W3  40 blocks -> 0, pruned=4
  floor intact: 21 uploads + 5 screenshots keeps 3

Ref: https://platform.claude.com/docs/en/api/overview#request-size-limits
2026-09-17 13:52:36 +05:30
Meowrasaki
5768f78e4f fix(agent): count image BLOCKS and reserve user uploads against the limit
Two gaps in the limit-triggered eviction, both found by review on #113517.

The ceiling is a count of API image blocks, but the send path counted tool
MESSAGES. Eight tool results carrying three screenshots each are 24 blocks
against a 20-block ceiling; counting messages saw 8 and evicted nothing.

User uploads occupy the same per-request budget and were not counted at all, so
a mixed session sailed past the limit. They must be reserved against the ceiling
while never being rewritten -- the user attached them by hand.

Reserving uploads introduces a failure of its own: once uploads alone reach the
ceiling, a pure limit rule retires every screenshot and leaves the model blind
on the frames it was just asked about, which is worse than the stricter
dimension cap the limit exists to avoid. keep_newest therefore becomes a FLOOR
rather than a trigger.

Both stages now weight by image-part count and reserve non-tool image blocks.
Three regression tests cover the block accounting, the reserved-upload
accounting, and the floor; all three fail when either fix is reverted.
2026-09-17 13:52:36 +05:30
teknium1
c37753c424 fix(agent): round the eviction overshoot up to whole batches so the image limit actually holds
Review fix on top of the limit-triggered eviction. Both send-path stages ran
statelessly on a fresh clone every request, and each retired a FIXED single
batch (`min(batch, count - keep)`) once over the limit. That caps total
eviction at 8 per stage forever: once a session passes 20 + 8 (+ 8 from the
adapter) images, the outbound request grows unbounded again (24 image blocks
at 40 vision calls, 44 at 60 in the issue's own replay), so the documented
20-image threshold and the 24 MB byte budget were not enforced above the
first window. The PR's 40/60-row "3 collapses" result was this saturation,
not a held frontier under the limit.

Retire `ceil(overshoot / batch) * batch` instead, extending by whole batches
while the surviving payload still exceeds the byte budget. The frontier is
still a step function of the image count (one move per batch), and the
request now stays at or below the limit at every size:

  images  collapses  max outbound images   (PR head -> fixed)
      40      3 -> 4        24 -> 20
      60      3 -> 6        44 -> 20
  2 MB x 30   3 -> 4    40 MB -> 22 MB

Both frontier tests now span three batch windows and assert the outbound
count never exceeds the limit; both are red on the previous head.
2026-09-17 13:52:36 +05:30
Meowrasaki
66a281c330 fix(agent): evict images at the provider limit, not on a keep-newest count
Both send-path image evictions retired the oldest payload as soon as more than
three images were present. Retiring an image edits a message the provider has
already cached, and Anthropic matches its prompt cache on an exact byte prefix,
so each retirement re-wrote the entire conversation. Because the window was a
count, one more message retired on every new image and nearly every turn became
a full-prefix miss.

Observed on a real 104-request session: requests reported cache_read=16,076 --
system prompt plus tool schemas and nothing else -- with cache_write up to
127,720, against a steady-state read of ~145,000. Those requests produced 62
percent of the session's entire cache-write volume.

Trigger eviction on the API's own per-request image and byte limits instead, and
retire a batch when it fires. Under the limit nothing is rewritten and the
request is append-only, so the cached prefix survives; over it, the cost is one
slower turn per batch rather than one per image. This matches the documented
behaviour of the reference Anthropic client.

Simulated across session sizes, prefix collapses fall from 3/8/18/38/58 to
1/1/1/3/3 at 5/10/20/40/60 vision calls.

Compaction keeps the count-based window: it commits its rewrite into the
canonical transcript once and has no prefix to preserve.
2026-09-17 13:52:36 +05:30
joaomarcos
f59d973af6 fix(compression): stall cooldown covers one idle window; drifted-prompt INFO once per session
After a no-progress stall the ladder's first rung (60s) was shorter than the
default idle stall window (120s), so the next oversized automatic turn
re-entered the same silent summary route about a minute after burning the
whole window (#112420: "made no progress ... continuing without
compression" followed by "context compression started" 62s later).
record_timeout_failure now floors every timeout-class cooldown at the
configured idle window; a window below the rung leaves the ladder untouched.

The "Compaction rebuilt a drifted system prompt" line fired on every compact
of a long session (19x/day observed). The rebuild stays mandatory; only the
first drift per session is INFO, later ones DEBUG.

Both hunks re-applied from PR #112504 (@JoaoMarcos44); its no-LLM prune on
stall is a product decision and was not ported.

Part of #112420
2026-09-16 17:12:33 -07:00
teknium1
5670f846cf fix(compression): mark failed skill-tool results and drop the duplicated stub suffix
Follow-up to the salvaged #112719 commit:

- `_skill_result_failure_suffix()` is shared by the `skill_manage` and
  `skills_list` stubs instead of two inline copies, and also fires on a
  payload that says `success: false` without an `error` string (previously
  such a result still compressed into a success-shaped line — the exact
  class the issue reports). The error preview is whitespace-collapsed so
  the stub stays one line even when the tool error carries newlines.
- The `skills_list` comment no longer names a `query` arg the schema does
  not have (`SKILLS_LIST_SCHEMA` takes `category` only).
- Tests trimmed to two invariants in the existing summarizer module
  (`tests/agent/test_context_compressor.py`): skill_manage names its ops
  (operations array and legacy flat shape) and keeps FAILED visible with
  the error text; skills_list renders category/count and marks failure,
  with skill_view as the unchanged control. The standalone 9-test file
  from the contributor PR is folded into these.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 16:59:43 -07:00
Konstantin Khlopkov
031654650f fix(compression): skill-tool summaries carry skill names and failure outcome
_sum_named read a top-level `name` arg that only skill_view has: skill_manage
names live in operations[] (or the legacy flat shape) and skills_list has none,
so every compressed line rendered `name=?`. The stub also ignored the result
payload, collapsing failed batches into success-shaped lines and dropping the
error text. Dedicated summarizers keep the op list and surface FAILED plus the
leading error text; _sum_named is gone.

Fixes #112710
2026-09-16 16:59:43 -07:00
teknium1
3f59b5c594 refactor(agent): /context breakdown and native-compaction retention use the canonical token estimator
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
2026-09-13 05:09:43 -07:00
Teknium
fc3d60af09 feat(compression): provider-scoped model_thresholds keys ("<provider>:<substr>")
A bare `astra: 0.85` in compression.model_thresholds was written for the Codex
OAuth route, where Astra is capped at 272K and 50% would compact at ~136K. The
key is substring-matched on the model name alone, so it also fired on
openai/gpt-6-astra via OpenRouter and Nous, where the window is 1.1M: the user's
0.5 global threshold was silently replaced by 0.85 and the session sat at 620K
(~59%) without compacting.

Keys may now carry a provider prefix: `"openai-codex:astra": 0.85` applies only
when the session's provider is openai-codex; bare keys keep their route-agnostic
behaviour. Ranking is by model-substring length with scope as the tie-break, so
`astra-900k` still outranks `openai-codex:astra` for the 900K picker. The
provider flows through ContextCompressor (ctor + update_model), the ContextEngine
base class and the TUI hot-reload path, so a /model switch between routes
re-scopes the override.
2026-09-11 02:07:48 -07:00
gaoanze888
8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
kshitijk4poor
26f4a674e0 fix(agent): a /steer row is human input for every user-turn predicate
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.

Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
2026-09-09 13:08:25 +05:30
Teknium
e9313f6458 fix: let provider evidence adjudicate past-window preflight estimates 2026-09-07 14:11:41 -07:00
kshitijk4poor
ed063034ad test(compressor): pin custom_providers threading at the get_model_context_length seam
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
2026-09-08 02:27:01 +05:30
dhruv kejriwal
71516214c3 fix(compressor): thread custom_providers into context-length resolution
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
2026-09-08 02:27:01 +05:30
Teknium
77086b1439 refactor: isolate compression summary dispatch 2026-09-07 05:59:08 -07:00
Teknium
be58c276ee feat(compression): per-image token cost learned from the provider's own usage (#70328, supersedes #70463)
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).

The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.

- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
  anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
  ~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
  read the same bound value, so trigger and walk agree; the per-message memo now caches text
  tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.

evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.

Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
2026-09-06 14:19:42 -07:00
Joey
820106d4a5 fix(compaction): count api_content in tail budget 2026-09-06 14:04:05 -07:00
Teknium
562e6e4824 fix(compression): opaque encrypted_content costs 0 in every local token estimate
Codex Responses reasoning and compaction items carry ciphertext the provider prices by its own
token count, never by bytes; a single native compaction checkpoint is ~5M chars, which the
bytes/4 estimator turned into ~1.29M "tokens" against a 204K threshold (#100611). #104192 deferred
that decision for one request; this removes the mis-pricing at the source so the preflight
estimator and the tail-budget walk agree (a mismatched size class protects blob-heavy rows as
"small" and compaction re-fires). Only real usage prices these items, and the usage anchor carries
that price forward.

evals/native_compaction/ab_checkpoint_preflight.py: preflight estimate after checkpoint
1,292,413 -> 58 rough tokens; the over-threshold negative arm still compresses.
2026-09-06 13:21:17 -07:00