Commit Graph

500 Commits

Author SHA1 Message Date
kshitijk4poor
fb86bc708d fix(context): redact full tool-call args before the summarizer cut
62ceddd342 cut raw args to HEAD+4096 before redaction to save time. The
PEM redaction pattern only matches a complete BEGIN...END block. A long key
whose END fell past the cut stayed unredacted, and once an earlier key was
redacted and the text shrank, its body landed in the 1200-char head that
goes into the persisted summary. Go back to the BASE order: redact the full
args, then apply the MAX/HEAD cut. This is a cold path (once per summarized
call per compaction), and _SUMMARY_INPUT_MAX_CHARS still bounds the prompt.
Extend the kept canonical-args test with a two-PEM input that leaks on
62ceddd342 and passes now.
2026-09-27 18:38:58 +05:30
kshitijk4poor
be834681dc fix(context): measure pruned regions, bound arg redaction, drop dead counters
Follow-ups to making tool-call args byte-exact:
- _record_compression_regions measured canonical_messages slices while
  compress_start/compress_end are indices into the pruned copy that head/tail
  are assembled from; measure the pruned rows actually sent, as before. This
  also removes the only canonical slicing, so blank-echo classification drift
  between the two copies can no longer misalign anything.
- _render_tool_call_for_summary redacted the full (now unbounded) args before
  cutting to 1200 chars; cut to head+4096 first. Output unchanged for args
  within that window.
- pressure_hits always equalled demoted once arg truncation left; fold it.
- Drop the fixture-only tautological assert in the guardrail test helper.
- Reword stale compress()/compression_marker docstrings that still described
  canonical head/tail and compressor-written arg markers.
2026-09-27 18:38:58 +05:30
kshitijk4poor
66240ddc59 fix(context): restore marker import in tests, drop unused compressor imports
The arg-truncation removal deleted the only uses of the marker constants in
agent/context_compressor.py (ruff F401), and the test module had been
importing _COMPRESSION_MARKER_PREFIX through it, so test_context_compressor
line 137 raised NameError. Import it from its home, agent.compression_marker.
2026-09-27 18:38:58 +05:30
kshitijk4poor
f6ce8bb23b fix(context): assemble compaction head/tail from the pruned copy (#61932)
The salvaged lossless-history change rebuilt the carried head/tail from
canonical history, which undid _pressure_demote_tail's tool-result
shrinking and re-broke #61932 (an all-oversized tail could no longer
compress). Pruning no longer rewrites tool_calls, so the pruned copy's
arguments are already byte-identical to canonical history: assemble the
head and tail from the pruned copy, keeping tool-result demotions and
exact tool-call arguments at once. Docs updated to match.
2026-09-27 18:38:58 +05:30
JoaoMarcos44
a7baa5f5eb fix(context): keep compaction history lossless
(cherry picked from commit d51c8f4f5096badfd0beddd78646617643f6028f)
2026-09-27 18:38:58 +05:30
kshitijk4poor
ad4e4496c2 fix(compression): drop duplicated snapshot sentence from task prompt
The historical_task instructions already explain that the compressor inserts
a bounded, redacted snapshot after generation; repeating it in the reverse
signal paragraph only adds prompt tokens.
2026-09-27 18:27:37 +05:30
Charan Rathore
f7be32556d fix(compression): avoid long verbatim task quotes in summaries
(cherry picked from commit 006e1d286126835124c3d837c5bd094291c89716)
2026-09-27 18:27:37 +05:30
Adolanium
e46d4c0ade fix(compression): let the next summary build on a fallback handoff
A deterministic fallback summary replaced the older handoff in the transcript but never updated _previous_summary. The next compaction kept the stale in-memory summary and dropped the fallback row from its window, so the fallback's user asks, files and last dropped turns never reached the summarizer. Store the fallback body in _previous_summary the same way a normal summary is stored.

(cherry picked from commit 35417d2e1ffbb775c3eaff17b26623896afa56c1)
2026-09-27 18:20:40 +05:30
kshitijk4poor
d8be403097 fix(compression): match the Codex stall marker on the raw error text
The stall check relied on the lowercased error string, which only works
while CODEX_STREAM_STALL_MARKER happens to be all-lowercase. Match against
str(e) so the shared marker stays authoritative regardless of case, and
fold the two duplicate #124077 comments into one explaining the split
(stall -> retry-ladder timeout; transport timeouts stay terminal).
2026-09-27 18:19:07 +05:30
kshitijk4poor
8c9f2162c7 refactor(compressor): key the Codex stall on the stream guard's shared marker
Co-authored-by: happy5318 <5318happy@users.noreply.github.com>
(cherry picked from commit 83c0a8bb9c7d0e611bafe436d2e4da77144c26ec)
2026-09-27 18:19:07 +05:30
kshitijk4poor
3a1a45dbf1 fix(compression): keep transport timeouts terminal; reclassify only the Codex stall
The cherry-picked fix set streaming_closed=False for every timeout, which
also stripped the terminal abort-and-preserve-session behaviour from real
network timeouts (openai APITimeoutError, httpx Read/ConnectTimeout), the
deliberate #29559/#25585/#94448 design. Narrow it: only a TimeoutError whose
message says "stalled" (the Codex aux stream guard) becomes a timeout that
takes the 60/300/900s ladder and does not arm _last_summary_network_failure.
Other timeouts classify exactly as on main.

Co-authored-by: happy5318 <5318happy@users.noreply.github.com>
(cherry picked from commit e013c38a017b4709d4598a0a07c71ea26519312c)
2026-09-27 18:19:07 +05:30
happy5318
56ee7d6f09 fix(compression): classify a Codex stream-guard stall as a timeout, not a terminal network failure (#124077)
## Thinking Path
When the Codex auxiliary stream guard aborts a compaction summary mid-stream
it raises `TimeoutError("Codex auxiliary Responses stream stalled: no new
output for 60.0s ...")`. The message contains neither "timeout" nor
"timed out", so `_classify_summary_failure` returned `timeout=False` while
`_is_connection_error` (which matches the type name "Timeout") returned
`streaming_closed=True`. The terminal network-failure flag then armed an
unconditional abort (`_TERMINAL_SUMMARY_FAILURES`), bypassing the retry
ladder and the deterministic fallback summary — on turn-start preflight
compression that ends in "Auto-resetting session after compression
exhaustion", wiping the session.

### What Changed
`agent/context_compressor.py` `_classify_summary_failure`:
- `timeout` is now computed first, and additionally matches `isinstance(e,
  TimeoutError)` and the "stalled" message shape (the actual text the Codex
  guard emits).
- `streaming_closed` is `_is_connection_error(e) and not timeout` — a
  timeout keeps its retry-ladder semantics and can never arm the terminal
  network-failure abort.

### Tests
New `TestSummaryFailureClassification124077` in
`tests/agent/test_context_compressor.py`:
- classify: Codex stall → `timeout=True, streaming_closed=False`.
- classify: plain `ConnectionError` stays `streaming_closed=True` (no
  regression on the premature-close class).
- classify: a "timed out" message on a non-TimeoutError type stays a
  timeout and is excluded from `streaming_closed`.
- integration: the stalled-summary path in `_generate_summary` does NOT arm
  `_last_summary_network_failure`.

### Verification
- RED/GREEN double proof via git stash: pre-fix 3 failed, post-fix 4/4 pass.
- Regression: 15 existing failure-classification tests pass
  (network_failure / premature_stream / empty_content / auth / truncation).
- ruff clean on changed files.

### Notes
Local test env: this checkout's venv python is a symlink into
`<home>/hermes-agent/.hermes-runtime/...`, so stdlib `zoneinfo` first-import
and pydantic's plugin `distributions()` scan under the real-home IO guard
needed the collection-time warmups at the top of the test file. CI
interpreters are not symlinked into the home — those two warm-up blocks are
no-ops there.

## Related
#124077 (issue). Family: #124078 (the same stall's template trigger),
#108104 (`auxiliary.compression.no_progress_timeout`).

(cherry picked from commit c5399fbdee44f2db1172f22632cbf348a2c7cab5)
(cherry picked from commit bce26a99791c09911d8a008a1922c512f5c1fcea)
2026-09-27 18:19:07 +05:30
kshitijk4poor
a68e29a2c7 refactor(compression): drop the redundant dict guard in the replay check 2026-09-27 00:48:22 +05:30
kshitijk4poor
6b8f2b4e36 refactor(compression): drop the in-memory merged-replay flag
_INFLIGHT_REPLAY_MERGED_KEY was set at exactly one site, on a summary
carrier that has just had the replay header and a non-empty task appended
after its (last) end marker, so the content predicate
_has_merged_inflight_replay is always true wherever the flag is. Keeping
both left two sources of truth for one fact, and the flag is the one lost
on SessionDB reload -- the bug this stack fixes. Detect the merge from
content only so the in-memory and reloaded paths share one code path.

The cron reappend test now asserts the merge layout through the predicate
instead of the private dict key.

Co-authored-by: ppazosp <pablopazosp3@gmail.com>
2026-09-27 00:48:22 +05:30
kshitijk4poor
3d7629dbe5 fix(compression): anchor merged-replay detection on the last summary marker
_has_merged_inflight_replay partitioned on the FIRST end marker. A
merged-into-tail carrier can embed an older carrier (marker + replay
header + task) in its prior-context block, ahead of the new summary's own
marker, so a carrier with nothing after its real boundary was misread as
already carrying the active request. _find_inflight_user_task would stop
there and the reappend would drag the new summary body into the task
text. rpartition anchors on the real boundary; summary bodies never keep
an inner marker and the _force_user_leading layout has exactly one, so
every legitimate layout still matches.

Also compute the stripped remainder once and use removeprefix instead of
a manual length slice.

Co-authored-by: ppazosp <pablopazosp3@gmail.com>
2026-09-27 00:48:22 +05:30
ppazosp
96f087eef5 fix(compression): retain non-text active requests when splitting turns
(cherry picked from commit 0cbbee6bf66aa091312fde3df6dbfd23ff75c3c1)
2026-09-27 00:48:22 +05:30
ppazosp
e81be5b66a fix(compression): preserve the active request after SQLite reload
(cherry picked from commit 6ee1d087caf3f158d7379c8e4ca09a67c9c0e684)
2026-09-27 00:48:22 +05:30
kshitijk4poor
40cb28e60c fix(agent,slack): close remaining compressor elision gaps from review
- salvage summary cap and _bound_oversized_record still composed bare
  truncation idioms in model-visible text; route them through elide /
  elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
  apostrophe when the user text held both quote kinds and hiding it from
  the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
  over-budget marker line in _build_verbatim_user_section and the Slack
  nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
  one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
  stop's repeated 1200.

Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
2026-09-26 23:49:25 +05:30
kshitijk4poor
21edc762d5 fix(agent): route summary-snapshot and memory-context truncations through elide()
#121548's sweep missed three model-visible elisions whose wording differs
from the bare idiom the tokenizer test looked for:

- context_compressor fallback "...[previous summary snapshot truncated]"
- context_compressor fallback "...[fallback summary truncated]"
- context_engine memory-provider "...[memory provider context truncated]..."

They are the same imitable shape, so a model copying them into a durable
write would bypass the dispatch-boundary guard. Route them through
elide()/elide_middle() (caps still held; the helpers size against the
marker) and widen the no-idiom invariant to any "...[<words> truncated]"
variant so a reworded marker cannot slip back in.
2026-09-26 23:49:25 +05:30
ahisblessed
77b9ea0831 fix(agent): route every model-visible elision through the non-imitable compression marker (#121548)
The #83714 fix hardened only the tool-call args renderer; five other
renderers (plus three same-idiom siblings) still composed the bare
truncation marker, which the model imitated from replayed context into
new durable writes (#83435).

This routes all model-visible elision in agent/ through shared helpers
in agent/compression_marker.py:

- elide(text, limit): head + marker cap with accurate omitted/total counts
- elide_middle(text, head, tail): kept head/tail with the middle elided

The elision marker's first sentence is byte-identical to the args marker's,
so the existing _COMPRESSION_MARKER_RE (and the dispatch-boundary guard in
tool_dispatch_helpers) rejects a copied marker regardless of which renderer
leaked it; only the tool-call-specific second sentence is dropped so the
marker still fits small caps (the clarify summary cap is 199 chars).

Routed sites:
- context_compressor.py: static-fallback turn renderer, _sum_clarify(),
  summarizer prompt builder (middle elision), active-task snapshot,
  lean user-message quotes (2 sites)
- skill_preprocessing.py: inline-shell output embedded into skill bodies
- lsp/reporter.py: truncate() for <diagnostics> tool output
- verification_stop.py: verify-on-stop nudge output summary

Regression test asserts the imitable idiom appears nowhere in agent/
strings (comments excluded), so a new open-coded renderer cannot land
silently.

Fixes #121548

(cherry picked from commit 69c63d50b4ababedf4fb4ff88c08f5d1e144a79f)
2026-09-26 23:49:25 +05:30
kshitijk4poor
3aa8a7f393 docs(compression): say the newest-row tail floor applies to every row
The cut clamp n - max(walk_floor, 1) runs for every caller of
_find_tail_cut_by_tokens (batch compress and micro-compaction) and
keeps the newest row whatever its role. The old comment described it
only for tool rounds, so a reader could take it for tool-round-specific
logic.
2026-09-26 23:45:36 +05:30
kshitijk4poor
3231237448 fix(compression): keep the pending round's tool-call args verbatim too
The spare covered only the pending round's tool results. The owning
assistant(tool_calls) row sits at spared.start - 1, below the lowered
prune boundary, so pass 3 (range(prune_boundary)) and pass 4's
_shrink_at still truncated its arguments: a pending 60 KB write_file
call came back as 456 chars next to its intact, unread result. The
model then reads its result beside a mangled copy of the call it just
made.

_spared_pending_tool_round now includes the owning assistant row, so
the boundary and _shrink_at skip it along with the results. That row
also counts toward the hard-share check, so the >20% escape hatch
(#61932) still applies to the whole round.

The existing pending-round test now gives the pending call >500-char
args and asserts the assistant's tool_calls survive unchanged (red
before, green after).

Co-authored-by: MongLong0214 <97578200+MongLong0214@users.noreply.github.com>
2026-09-26 23:45:36 +05:30
MongLong0214
bb17b1f74c fix(compression): keep the pending round's images when it fits
Compaction spared a pending tool round's text results but not its
images. The image-retirement pass kept only the newest three image
results across the transcript, so four parallel vision calls in one
unread round lost the oldest image before the model saw any of them.
On a single-prompt run the final media pass lost three of the four:
compaction re-appends the task as a user row after the round, and the
pending round was looked up after that row was added, so nothing was
spared.

Both image passes now skip a pending round that fits the hard share,
and the pending round is found before the task row is re-appended.
Completed rounds and a pending round over the hard share keep the
existing policy.

(cherry picked from commit eb7661f4365f009d5f9ef85e26f2f4b6ac83632e)
2026-09-26 23:45:36 +05:30
MongLong0214
65f185c594 fix(compression): keep the pending tool round when two steers follow it
The pending-round check skipped only one trailing /steer row. Two can
land in the same iteration: one is appended when the tool batch ends,
and a steer sent after that is drained before the next request and
inserted right after the newest tool result. The transcript then ends
tool -> steer -> steer, so the check stopped on the second steer row,
found no pending round, and preflight compaction replaced the unread
tool output with the one-line pressure stub.

The check now skips every contiguous trailing steer row, identified the
same way as before, and still stops on any other user row. The existing
regression gains a case with two steer rows after the round.

(cherry picked from commit 455a1a3868b3843d15401781363ce796bae96b5b)
2026-09-26 23:45:36 +05:30
MongLong0214
18eb09f5db fix(compression): keep the unanswered tool round verbatim through mid-turn compaction
Preflight compaction can fire right after a tool round, before the model
has read its results. The protected-tail passes then treated that round
like any old output: pass 2 and the pressure pass (#61932) replaced its
results with one-line stubs, and when the newest row alone exceeded the
tail ceiling the cut landed at the end of the transcript and summarised
the round with its whole turn. The model then re-ran the call, side
effects included, or answered without the output.

_pending_tool_round finds the tool results the transcript ends with,
skipping one trailing /steer row (a steer is delivered after the newest
result, before the next API call). Both passes spare that round, and the
tail cut aligns from the row before the end so the group stays whole.
The one exception is a round that alone exceeds the tail's hard share
(20%) of the input budget, the window minus the output reservation: it
still gives way, as #61932 requires. _effective_input_window is
extracted from _compute_threshold_tokens so both use the same budget;
thresholds are unchanged.

(cherry picked from commit 9d74e22379cd7dc39636c522175772d0165db4ac)
2026-09-26 23:45:36 +05:30
kshitijk4poor
74dc4ac1e4 fix(agent): a 403/402 saying "overloaded" is not an overload strike
An auth/quota error whose text also says "overloaded" went through the
overload branch: it bumped the durable sustained-overload streak even though
it always aborts as summary_auth_failure (#29559). After the credential was
fixed, the next real 503 then inherited that streak and committed the lossy
fallback immediately, skipping the one-abort grace (#115906).

Compute the access/quota classification once and keep such errors out of the
overload branch. With that, the overload-escalation path never runs for an
auth error, so the keep_auth escape hatch (and its hard-coded flag-name
compare) in _clear_terminal_summary_failures is dead and goes away.

The existing [403_overloaded] case now also asserts the streak stays 0.

Co-authored-by: Yuan Li <dskwelmcy@163.com>
2026-09-26 22:42:38 +05:30
kshitijk4poor
d6f09e0eb7 refactor(agent): one helper to clear terminal summary-failure flags
The two-line clear-all loop over _TERMINAL_SUMMARY_FAILURES had three
copies (init, summary success, overload escalation), and the escalation
copy now needs to spare the auth/quota flag. Extract
_clear_terminal_summary_failures(keep_auth=...) so the three sites share
one definition and the #29559 exception lives in one place.
2026-09-26 22:42:38 +05:30
kshitijk4poor
4f134d4c37 fix(agent): always write the row when resetting the overload budget
_reset_consecutive_overload_aborts skipped the DB write unless the
in-memory count was armed (or settle=True). The in-memory count is not
authoritative when two agents share a session: agent A can bump the row
to 2 while agent B still holds 0, so B's summary success or runtime
switch left A's streak alive and the next overload escalated early.
Drop the settle parameter and always persist; one helper, one contract.
2026-09-26 22:42:38 +05:30
kshitijk4poor
507e856745 fix(agent): keep the auth/quota abort when an overload escalates
The escalation clear-all of terminal summary flags also cleared the
auth/quota flag the SAME error had just set. A 403/402 or quota-worded
error that also says "overloaded"/"at capacity" classifies as both, so on
strike 3 it committed the lossy fallback instead of aborting - breaking
the #29559 invariant that access/quota failures always preserve the
session. Keep the auth flag when the current error is access/quota; a
stale auth flag from an earlier failure is still cleared.

The kept escalation test is parametrized with a 403 "provider overloaded"
case that must abort all three times with summary_auth_failure.
2026-09-26 22:42:38 +05:30
kshitijk4poor
7ac231fcdb docs(agent): keep the overload-escalation rationale on the constant only
The why (retry can win vs. deferred compression_exhausted wipe, and why the
streak is durable) was spelled out in full at the constant, in
_on_summary_failure, in record_completed_compaction and in the durable
getter docstring. Keep it once on _CONSECUTIVE_OVERLOAD_ABORT_ESCALATION
and cut the other copies to one line, matching the sibling fallback-streak
getter. Comment-only.
2026-09-26 22:42:38 +05:30
kshitijk4poor
86bf8886c2 refactor(agent): one helper for the three sustained-overload budget resets
The zero-and-persist pair was pasted at the summary-success, runtime-switch
and completed-boundary sites, and every healthy compaction wrote
compression_overload_streak=0 twice even when it was already 0.
_reset_consecutive_overload_aborts() skips the UPDATE when the in-memory
streak is 0; record_completed_compaction passes settle=True and still
writes unconditionally because another agent on the same session may have
bumped the durable row since this object last read it.
2026-09-26 22:42:38 +05:30
kshitijk4poor
818faaf16d fix(agent): label the escalated overload commit summary_overload_degraded
On a compressor with a distinct summary_model, the first overload takes the
aux->main one-shot retry, which pre-sets telemetry failure_class to
aux_model_fallback; the `telemetry.get(...) or` chain then hid the
escalation and operators never saw the documented
summary_overload_degraded class. The degraded flag now wins. It is always
initialised (__init__ and _begin_compress_attempt, the only path into
_fallback_summary_for_window), so the defensive getattr goes too; no test
builds a compressor via __new__ and reaches this method without compress().
The kept fresh-bind test now asserts the label instead of excusing it.
2026-09-26 22:42:38 +05:30
kshitijk4poor
5021e3633b fix(agent): let an escalated overload clear stale terminal summary flags
Only a successful summary clears the network/empty/truncated/auth failure
flags, and _abort_on_summary_failure aborts on ANY of them. One earlier
streaming_closed failure on a long-lived (gateway-cached) compressor
therefore kept every later sustained-overload attempt aborting past the
3-strike escalation, so the #123167 compression_exhausted wipe persisted.
When the escalation fires, the latest failure class decides: clear the
other terminal flags so compress() commits the degraded fallback.
2026-09-26 22:42:38 +05:30
kshitijk4poor
50066dda6e fix(state): bump the summary-overload streak atomically
The durable overload budget was a read-modify-write (memory += 1, then
set the column), so two agents compressing the same session concurrently
could each write N+1 and lose a strike, delaying the fallback escalation.
Bump the column with one UPDATE ... RETURNING inside _execute_write and
take the returned row value as authoritative; memory-only counting is
kept for unbound compressors. No try/except or legacy-setter shim: the
existing _durable_read plumbing already degrades on unsupported DBs.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 22:42:38 +05:30
Yuan Li
2305385042 fix(agent): persist the summary-overload budget in session state so fresh bindings inherit it
Review P1 on #123186 (ehz0ah): the N=3 sustained-overload budget was
object-local. The gateway binds a fresh compressor to the same session on
every turn / cache eviction, and restart / API-server requests construct
one too, so each fresh instance restarted the budget at zero and a
sustained summary-provider outage walked the session back into
compression_exhausted and auto-reset — the exact wipe the PR exists to
prevent. Reproduced by the reviewer with three fresh compressors bound to
one session: every attempt ended at counter=1.

Persist the streak as sessions.compression_overload_streak through the
same durable channel the fallback streak and recovery deadline already
use (#100185): SessionCompressionMixin get/set pair over the sessions
row, loaded in the compressor's durable-load block, written on every
change (overload abort, successful summary, committed boundary, runtime
model switch).

- Compression rotation carries the streak to the child row at the
  boundary, same as the fallback streak: the parent's value is read
  before the bind, re-applied, and persisted onto the fresh child row.
- A completed compaction boundary settles the budget to 0 — including
  the committed degraded fallback, so a recovered provider regains a
  full budget instead of the session staying degraded forever.
- Schema event appended to SCHEMA_HISTORY["sessions"] (seq 28, after
  compression_recovery_deadline) so salvage replay maps the new column.
- Fresh-instance regression: two aborts on one compressor, then a fresh
  compressor bound to the same session inherits streak=2 and the third
  session-wide attempt commits the fallback; plus a rotation carry-over
  test. Both fail red against the memory-only implementation.

(cherry picked from commit 3742fee8d03d7077d23b0dac3a3c79fe48e102ac)
2026-09-26 22:42:38 +05:30
Yuan Li
a5e8ad9f02 fix(agent): bound sustained summary-overload aborts so they cannot guarantee a session wipe
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.

After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.

Fixes #123167

(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
2026-09-26 22:42:38 +05:30
brooklyn!
1674499d00 fix(compression): archive the rows the compressor held, not every row up to the newest
A gap below the newest held id, and turns appended above an unpersisted current
turn, were summarized away without being read. Name those held ids and clone
the rest.
2026-09-24 11:57:56 -05:00
kshitijk4poor
b7f91258b9 fix(agent): tools-available note on dropped-tools continuation; keep legacy stub recognized (#74990)
The dropped-tools partial-stream continuation now also states the cut was a
transport interruption and tools remain available. The pre-rewording
network-stub text stays in the compressor's synthetic-turn set so
crash-persisted nudges from older sessions are not mistaken for user turns.
2026-09-24 22:27:18 +05:30
kshitijk4poor
2492193c44 fix(compression): a failed watermark read takes prune's logged no-op path
_archive_watermark_for sat outside prune's try, so a raising
get_active_message_watermark/get_message_role escaped prune_tool_results_only and
surfaced only at debug level in turn_preflight, skipping the "keeping the original
transcript" warning. Read it inside the existing try, with StaleHeldHistory caught
ahead of the generic handler. Also add the missing E302 blank line.
2026-09-24 18:01:21 +05:30
kshitijk4poor
40051c2a13 fix(compression): a stale prune/micro-compaction generation aborts instead of publishing beside the winner
Prune and micro-compaction hold no compression lease. When the newest exact row of the history they
hold is no longer active, another compaction (a /compress on this or another surface, or an earlier
pass) has already committed. `held_archive_watermark` then falls back to the start watermark, which is
right for the in-place commit (its lease rules out overlap) but for a lease-less caller archives the
winner's rows and clones them back as a "concurrent tail": two summary generations live in one session.

`_archive_watermark_for` now asks `held_archive_watermark` to raise `StaleHeldHistory` on that branch.
The prune commit returns its input unchanged (`prune:stale_generation`), micro-compaction checks before
the slow summary call and skips the pass (telemetry `stale_generation`), and its commit re-checks so a
compaction that lands during the summary call is not overwritten either. The in-place commit keeps its
fallback.

Raised on #120634 by @ehz0ah (agent/conversation_compression.py:3623 thread) and independently by
@JoaoMarcos44 in #120821.

Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-24 18:01:21 +05:30
John Paul Soliva
3f7e0fd072 fix(compression): proactive prune and micro-compaction keep turns another surface appended
Both commits call archive_and_compact with no watermark, which archives
every active row. A turn another surface appended to the same session
after this process loaded it (a Desktop session continued from Telegram)
was archived with the rest as compacted (active=0, compacted=1): marked
summarized away though no summary holds it. The display, REST with
include_compacted and session_search still show it, but it is gone from
the model's history, so the agent forgets a turn the user can still see.
Micro-compaction also runs a slow aux summary call before its commit, so
a turn that arrived during it went the same way.

Both now cap the archive at the newest row the process held, the rule the
in-place compaction commit applies, so those rows take the concurrent-
append path and are cloned after the new set. Micro-compaction captures
the held history and the store's watermark before the summary call and
caps at commit. The cap moves into held_archive_watermark(session_db,
session_id, ...), which the compressor paths can call; _held_watermark
stays as the in-place commit's wrapper.

A store without get_active_message_watermark or get_message_role keeps
today's archive-everything commit, as a store without archive_and_compact
already skips the prune. Micro-compaction's DB sync runs under a broad
except, so an unguarded call there would have failed silently.

Held lists without row ids (the gateway's replay dicts) keep today's
behaviour, as in the in-place commit. Both features are opt-in
(compression.proactive_prune_tokens > 0, compression.micro_compact).

(cherry picked from commit 8d67bee760db0fc65d4330df15198ddfd5f24480)
2026-09-24 18:01:21 +05:30
kshitijk4poor
159f716313 refactor(agent): derive the pruned-marker matcher from the compressor template
The dispatch-boundary detector re-typed the marker's wording next to the
producer's template and hid the import behind functools.lru_cache, citing a
context_compressor -> prompt_builder -> tool_dispatch_helpers cycle that does
not exist (prompt_builder only mentions this module in a comment; both import
orders succeed). Two copies of one string means a template edit silently
disables the guard.

Move the prefix/template into a dependency-free leaf, agent/compression_marker,
and build the regex from the template's first sentence with the count
placeholders swapped for \d[\d,]*. context_compressor re-exports the names, so
its callers and tests are unchanged. The matcher keeps the intended behaviour:
prefix-only mentions do not match; a minted marker, flat or nested, does. The
old \s+ tolerance for multiple spaces is dropped because the producer never
emits one, so only a template-shaped copy counts.

tool_dispatch_helpers stays light: importing it alone still does not load
context_compressor (auxiliary_client, context_engine, ...).
2026-09-24 16:29:19 +05:30
teknium1
ab2a206e5e fix(micro-compaction): show each input once after a superseded marker
Superseding a stale micro marker joins the now-adjacent user turns into one
model-facing row, while the originals stay in display history as compacted
rows, so resumed display history painted every merged input twice (and one
more time per later pass). Flag the join display_metadata.model_only and skip
it in every display projection: resume dedupe, the indexed and legacy
get_messages pages, and the prompt timeline. The model payload is unchanged.
2026-09-23 10:44:29 -07:00
kshitijk4poor
94e8a255df refactor(compression): one effective-cap helper; drop dead _config_threshold_percent fallback
`min(threshold_tokens_cap, context_length)` behind the same `cap is not None and cap > 0`
guard was written in both `_derive_trigger` and `_apply_threshold_tokens_cap`, contradicting
`_derive_trigger`'s "one place the trigger math lives". Both now call
`_effective_threshold_cap(context_length)`. The `getattr(self, "_config_threshold_percent",
...)` fallback was dead: `__init__` sets the attribute unconditionally and no bare-`__new__`
test instance reaches `_derive_trigger`. Behaviour-preserving; the cap tests still fail when
the helper is mutated to return None.
2026-09-23 00:10:27 +05:30
kshitijk4poor
e4f76b8277 refactor(agent): inline the single-use _finish closure in _sample_summary_records
`_finish` (agent/context_compressor.py:3586-3588) only wrapped
`_coverage` and re-derived `shown` from the final `selected`; it was
called once, at the return. With `_merged` hoisted above the slice loop
the classmethod now has two closures (`_merged`, `_render`) plus
`_coverage` instead of four. Same expressions, same evaluation order
(`_render` first, then the coverage counters), so the output and
coverage dict are byte-identical — verified on the 47-shape capture and
probe_C2 (0 violations, base identity x4 True).
2026-09-22 13:40:13 +05:30
kshitijk4poor
8e12cd92d7 refactor(agent): integer head/tail split in _bound_oversized_record
`_bound_oversized_record` (agent/context_compressor.py:3485-3488) split
`remaining` with float arithmetic (`int(remaining * 0.5)`) and guarded
the tail slice with `if tail_len else ""`. `remaining = limit -
marker_reserve` is >= 1 because the line above already returned when
`limit <= marker_reserve`, so `head_len = remaining // 2 <= remaining - 1`
and `tail_len = remaining - head_len >= 1` always: the guard can never
fire. Use `remaining // 2` and drop the dead guard.

Byte-identical: `int(r * 0.5) == r // 2` holds for every non-negative
int (checked exhaustively for r in 0..10**6), and the guard was dead.
Capture of 800 bounded-record cases (4 record shapes x remaining 1..200)
plus the 47-shape sampling capture is identical before/after.
2026-09-22 13:40:13 +05:30
kshitijk4poor
e17d49174d refactor(agent): drop unreachable trailing-gap branch from lean-sampling _render
The `if cursor < len(records)` block at the end of `_render`
(agent/context_compressor.py:3580-3583) duplicated the gap-marker
arithmetic of the in-loop branch but had no path to it.

Invariant: `_render` is only reached with non-empty `records`
(empty returns early), and n >= 1, so the slice-build loop always runs
its last iteration, which anchors `end = len(records)` and
`start = end - 1 >= 0`, hence `end > start` and the final slice ends at
`len(records)`. `_merged` keeps `max(end)` for the last interval, and the
extension pass only grows the newest slice backward (`(s - 1, e)`) or
older slices forward but never past the next slice's start, so the last
slice's end stays `len(records)`. Therefore `cursor == len(records)`
after the loop for every input.

Byte-identical across the 47-shape capture (40 probe_C2 shapes + 7 edge
shapes incl. single-record and empty); probe_C2 -> 0 violations;
under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
56b63ecf82 refactor(agent): merge lean-sampling slices once via _merged instead of inline
The slice-build loop in _sample_summary_records hand-rolled the same
"overlapping/touching slice -> extend previous" fold that the `_merged`
closure 25 lines below implements (agent/context_compressor.py:3563-3568
vs :3590-3597). Hoist `_merged` above the loop, append raw (start, end)
pairs, and merge once before the extension pass; the first `_render`
no longer wraps `selected` in a redundant `_merged` since it is already
merged. One interval-merge implementation instead of two.

Output is byte-identical: the merge is applied once to the same ordered
slice list, so the extension pass starts from the same `selected`.
Verified with a 47-shape capture (40 probe_C2 shapes + 7 edge shapes)
diffed before/after: identical sampled text and coverage; probe_C2
40 shapes -> 0 violations; under-cap base identity x4 True.
2026-09-22 13:40:13 +05:30
kshitijk4poor
f3642a27d8 docs(agent): say why the extension pass re-renders after the cap pre-check
The comment claimed "marker widths shrink as records leave a gap" — but
shrinkage can only make the render smaller, which the pre-check already
tolerates. The exact re-render exists for the opposite case: a gap's
first index can gain a digit or thousands separator (999 -> 1,000, +2
chars) when the added record moves a marker, and that growth can push a
render sitting at cap over it. Comment only; no behaviour change.

Finding: $D/regate/C.md suggestion 1 (agent/context_compressor.py:3619).
2026-09-22 13:40:13 +05:30
kshitijk4poor
be63d5bcd3 fix(agent): hand out lean-sampling headroom round-robin across slices
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.

Before -> after (records per slice, fill unchanged):
  10000x100  [1,1,1,1,1,1,1,8]      -> [1,2,2,2,2,2,2,2]      fill 0.944
   8000x100  [2,2,2,2,2,2,2,5]      -> [2,2,2,2,2,3,3,3]      fill 0.957
   4000x300  [4,4,4,4,4,4,4,11]     -> [4,5,5,5,5,5,5,5]      fill 0.987
   2000x600  [9,9,9,9,9,9,9,15]     -> [9,9,10,10,10,10,10,10] fill 0.995
    500x2000 [37,...,37,39]         -> [37,...,37,38,38]      fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.

Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).

Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
2026-09-22 13:40:13 +05:30