168 Commits

Author SHA1 Message Date
unsupportedpastels
7a2ee65bd2 fix(compression): classify complete legacy notice fragments
Preserve genuine selections around comma-normalized historical notices and require closed quoted question/choice fields before recognizing a notice suffix. Cover real historical producer outputs, embedded terminators, prefix lookalikes, and explicit status through public compression dispatch.
2026-10-01 02:28:11 +00:00
unsupportedpastels
bb873ade25 fix(compression): recognize normalized legacy headless notices
Match the complete single-query notice rather than its prefix, and recognize comma-normalized headless notice envelopes before filtering individual selections. Preserve current explicit status and independent non-notice selections.

Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
2026-10-01 01:27:27 +00:00
unsupportedpastels
ddc55890f0 fix(compression): match legacy clarify notices by producer shape
Match the classic CLI cancel notice as its whole text so real answers that start "The user cancelled" survive, add the hermes chat -q single-query notice, and drop only notice items from a legacy multi-select so other selections are kept.
2026-09-30 00:32:09 +00:00
unsupportedpastels
0145fc95b4 fix(compression): treat legacy clarify cancel notice as non-response
The pre-#127760 classic CLI Ctrl+C path persisted "The user cancelled. ..." as a status-less user_response. Add it to the legacy sentinel prefixes, now a named module constant, and to the regression parameterization. Drop a leftover debug print.
2026-09-29 21:09:28 +00:00
unsupportedpastels
65f5d631ed fix(compression): distinguish legacy clarify notices from answers
Keep explicit status authoritative and restore sentinel recognition only for persisted pre-status responses. Exercise repeated pruning and public compaction with both legacy shapes, multi-select answers, and large question payloads.

Co-authored-by: Yuan Li <dskwelmcy@163.com>
2026-09-29 18:59:53 +00:00
Yuan Li
a50078d2d3 test(compression): cover legacy clarify result shapes
(cherry picked from commit 1e4c3014b6f246f3d264efc86c1f071be81f7637)
2026-09-29 18:57:12 +00:00
Siddharth Balyan
5eea87882a Clarify: one question shape and one result shape on every surface (skip, cancel, timeout and undelivered are told apart) (#127760)
* refactor(desktop): split clarify-tool.tsx into a clarify/ folder

Pure moves, no behaviour change. Parsing, the question-card core
(shell, choice rows, question block), the delivery watchdog, and each
pending/settled card get their own file so the core can be reused.

* feat(clarify): one questions[] shape with per-question status and one outcome

The clarify tool now takes only questions=[{question, choices?, multi_select?}].
A wrong shape is a tool error that names the right one, so a model that
sends the old top-level question/choices corrects itself on the next call.

Every result has the same shape on every surface: each response carries
status (answered, skipped, unanswered) with user_response null unless
answered, and the result carries one outcome (submitted, cancelled,
timed_out, undelivered) plus an optional surface-written notice. The
timeout and cancel sentences that used to pose as the user's answer are
gone, and the compressor reads status instead of matching their prefixes.

The callback contract is callback(questions) -> {answers, outcome, notice?}
(None = skipped, missing qid = unanswered). tui_gateway settles a batch the
same way for the last lock, a cancel, an interrupt and the deadline, and
clarify.lock keeps a null answer as a skip. A multi-select answer that is
not a JSON array counts as one typed ("Other") answer. The tool-row preview
reads the first question.

* feat(clarify): every surface asks through the one question card

CLI: the batch panel is the only panel; Enter on an empty field skips a
question, Ctrl+C cancels and keeps the locked answers, the deadline returns
timed_out. -q and -z return undelivered with a no-user notice.

Ink TUI: the single-question prompt is gone; the card supports multi-select
(Space or a digit toggles, Enter locks, typed Other text joins the picks),
an empty submit skips, Esc cancels, and "Batch" names are dropped.

Desktop: the single-question card is deleted and its keyboard handling moved
into the one card (arrows, letters and digits, auto-advance, Enter picks then
confirms). Confirm enables at one answer; blank questions lock null; Skip and
a composer message cancel; the settled card shows No answer for unanswered
questions. The bots room card follows the same contract.

Messaging: one card per question with a "Reply skip to skip" line in every
locale; timeouts and delivery failures map to timed_out and undelivered.

* fix(clarify): cancelled for a released messaging card, labels win over the skip word

A prose reply to a messaging choice card, /new and session end released
the wait with "", which read as timed_out. They now resolve with a
CANCELLED marker and the tool gets outcome cancelled.

A typed reply that matches a choice label ("Skip") now resolves to that
choice; the skip word applies only when no choice matches.

The bots room card sends the picked labels as they are; the tool already
strips the recommendation label. Unused OUTCOMES and a new docstring and
comment are removed.

* fix(tui): re-editing a multi-select answer keeps its picks

The Ink card stored a multi-select answer as display text ("A, B"), so
revisiting the question put the whole string into Other and re-locked it
as one item. The answers map now keeps the raw JSON array, the card
restores the picks on revisit, and only the display lines format it.

Comments that earlier edits reworded are restored to their original
words, minus the phrases the change made wrong. The compute-host clarify
lock accepts None for a skipped question.

* test(clarify): align three checks with the one-answer confirm and the regenerated keys

The desktop Cmd+Enter test now expects a send with one answer (the blank
question locks null), the compressor test expects the single-answer summary
the code gives, and locales/_keys.desktop.json is regenerated for the
removed and added clarify strings.

* test(tui): the one-question card header is singular
2026-09-29 18:33:44 +05:30
Siddharth Balyan
98278833e2 Compaction follows compression.threshold again: no default 256K token cap (#117915, #125235) (#126064)
* fix(compression): compaction follows the ratio again; no default token cap

compression.threshold_tokens defaulted to 256000 since #115986, so every
window above ~341K compacted at 256K regardless of compression.threshold
or model_thresholds: a 1M-window model compacted at 25% of its window,
and threshold: 0.8 still compacted at 256K. No single token count suits
windows from 64K to 1M+, so the cap goes back to opt-in (null) and the
ratio decides, as it did before #115986.

The template seeder and `hermes doctor --fix` copied the 256000 default
into config.yaml, where it reads as a user choice and would keep capping
those installs. Config v47 drops threshold_tokens only when it equals
256000; any other explicit cap and an explicit null are preserved.

The rest of #115986 stays: the model-switch warning quoting the real
trigger and the shared _derive_trigger are correct with any default.
Tests that encoded 256K now derive expectations from DEFAULT_CONFIG.

* docs(compression): threshold_tokens is an optional cap, default null

User guide, developer guide and delegation page described the 256K
default cap; they now describe the ratio trigger as the default and
threshold_tokens as an opt-in cost ceiling.
2026-09-28 07:29:40 +00:00
kshitijk4poor
fb86bc708d fix(context): redact full tool-call args before the summarizer cut
62ceddd342 cut raw args to HEAD+4096 before redaction to save time. The
PEM redaction pattern only matches a complete BEGIN...END block. A long key
whose END fell past the cut stayed unredacted, and once an earlier key was
redacted and the text shrank, its body landed in the 1200-char head that
goes into the persisted summary. Go back to the BASE order: redact the full
args, then apply the MAX/HEAD cut. This is a cold path (once per summarized
call per compaction), and _SUMMARY_INPUT_MAX_CHARS still bounds the prompt.
Extend the kept canonical-args test with a two-PEM input that leaks on
62ceddd342 and passes now.
2026-09-27 18:38:58 +05:30
kshitijk4poor
66240ddc59 fix(context): restore marker import in tests, drop unused compressor imports
The arg-truncation removal deleted the only uses of the marker constants in
agent/context_compressor.py (ruff F401), and the test module had been
importing _COMPRESSION_MARKER_PREFIX through it, so test_context_compressor
line 137 raised NameError. Import it from its home, agent.compression_marker.
2026-09-27 18:38:58 +05:30
kshitijk4poor
697f07a56f test(context): pin byte-exact old tool-call args with one red-on-base test
The PR's two tests passed on the base code (the prune boundary never
reached their calls). Replace them with one invariant test that is red on
base: six old 3000-char write_file arguments must survive the prune
unchanged. Drop the PR's canonical-state test file (tail-canonical shape
contradicts #61932's pressure demotion).
2026-09-27 18:38:58 +05:30
JoaoMarcos44
a7baa5f5eb fix(context): keep compaction history lossless
(cherry picked from commit d51c8f4f5096badfd0beddd78646617643f6028f)
2026-09-27 18:38:58 +05:30
kshitijk4poor
cd5dcd1458 test(compression): parametrize the transport-error case over exception instances
Why: the test parametrized a label string that a ternary in the body mapped
back to an exception, and its name still said "api_timeout" although it also
covers an APIConnectionError carrying the stall marker. Parametrize the two
exception instances directly (with ids) and rename the test to
test_transport_errors_stay_terminal_network_failure. No behaviour change; still
one parametrized test.
2026-09-27 18:19:07 +05:30
kshitijk4poor
f05d11c74b test(compression): pin the TimeoutError guard on the Codex stall reclassification
The isinstance(e, TimeoutError) guard was untested: an APIConnectionError
whose text contains the stall marker must stay a terminal network failure
(#29559/#94448). Parametrize the existing api-timeout test with that input
(red when the guard is removed), and move both #124077 tests into
TestStreamingClosedFailure reusing _fail_on_main instead of a duplicate helper.
2026-09-27 18:19:07 +05:30
kshitijk4poor
8c9f2162c7 refactor(compressor): key the Codex stall on the stream guard's shared marker
Co-authored-by: happy5318 <5318happy@users.noreply.github.com>
(cherry picked from commit 83c0a8bb9c7d0e611bafe436d2e4da77144c26ec)
2026-09-27 18:19:07 +05:30
kshitijk4poor
bf8099dab5 test(compression): trim #124077 tests to two invariants, drop local-env warm-ups
The zoneinfo/pydantic plugin warm-up imports were author-machine workarounds
that ran at collection time in CI; remove them. Collapse the class to one
helper and two tests: the Codex stall takes the timeout ladder without the
terminal network-failure flag, and APITimeoutError still sets it.

Co-authored-by: happy5318 <5318happy@users.noreply.github.com>
(cherry picked from commit 1cda73e0ac2d073066741a964b5166b16e9caf34)
2026-09-27 18:19:07 +05:30
happy5318
56ee7d6f09 fix(compression): classify a Codex stream-guard stall as a timeout, not a terminal network failure (#124077)
## Thinking Path
When the Codex auxiliary stream guard aborts a compaction summary mid-stream
it raises `TimeoutError("Codex auxiliary Responses stream stalled: no new
output for 60.0s ...")`. The message contains neither "timeout" nor
"timed out", so `_classify_summary_failure` returned `timeout=False` while
`_is_connection_error` (which matches the type name "Timeout") returned
`streaming_closed=True`. The terminal network-failure flag then armed an
unconditional abort (`_TERMINAL_SUMMARY_FAILURES`), bypassing the retry
ladder and the deterministic fallback summary — on turn-start preflight
compression that ends in "Auto-resetting session after compression
exhaustion", wiping the session.

### What Changed
`agent/context_compressor.py` `_classify_summary_failure`:
- `timeout` is now computed first, and additionally matches `isinstance(e,
  TimeoutError)` and the "stalled" message shape (the actual text the Codex
  guard emits).
- `streaming_closed` is `_is_connection_error(e) and not timeout` — a
  timeout keeps its retry-ladder semantics and can never arm the terminal
  network-failure abort.

### Tests
New `TestSummaryFailureClassification124077` in
`tests/agent/test_context_compressor.py`:
- classify: Codex stall → `timeout=True, streaming_closed=False`.
- classify: plain `ConnectionError` stays `streaming_closed=True` (no
  regression on the premature-close class).
- classify: a "timed out" message on a non-TimeoutError type stays a
  timeout and is excluded from `streaming_closed`.
- integration: the stalled-summary path in `_generate_summary` does NOT arm
  `_last_summary_network_failure`.

### Verification
- RED/GREEN double proof via git stash: pre-fix 3 failed, post-fix 4/4 pass.
- Regression: 15 existing failure-classification tests pass
  (network_failure / premature_stream / empty_content / auth / truncation).
- ruff clean on changed files.

### Notes
Local test env: this checkout's venv python is a symlink into
`<home>/hermes-agent/.hermes-runtime/...`, so stdlib `zoneinfo` first-import
and pydantic's plugin `distributions()` scan under the real-home IO guard
needed the collection-time warmups at the top of the test file. CI
interpreters are not symlinked into the home — those two warm-up blocks are
no-ops there.

## Related
#124077 (issue). Family: #124078 (the same stall's template trigger),
#108104 (`auxiliary.compression.no_progress_timeout`).

(cherry picked from commit c5399fbdee44f2db1172f22632cbf348a2c7cab5)
(cherry picked from commit bce26a99791c09911d8a008a1922c512f5c1fcea)
2026-09-27 18:19:07 +05:30
ahisblessed
77b9ea0831 fix(agent): route every model-visible elision through the non-imitable compression marker (#121548)
The #83714 fix hardened only the tool-call args renderer; five other
renderers (plus three same-idiom siblings) still composed the bare
truncation marker, which the model imitated from replayed context into
new durable writes (#83435).

This routes all model-visible elision in agent/ through shared helpers
in agent/compression_marker.py:

- elide(text, limit): head + marker cap with accurate omitted/total counts
- elide_middle(text, head, tail): kept head/tail with the middle elided

The elision marker's first sentence is byte-identical to the args marker's,
so the existing _COMPRESSION_MARKER_RE (and the dispatch-boundary guard in
tool_dispatch_helpers) rejects a copied marker regardless of which renderer
leaked it; only the tool-call-specific second sentence is dropped so the
marker still fits small caps (the clarify summary cap is 199 chars).

Routed sites:
- context_compressor.py: static-fallback turn renderer, _sum_clarify(),
  summarizer prompt builder (middle elision), active-task snapshot,
  lean user-message quotes (2 sites)
- skill_preprocessing.py: inline-shell output embedded into skill bodies
- lsp/reporter.py: truncate() for <diagnostics> tool output
- verification_stop.py: verify-on-stop nudge output summary

Regression test asserts the imitable idiom appears nowhere in agent/
strings (comments excluded), so a new open-coded renderer cannot land
silently.

Fixes #121548

(cherry picked from commit 69c63d50b4ababedf4fb4ff88c08f5d1e144a79f)
2026-09-26 23:49:25 +05:30
kshitijk4poor
74dc4ac1e4 fix(agent): a 403/402 saying "overloaded" is not an overload strike
An auth/quota error whose text also says "overloaded" went through the
overload branch: it bumped the durable sustained-overload streak even though
it always aborts as summary_auth_failure (#29559). After the credential was
fixed, the next real 503 then inherited that streak and committed the lossy
fallback immediately, skipping the one-abort grace (#115906).

Compute the access/quota classification once and keep such errors out of the
overload branch. With that, the overload-escalation path never runs for an
auth error, so the keep_auth escape hatch (and its hard-coded flag-name
compare) in _clear_terminal_summary_failures is dead and goes away.

The existing [403_overloaded] case now also asserts the streak stays 0.

Co-authored-by: Yuan Li <dskwelmcy@163.com>
2026-09-26 22:42:38 +05:30
kshitijk4poor
507e856745 fix(agent): keep the auth/quota abort when an overload escalates
The escalation clear-all of terminal summary flags also cleared the
auth/quota flag the SAME error had just set. A 403/402 or quota-worded
error that also says "overloaded"/"at capacity" classifies as both, so on
strike 3 it committed the lossy fallback instead of aborting - breaking
the #29559 invariant that access/quota failures always preserve the
session. Keep the auth flag when the current error is access/quota; a
stale auth flag from an earlier failure is still cleared.

The kept escalation test is parametrized with a 403 "provider overloaded"
case that must abort all three times with summary_auth_failure.
2026-09-26 22:42:38 +05:30
kshitijk4poor
818faaf16d fix(agent): label the escalated overload commit summary_overload_degraded
On a compressor with a distinct summary_model, the first overload takes the
aux->main one-shot retry, which pre-sets telemetry failure_class to
aux_model_fallback; the `telemetry.get(...) or` chain then hid the
escalation and operators never saw the documented
summary_overload_degraded class. The degraded flag now wins. It is always
initialised (__init__ and _begin_compress_attempt, the only path into
_fallback_summary_for_window), so the defensive getattr goes too; no test
builds a compressor via __new__ and reaches this method without compress().
The kept fresh-bind test now asserts the label instead of excusing it.
2026-09-26 22:42:38 +05:30
kshitijk4poor
5021e3633b fix(agent): let an escalated overload clear stale terminal summary flags
Only a successful summary clears the network/empty/truncated/auth failure
flags, and _abort_on_summary_failure aborts on ANY of them. One earlier
streaming_closed failure on a long-lived (gateway-cached) compressor
therefore kept every later sustained-overload attempt aborting past the
3-strike escalation, so the #123167 compression_exhausted wipe persisted.
When the escalation fires, the latest failure class decides: clear the
other terminal flags so compress() commits the degraded fallback.
2026-09-26 22:42:38 +05:30
kshitijk4poor
e046fdd323 test(agent): keep two sustained-overload invariant tests
Trim the salvaged stack to the two invariants the fix exists for: an
overload aborts twice then commits the fallback, and the budget survives
a fresh compressor bound to the same session. The dropped cases re-prove
existing abort_on_summary_failure / success-reset / rotation plumbing.
2026-09-26 22:42:38 +05:30
Yuan Li
2305385042 fix(agent): persist the summary-overload budget in session state so fresh bindings inherit it
Review P1 on #123186 (ehz0ah): the N=3 sustained-overload budget was
object-local. The gateway binds a fresh compressor to the same session on
every turn / cache eviction, and restart / API-server requests construct
one too, so each fresh instance restarted the budget at zero and a
sustained summary-provider outage walked the session back into
compression_exhausted and auto-reset — the exact wipe the PR exists to
prevent. Reproduced by the reviewer with three fresh compressors bound to
one session: every attempt ended at counter=1.

Persist the streak as sessions.compression_overload_streak through the
same durable channel the fallback streak and recovery deadline already
use (#100185): SessionCompressionMixin get/set pair over the sessions
row, loaded in the compressor's durable-load block, written on every
change (overload abort, successful summary, committed boundary, runtime
model switch).

- Compression rotation carries the streak to the child row at the
  boundary, same as the fallback streak: the parent's value is read
  before the bind, re-applied, and persisted onto the fresh child row.
- A completed compaction boundary settles the budget to 0 — including
  the committed degraded fallback, so a recovered provider regains a
  full budget instead of the session staying degraded forever.
- Schema event appended to SCHEMA_HISTORY["sessions"] (seq 28, after
  compression_recovery_deadline) so salvage replay maps the new column.
- Fresh-instance regression: two aborts on one compressor, then a fresh
  compressor bound to the same session inherits streak=2 and the third
  session-wide attempt commits the fallback; plus a rotation carry-over
  test. Both fail red against the memory-only implementation.

(cherry picked from commit 3742fee8d03d7077d23b0dac3a3c79fe48e102ac)
2026-09-26 22:42:38 +05:30
Yuan Li
a5e8ad9f02 fix(agent): bound sustained summary-overload aborts so they cannot guarantee a session wipe
One overload abort preserves the transcript so a retry can still win (#115906).
But when every summary attempt keeps aborting under a sustained outage the
transcript only grows until the session exits compression_exhausted, which the
gateway answers with an auto-reset discarding the ENTIRE transcript — bounded
middle-window loss becomes total session loss, deferred.

After 3 consecutive overload aborts in one session the overload stops counting
as a terminal summary failure and compress() commits the deterministic fallback
(failure_class=summary_overload_degraded) — the same bounded degrade the
repeated-stall ladder already takes (#112420). A successful summary resets the
budget; abort_on_summary_failure=true still hard-aborts every attempt.

Fixes #123167

(cherry picked from commit 2941aadffa71a3623aee26dfb1109106b0555741)
2026-09-26 22:42:38 +05:30
teknium1
21cc9de8c9 test: restore agent-lane aux/compressor regression guards dropped by #120071
Restored (repaired to drive production code, mocks only at the provider
boundary; each verified red when its fix is reverted):

- test_auxiliary_client.py::TestCodexAuxiliaryAdapterCompletedResponse::
  test_completed_response_with_null_output_does_not_crash
  A Codex-compatible host returning a completed Responses object with
  output=None yields an empty 'stop' turn instead of TypeError (#33368).
  Repaired: the old test patched the internal stream consumer; it now goes
  through the real completed-response path via the fake client.
- test_aux_progress_streaming.py::TestContentBearingProgress::
  test_content_free_frames_still_record_ttfp_timing
  Keepalive frames still fire the provider-response (TTFP) hook after the
  #96707 progress gating (#96945/#96963).
- test_context_compressor.py::TestStreamingClosedFailure::
  test_premature_stream_close_is_transient_network_failure[x2]
  httpcore 'incomplete chunked read' / 'response ended prematurely' are
  classified by the real classifier as transient: network-failure flag set
  (session preserved) and a cooldown shorter than a generic failure's
  (#18458). Repaired: no longer patches _is_connection_error, and compares
  cooldowns instead of pinning the 30s constant.
2026-09-23 10:34:54 -07:00
teknium1
2002f03e42 test: purge low-value tests, lane py02 (261 removed)
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
2026-09-23 03:15:26 -07:00
kshitijk4poor
600bee9201 test(agent): derive the lean-sampling marker bound from _SAMPLED_INPUT_SLICES
tests/agent/test_context_compressor.py:3397 asserted
`sampled.count("chars elided") < 7`, hardcoding the n-1 gap count that
the 8-slice constant implies and that the comment explained in prose.
Use `ContextCompressor._SAMPLED_INPUT_SLICES - 1` so the invariant
("the extension pass closed at least one initial gap") survives a
change to the slice count instead of silently becoming a change
detector. `_SAMPLED_INPUT_SLICES - 1 == 7` today, so the assertion
value is unchanged.
2026-09-22 13:40:13 +05:30
kshitijk4poor
c2a2a62697 test(agent): cover slices merged when the extension pass closes a gap
No fixture exercised the final `selected = _merged(selected)` in
_sample_summary_records: the re-gate mutant that dropped it survived
all 8 bounding tests and 40 random probe shapes. Add 20 x 8.5K records
(170K > 160K cap, headroom > one 8.5K gap) so the round-robin extension
closes the gaps between the newer slices. Assert consecutive shown
records are joined by exactly "\n\n" (an unmerged adjacent pair renders
as "...xxx[USER]: record-..." with no separator), remaining gaps carry a
correctly numbered marker, fewer than 7 markers survive, and
sampled_record_count equals the whole records shown.

RED with the final `_merged` call removed: AssertionError (14, 15);
GREEN at head.

Finding: $D/regate/C.md suggestion 2 (M-c, agent/context_compressor.py:3623).
2026-09-22 13:40:13 +05:30
kshitijk4poor
be63d5bcd3 fix(agent): hand out lean-sampling headroom round-robin across slices
The budget-extension pass in _sample_summary_records grew the newest
slice until it hit the cap and only then moved to older slices, so on
mid-size records all headroom went to one region: 10K x 100 records ->
records/slice [1,1,1,1,1,1,1,8] (newest 4.3x the mean). That contradicts
the design comment above (regions must not consume each other's budget)
and the docs' "evenly sampled". Wrap the per-slice loop in a
`while grew` round so each slice adds at most ONE whole record per
round (newest grows backward, others forward); the cap pre-check and
exact re-render check are unchanged.

Before -> after (records per slice, fill unchanged):
  10000x100  [1,1,1,1,1,1,1,8]      -> [1,2,2,2,2,2,2,2]      fill 0.944
   8000x100  [2,2,2,2,2,2,2,5]      -> [2,2,2,2,2,3,3,3]      fill 0.957
   4000x300  [4,4,4,4,4,4,4,11]     -> [4,5,5,5,5,5,5,5]      fill 0.987
   2000x600  [9,9,9,9,9,9,9,15]     -> [9,9,10,10,10,10,10,10] fill 0.995
    500x2000 [37,...,37,39]         -> [37,...,37,38,38]      fill 0.998
Re-gate probe (40 random shapes): 0 violations, worst fill 0.892 -> 0.930,
worst char drift 4.27 -> 1.83; under-cap output byte-identical to base.

Test: test_lean_sampling_oversized_middle_record_does_not_evict_tail now
asserts no slice holds more than mean+1 records (RED with the previous
newest-first order: [16,16,16,16,16,16,18]; GREEN: [16,16,16,16,16,17,17]).

Finding: $D/regate/C.md W1 (agent/context_compressor.py:3599-3623).
2026-09-22 13:40:13 +05:30
kshitijk4poor
33a1d46dfc feat(agent): record summary-input coverage telemetry (from #118390)
Lean sampling is the documented budget mechanism, but nothing reported
how much of a region actually reached the summarizer, which is the gap
the reporter narrowed #118362 to. _sample_summary_records now returns
record-level coverage counters alongside the bounded transcript and
_generate_summary writes them into the active compression telemetry
(summary_input_chars / sampled_chars / omitted_chars / record_count /
sampled_record_count / elided_record_count), so the existing
compression_attempt JSON log line carries them with no transcript
content. Char counters cover record content only, not separators or
markers. Elision markers also name the one-based record range they
stand for, so a reader can tell how many turns a gap hides, not only
how many characters.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-22 13:40:13 +05:30
kshitijk4poor
4410ced18f test: keep two behavioural lean-sampling tests
Replace the eight sampler tests with two invariants. The boundary test
now drives _generate_summary so it fails on main for the actual defect
(a retained section starting mid-record) instead of on a helper name,
and folds in the multi-paragraph case since blank lines inside a message
were what broke the string-split approach. Dropped: covers_multiple_
regions and generate_summary_lean_samples_structural_records (green on
main, no teeth), omission_markers_account_separators_exactly (asserted
only that a marker exists), newest_message_ending_in_blank_lines (moot
once records are structural), and preserves_oversized_last_record (a
record above the whole cap is unreachable through _serialize_for_summary
because _CONTENT_MAX bounds message bodies). The remaining oversized
test asserts behaviour — tail kept, oversized record bounded — not the
marker wording.
2026-09-22 13:40:13 +05:30
kshitijk4poor
a3a03d48d4 refactor(agent): sampler takes serialized records only
_sample_summary_input accepted either a flat string or a record
sequence and, for strings, rebuilt records with content.split("\n\n").
The PR's own third commit established that "\n\n" is not a record
delimiter (message bodies contain blank lines), so that fallback would
re-introduce split records for any caller that reached it. The only
production caller already passes records to _sample_summary_records;
drop the dual-signature wrapper and move the test callers to records.
2026-09-22 13:40:13 +05:30
joaomarcos
04fe735c57 fix: preserve structural record framing in lean summary sampling
Preserve record framing structurally so messages containing internal blank lines are not split into pseudo-records.

- agent/context_compressor.py:
  * Extract _serialize_records_for_summary() to retain turn boundaries as a sequence of records.
  * Add _sample_summary_records() to sample across serialized turn records directly without lossy double-newline splitting.
  * Correct separator accounting in omission markers at gap boundaries so separator counts match full serialized text.
  * Update _sample_summary_input() to support structural records while preserving flat-string fallback with trailing fragment pruning.
  * Wire _generate_summary() to use _serialize_records_for_summary() and _sample_summary_records() in lean mode.
- tests/agent/test_context_compressor.py:
  * Test multi-paragraph serialized messages maintain role boundaries and keep newest message paragraphs together.
  * Test newest message ending in blank lines does not produce empty tail anchor.
  * Test exact omission marker separator accounting.
  * Test end-to-end _generate_summary prompt assembly with multi-paragraph messages.

(cherry picked from commit f68f1bfcf037dc5ddecf6cc4032023cd88c496c8)
2026-09-22 13:40:13 +05:30
joaomarcos
723fc9ffb5 fix: bound oversized records and protect newest slice in lean summary sampling
(cherry picked from commit 1d0d938cc3572c992f3b6540486b98dbcc777429)
2026-09-22 13:40:13 +05:30
joaomarcos
3bdc07efef fix: sample complete summary records
(cherry picked from commit b557fc4534a4557f95540e29cb77ac8acfffa473)
2026-09-22 13:40:13 +05:30
fangliquanflq
af93d57ca7 fix(agent): preserve context when summary provider overloads 2026-09-19 10:35:44 -07:00
kshitijk4poor
39abfdf4ec fix(context-compressor): the marked-leaf guard must match the whole tail
Second review round on the previous commit, both findings reproduced:

- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
  about: a replayed leaf of head + marker followed by new content was never
  shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
  the marker to close the leaf, so that shape shrinks to head + marker with
  true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
  adds separator whitespace, so compact args with many keys could come back
  LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
  The helper now returns the caller's string unless the whole rewrite is a net
  reduction.

Tests cover both shapes plus a many-key compact payload.
2026-09-19 21:06:12 +05:30
kshitijk4poor
a5381a7dca fix(context-compressor): match the marker by position, and leave args byte-identical
Review findings on the previous commit, all reproduced:

- The "already marked" guard was a substring test, so a leaf that merely
  contains the marker — including one a model imitated into a new call, the
  #83714 failure mode itself — was exempt from shrinking forever. The marker
  is always written at `head_chars`, so the guard now tests that position: a
  1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
  JSON came back with inserted spaces and callers read it as a change
  (rewriting replayed history and counting a pressure hit for zero reclaim).
  Nothing replaced now returns the original string: a 547-char compact blob
  is byte-identical.
2026-09-19 21:06:12 +05:30
kshitijk4poor
1e216d1330 fix(context-compressor): only shrink args when it reclaims, never re-shrink
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:

- A leaf just over `head_chars` came back LONGER: a 628-char args blob
  shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
  compaction, so a leaf's marker was rewritten from the true count
  ("2,800 of 3,000 chars omitted") to a self-referential one
  ("223 of 423"), destroying the per-instance count property the marker
  relies on and churning those bytes on each pass.

A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
2026-09-19 21:06:12 +05:30
Daniel JB Clark
262a6436fa fix(context-compressor): stop injecting an imitable truncation marker into replayed tool_calls
Root cause for #83714 (write_file/patch_tool writing literal
"...[truncated]" into files, PR #83752's guard is the safety net, not
the fix): _truncate_tool_call_args_json() in the compression pass
shrinks long string values inside a PAST assistant message's
tool_calls[].function.arguments — the exact field that represents the
model's own prior generated output, replayed back to it verbatim on
every subsequent turn. The old marker, a bare "...[truncated]" suffix,
is indistinguishable from something the model itself could have
written (it's exactly the kind of terse ellipsis abbreviation models
already produce). A model conditioned on seeing itself "get away with"
that pattern in its own history imitates it in a new tool call,
writing the literal marker instead of real content.

This is the second bug from the same root text. The first (#11762,
MiniMax 400s from unterminated JSON) was fixed by shrinking inside the
parsed structure so the JSON stays valid, but kept the same visible
marker text — fixing the syntax problem while leaving the imitation
problem untouched.

Fix: replace the marker with one deliberately NOT shaped like prose a
model would write — distinctive non-ASCII delimiters, an explicit "not
part of the original tool call" disclaimer, and a per-instance
char-count that won't match the next omission point even if copied
verbatim. The shrunk value stays a plain string (not a nested object)
so the #11762 valid-JSON/matching-shape contract is unchanged — only
the marker text changed.

Checked context_compressor.py's other "...[truncated]" call sites
(_serialize_for_summary, _compact_fallback_turn, the user-message-only
one near _ACTIVE_TASK_MAX_CHARS) — none of them write into a value
that gets replayed as the main model's own assistant/tool_calls
history, so they don't share this priming risk and were left as-is.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:06:12 +05:30
kshitijk4poor
001206cbdc test(agent): default-cap matrix derives its expectations from DEFAULT_CONFIG
Hardcoded 256_000 / 204_000 literals would turn into change-detectors the day the shipped cap
moves; assert the contract instead — trigger == min(ratio-only trigger, cap) — for every window.
2026-09-19 16:28:57 +05:30
fangliquanflq
0946de78c0 test(agent): cover default compression cap across windows 2026-09-19 16:28:57 +05:30
fangliquanflq
489cb74630 fix(compression): bound default trigger at 256K tokens 2026-09-19 16:28:57 +05:30
KoNit-K
fdbcdef914 fix(agent): bound tail message floor by token budget 2026-09-19 16:21:37 +05:30
teknium1
cab9a27954 fix(tests): cover the 'no credentials were found' permanent-failure marker
An OAuth auxiliary provider with no login now raises 'no credentials were found'
(#114405 / #78996); the compressor must classify it as permanent instead of
retrying. Extend the existing missing-credential classifier test with that
wording so removing the marker from _SUMMARY_MISSING_CREDENTIAL_MARKERS goes red.
2026-09-18 11:11:44 -07:00
teknium1
5670f846cf fix(compression): mark failed skill-tool results and drop the duplicated stub suffix
Follow-up to the salvaged #112719 commit:

- `_skill_result_failure_suffix()` is shared by the `skill_manage` and
  `skills_list` stubs instead of two inline copies, and also fires on a
  payload that says `success: false` without an `error` string (previously
  such a result still compressed into a success-shaped line — the exact
  class the issue reports). The error preview is whitespace-collapsed so
  the stub stays one line even when the tool error carries newlines.
- The `skills_list` comment no longer names a `query` arg the schema does
  not have (`SKILLS_LIST_SCHEMA` takes `category` only).
- Tests trimmed to two invariants in the existing summarizer module
  (`tests/agent/test_context_compressor.py`): skill_manage names its ops
  (operations array and legacy flat shape) and keeps FAILED visible with
  the error text; skills_list renders category/count and marks failure,
  with skill_view as the unchanged control. The standalone 9-test file
  from the contributor PR is folded into these.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 16:59:43 -07:00
Teknium
3026f4a993 test(compression): trim batch clarify coverage to two invariants
Drop the sentinel-only batch test: a batch sentinel is already rejected by the
shared _is_clarify_non_response_sentinel list check that the existing sentinel
tests pin, so the case adds no new contract. Also add the contributor email
mapping for the cherry-picked commit so release CI can attribute it.
2026-09-09 09:21:55 -07:00
gaoanze888
8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
Teknium
e9313f6458 fix: let provider evidence adjudicate past-window preflight estimates 2026-09-07 14:11:41 -07:00