jev_cycles_report.py renders the cycles/freed/floor/end-state table from
jev_cycles.py outputs; README carries the exact commands, thresholds and
cost so the diminishing-returns and stuck/fallback findings can be
reproduced. Fallback records now count the real tool calls.
Closed-book arms scored the summary with its session_search pointer unused
(43% vs 79% on the same banks) and made external compactors look like wins.
Bare policy names stay available as an explicit opt-in floor.
Against summary + one session_search round-trip (78.9% @ 55K) Jev's default
arm (75.5% @ 115K) loses on recall and retains 2.1x the tokens; the earlier
closed-book comparison scored our summary with its recovery pointer unused.
Programmatic tool-result removal as the primary compaction leaves ~2x the
context billed every turn and tightens compaction cadence every cycle (each
one a prompt-cache break); the summary's one-time ~90% reduction keeps a
long conversation cache-warm. The recall gap points at summariser retention
of identifiers in assistant text, not at a cadence change.
jev_cycles.py compacts a lineage with Jev every time the estimate crosses a
threshold and records freed tokens, text floor and fitting stage per cycle.
Six runs show the floor (never-removed user/assistant text) is the binding
limit: freed-per-cycle decays to 0-8% and a 200K-window host is stuck after
0.42M tokens; >~500 tool calls between compactions cannot fit the 32K state.
Scorecard gains a decisive verdict section.
Adds an eval-only budget selector to the Jev arm (rank tool pairs by Jev
keep_result or by recency, keep until a token budget) so Jev's judgment can
be separated from the effect of keeping text verbatim, and records the
3-transcript scorecard: default Jev +32 pts recall at 2.1x retained tokens,
1/9 the compaction cost; at matched budget Jev ranking ties recency.
Adds a Python port of tamaratran/fast-jev-compaction as an eval policy
(`engine: jev`): TypeSafe's Jev decision model scores every tool call and
result over the whole history and stale ones are dropped or truncated, with
no summary and no rewriting of user/assistant text. Transport is OpenRouter's
Decisions API (~typesafe/jev-latest). A state that cannot fit the plugin's
25K-token ceiling is recorded as a fallback (the plugin's own behaviour)
rather than scored.
Every arm now reports what its compaction step cost (calls, tokens, USD —
Jev reports cost directly; summary calls are metered through the
compressor's call_llm binding and priced at OpenRouter list), and the run
writes the harness's own answer/judge token bill so the eval's spend is
visible. The question cache key includes --cap-tokens so different caps on
one transcript no longer share a bank.
Why: measure remaining tokens, compaction cost and recall accuracy of the
"decide, don't summarize" approach against current/lean on real Hermes
lineages before deciding whether any of it belongs in the compressor.
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same
transcripts: chunk-file reads until its REAL auto-compaction fires (verified
via compacted events in the rollout jsonl; peak 455-483K vs its 258K
window), then quizzes post-compaction with the identical question banks and
judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean
closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access
over its rollout history — the session_search differentiator, measured.
lean+recovery 68.3% avg recall @ 49K retained vs current 45.8% @ 162K —
+22.5pts at 0.30x tokens. Anchor index moved GUI needle-fact recall
23.3->60.0 closed-book, 46.7->80.0 with recovery.
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
file paths, error strings, handles, URLs from the compacted region into a
bounded indexed summary section. LLM-free, so needle identifiers cannot be
paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
summarizer input carries ONLY the compacted region (head/tail sentinels
never reach the serialized turns body) in both legacy and lean modes.
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
mine anchor identifiers
Measures recall accuracy vs tokens retained across compaction policies.
Real transcripts in, LLM-generated recall exam from the summarized region,
per-policy answer+judge passes, scorecard out.