Commit Graph

16 Commits

Author SHA1 Message Date
teknium1
00919d1e8a evals(compaction): document the Jev repeated-compaction recipe and add the table renderer
jev_cycles_report.py renders the cycles/freed/floor/end-state table from
jev_cycles.py outputs; README carries the exact commands, thresholds and
cost so the diminishing-returns and stuck/fallback findings can be
reproduced. Fallback records now count the real tool calls.
2026-09-19 11:58:37 -07:00
teknium1
085ed46b63 evals(compaction): default arm is current+recovery, the production compaction path
Closed-book arms scored the summary with its session_search pointer unused
(43% vs 79% on the same banks) and made external compactors look like wins.
Bare policy names stay available as an explicit opt-in floor.
2026-09-19 11:58:37 -07:00
teknium1
dba031b5cb evals(compaction): add the shipping current+recovery arm to the Jev scorecard
Against summary + one session_search round-trip (78.9% @ 55K) Jev's default
arm (75.5% @ 115K) loses on recall and retains 2.1x the tokens; the earlier
closed-book comparison scored our summary with its recovery pointer unused.
2026-09-19 11:58:37 -07:00
teknium1
6befc26f46 evals(compaction): correct the Jev scorecard verdict — do not adopt the prune-first rule either
Programmatic tool-result removal as the primary compaction leaves ~2x the
context billed every turn and tightens compaction cadence every cycle (each
one a prompt-cache break); the summary's one-time ~90% reduction keeps a
long conversation cache-warm. The recall gap points at summariser retention
of identifiers in assistant text, not at a cadence change.
2026-09-19 11:58:37 -07:00
teknium1
d659843dbf evals(compaction): Jev repeated-compaction simulator, cycle data and verdict
jev_cycles.py compacts a lineage with Jev every time the estimate crosses a
threshold and records freed tokens, text floor and fitting stage per cycle.
Six runs show the floor (never-removed user/assistant text) is the binding
limit: freed-per-cycle decays to 0-8% and a 200K-window host is stuck after
0.42M tokens; >~500 tool calls between compactions cannot fit the 32K state.
Scorecard gains a decisive verdict section.
2026-09-19 11:58:37 -07:00
teknium1
770c47779f evals(compaction): matched-budget Jev/recency arms + 2026-09-19 fast-jev scorecard
Adds an eval-only budget selector to the Jev arm (rank tool pairs by Jev
keep_result or by recency, keep until a token budget) so Jev's judgment can
be separated from the effect of keeping text verbatim, and records the
3-transcript scorecard: default Jev +32 pts recall at 2.1x retained tokens,
1/9 the compaction cost; at matched budget Jev ranking ties recency.
2026-09-19 11:58:37 -07:00
teknium1
139047ae2a feat(evals): fast-jev-compaction arm + compaction cost metering in the compaction eval
Adds a Python port of tamaratran/fast-jev-compaction as an eval policy
(`engine: jev`): TypeSafe's Jev decision model scores every tool call and
result over the whole history and stale ones are dropped or truncated, with
no summary and no rewriting of user/assistant text. Transport is OpenRouter's
Decisions API (~typesafe/jev-latest). A state that cannot fit the plugin's
25K-token ceiling is recorded as a fallback (the plugin's own behaviour)
rather than scored.

Every arm now reports what its compaction step cost (calls, tokens, USD —
Jev reports cost directly; summary calls are metered through the
compressor's call_llm binding and priced at OpenRouter list), and the run
writes the harness's own answer/judge token bill so the eval's spend is
visible. The question cache key includes --cap-tokens so different caps on
one transcript no longer share a bank.

Why: measure remaining tokens, compaction cost and recall accuracy of the
"decide, don't summarize" approach against current/lean on real Hermes
lineages before deciding whether any of it belongs in the compressor.
2026-09-19 11:58:37 -07:00
Teknium
4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
Teknium
0ee5ae61e2 docs(evals): real Codex CLI head-to-head arm + results
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same
transcripts: chunk-file reads until its REAL auto-compaction fires (verified
via compacted events in the rollout jsonl; peak 455-483K vs its 258K
window), then quizzes post-compaction with the identical question banks and
judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean
closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access
over its rollout history — the session_search differentiator, measured.
2026-08-15 18:41:24 -07:00
Teknium
e2990428e7 docs(evals): ship transcript-building scripts + full eval detail in-repo
- scripts/reconstruct_lineage.py: rebuild full uncompacted lineage
  transcripts from a state.db COPY (descendant-tree walk, content-hash
  dedupe, synthetic-artifact strip, system_prompts hash resolution)
- scripts/replay_lineage.py + scripts/build_html_report.py: replay a 500K
  prefix through any checkout's compressor and render before/after
  side-by-side with compaction artifacts color-coded
- README: transcript-building workflow, scoping-tripwire section
- results/SCORECARD-2026-08-15.md: full per-transcript scorecards, all 4
  exam question banks, survival analysis, methodology + caveats
2026-08-15 17:40:43 -07:00
Teknium
9146f4c851 fix(evals): explicit encoding on all file I/O (ruff PLW1514) 2026-08-15 17:07:46 -07:00
Teknium
44536bd41b docs(evals): 4-transcript compaction scorecard
lean+recovery 68.3% avg recall @ 49K retained vs current 45.8% @ 162K —
+22.5pts at 0.30x tokens. Anchor index moved GUI needle-fact recall
23.3->60.0 closed-book, 46.7->80.0 with recovery.
2026-08-15 17:01:22 -07:00
Teknium
c4bbb14e52 feat(compression): mechanical anchor index + region-scoping tripwire
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
  file paths, error strings, handles, URLs from the compacted region into a
  bounded indexed summary section. LLM-free, so needle identifiers cannot be
  paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
  golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
  summarizer input carries ONLY the compacted region (head/tail sentinels
  never reach the serialized turns body) in both legacy and lean modes.
2026-08-15 17:01:22 -07:00
Teknium
7a82457ede feat(compression): digest noise filter + FTS5 recovery sim + digest-aware query hints
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
  digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
  session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
  mine anchor identifiers
2026-08-15 17:01:22 -07:00
Teknium
8fe9025abd feat(compression): lean tail mode + recovery-aware eval arm
Lean mode (tail_mode='lean', default stays 'legacy'):
- tail budget = clamp(2.5% of window, 10K, 25K) instead of 0.20*window
- stale tail tool results demoted to session_search recovery stubs
- chunked identifier-preserving digests of the compacted region (map-reduce,
  pristine pre-prune tool contents)
- verbatim user messages embedded in summary (codex retention-by-role rule)
- deterministic session_search recovery footer

Eval: policies matrix gains lean + a '+recovery' arm giving the answerer one
simulated session_search round-trip against the archived region.
2026-08-15 17:01:22 -07:00
Teknium
33242d5ee0 feat(evals): compaction recall eval harness
Measures recall accuracy vs tokens retained across compaction policies.
Real transcripts in, LLM-generated recall exam from the summarized region,
per-policy answer+judge passes, scorecard out.
2026-08-15 17:01:21 -07:00