evals(compaction): matched-budget Jev/recency arms + 2026-09-19 fast-jev scorecard

Adds an eval-only budget selector to the Jev arm (rank tool pairs by Jev
keep_result or by recency, keep until a token budget) so Jev's judgment can
be separated from the effect of keeping text verbatim, and records the
3-transcript scorecard: default Jev +32 pts recall at 2.1x retained tokens,
1/9 the compaction cost; at matched budget Jev ranking ties recency.
This commit is contained in:
teknium1
2026-09-19 10:08:52 -07:00
committed by Teknium
parent 139047ae2a
commit 770c47779f
3 changed files with 145 additions and 5 deletions

View File

@@ -112,6 +112,13 @@ class JevOptions:
max_state_tokens: int = 25_000
max_request_tokens: int = 30_000
truncate_head_chars: int = 300
# Eval-only extension (not in the plugin): instead of thresholding, keep
# whole call+result pairs in rank order until `result_budget_tokens` of
# tool content is retained; everything else is dropped. `select="jev"`
# ranks by Jev's keep_result, `select="recency"` by position and never
# calls Jev — the control that tells whether Jev's ranking carries signal.
select: Optional[str] = None
result_budget_tokens: int = 0
@dataclass
@@ -475,6 +482,36 @@ class JevCompactor:
return {c.id: {"keep_call": noul(f"call_{c.id}"), "keep_result": noul(f"result_{c.id}")} for c in batch}
def _budget_decisions(self, messages, calls, candidates, answers) -> List[Decision]:
"""Keep ranked call+result pairs until the tool-content budget is spent; drop the rest."""
opt = self.opt
def pair_tokens(c: ToolCall) -> int:
return (c.result_chars + len(_json(c.input))) // 4
if opt.select == "jev":
ranked = sorted(candidates, key=lambda c: -answers.get(c.id, {}).get("keep_result", 0.0))
else:
ranked = sorted(candidates, key=lambda c: -c.result_index)
kept, spent = set(), 0
for c in ranked:
t = pair_tokens(c)
if spent + t > opt.result_budget_tokens:
continue
kept.add(c.id)
spent += t
out = []
for c in calls:
a = answers.get(c.id, {})
base = dict(id=c.id, tool=c.tool, keep_call=a.get("keep_call", 1.0), keep_result=a.get("keep_result", 1.0))
if c.pinned:
out.append(Decision(**base, action="keep", reason="pinned"))
elif c.id in kept:
out.append(Decision(**base, action="keep", reason="kept"))
else:
out.append(Decision(**base, action="drop_call", reason="call_dropped"))
return out
def compress(self, messages: List[Dict[str, Any]], current_tokens: int = 0, force: bool = True):
t0 = time.time()
opt = self.opt
@@ -483,16 +520,19 @@ class JevCompactor:
answers: Dict[str, Dict[str, float]] = {}
fitted = {"tokens": 0, "stage": ""}
batches: List[List[ToolCall]] = []
if candidates:
if candidates and opt.select != "recency":
fitted = fit_state(messages, calls, opt)
batches = batch_calls(candidates, fitted["tokens"], opt)
with ThreadPoolExecutor(max_workers=self.concurrency) as pool:
for result in pool.map(lambda b: self._ask_batch(fitted["state"], b), batches):
answers.update(result)
self.decisions = [
decide_call(c, *(answers.get(c.id, {}).get(k, 1.0) for k in ("keep_call", "keep_result")), opt)
for c in calls
]
if opt.select:
self.decisions = self._budget_decisions(messages, calls, candidates, answers)
else:
self.decisions = [
decide_call(c, *(answers.get(c.id, {}).get(k, 1.0) for k in ("keep_call", "keep_result")), opt)
for c in calls
]
kept = apply_decisions(messages, self.decisions, calls, opt.truncate_head_chars)
reasons = [d.reason for d in self.decisions]
self.stats = {

View File

@@ -65,6 +65,17 @@ POLICIES: Dict[str, Dict[str, Any]] = {
"engine": "jev",
"jev": {"keep_threshold": 0.15},
},
# Matched-budget pair (eval-only extension): keep 60K tokens of tool
# call+result pairs ranked by Jev's keep_result vs. ranked by recency.
# Same retained size, so the recall gap is Jev's judgment alone.
"jev_top60k": {
"engine": "jev",
"jev": {"select": "jev", "result_budget_tokens": 60_000},
},
"recent_top60k": {
"engine": "jev",
"jev": {"select": "recency", "result_budget_tokens": 60_000},
},
}

View File

@@ -0,0 +1,89 @@
# fast-jev-compaction vs Hermes compaction — 3-transcript scorecard (2026-09-19)
Question asked: does https://github.com/tamaratran/fast-jev-compaction ("replace the
compaction summary with Jev decisions: score every tool call/result, drop or truncate
stale ones, keep everything else verbatim") beat our compressor on remaining tokens,
compaction cost and recall accuracy?
Harness: `evals/compaction/runner.py` with the new `engine: jev` arm
(`evals/compaction/jev_arm.py`, a Python port of the plugin over OpenRouter's Decisions
API, `~typesafe/jev-latest` → served as `typesafe/jev-1.13-20260917`). Three real 500K-token
lineage prefixes from state.db (PR review campaign, system-prompt token analysis, SIGSEGV
fix), 15-question recall exam each, same bank for every arm, answered and judged by the
configured `auxiliary.compression` route (gemini-3.8-flash via Nous). A fourth transcript
(541 tool calls in 500K) could not be fitted into Jev's 25K-token state ceiling even at the
last fitting stage — the plugin throws there and Claude Code falls back to its built-in
summary; recorded as `jev_fallback`, not scored.
## Results (recall % @ retained tokens; compaction cost and wall time per event)
| policy | prreview | sysprompt | sigsegv | AVG | compaction $ | compaction s |
|---|---|---|---|---|---|---|
| current (main) | 36.7 @ 40K | 50.0 @ 64K | 43.3 @ 61K | **43.3 @ 55K** | $0.061 | 36.9 |
| lean | 33.3 @ 40K | 53.3 @ 65K | 36.7 @ 61K | 41.1 @ 55K | $0.062 | 33.4 |
| jev (plugin defaults) | 70.0 @ 50K | 63.3 @ 112K | 93.3 @ 181K | **75.5 @ 115K** | $0.007 | 1.4 |
| jev_tail40 (40 pinned rows) | 73.3 @ 71K | 63.3 @ 179K | 93.3 @ 213K | 76.6 @ 154K | $0.006 | 1.4 |
| jev_t15 (threshold 0.15) | 93.3 @ 372K | 76.7 @ 332K | 100.0 @ 484K | 90.0 @ 396K | $0.007 | 1.4 |
| jev_top60k (Jev-ranked, 60K tool budget) | 70.0 @ 113K | 70.0 @ 173K | 93.3 @ 243K | 77.8 @ 176K | $0.007 | 1.5 |
| recent_top60k (recency-ranked, same budget) | 76.7 @ 114K | 63.3 @ 174K | 93.3 @ 245K | 77.8 @ 177K | $0 | 0.0 |
`current` and `lean` are the same code path on today's main (lean tail is the default), so
their 2–7 pt spread on identical context is the exam noise floor (15 questions ≈ ±3.3 pts).
Per-question paired comparison, jev vs current across 45 questions: 17 wins, 1 loss, 27 ties.
## Findings
1. **Jev's default arm is +32 pts recall (75.5 vs 43.3) at 2.1× the retained tokens
(115K vs 55K), for 1/9 the compaction cost ($0.007 vs $0.061) in 1/25 the time
(1.4 s vs 37 s).** On the one transcript where the sizes are comparable (prreview,
50K vs 40K) it still wins 70.0 vs 36.7.
2. **The recall gain is verbatim text, not Jev's judgment.** At the plugin's 0.5 threshold
Jev's `keep_result` never exceeded 0.20 (median 0.15) and `keep_call` topped out at 0.50,
so it dropped 100% of the 851 unpinned candidates across all three transcripts
(kept=0, result-truncated=0). The default `jev` arm is therefore behaviourally identical
to "delete every old tool call + result, keep every user/assistant row verbatim". The
facts the summary loses (delegation ids, root causes, config keys, exact error strings)
sat in assistant text the whole time.
3. **At a matched budget Jev's ranking ties plain recency: 77.8 vs 77.8.** Keeping 60K
tokens of tool pairs ranked by `keep_result` (jev_top60k) vs ranked by position
(recent_top60k) gives the same average; per transcript it is +6.7 / −6.7 / 0, inside the
noise floor. Lowering the threshold to 0.15 (jev_t15) reaches 90% but retains 396K of
500K — that is not compaction.
4. **The state ceiling does not fit Hermes scale.** Jev's 32K window forces the whole
history into 25K tokens; at 500K every transcript needed the harshest fitting stages
("old calls compacted/merged", "old messages collapsed") and one of four could not fit at
all. The plugin is designed for Claude Code's ~200K compaction point; a 1M-window Hermes
session compacting at 500K+ will fall back to the summary regularly, and once the tool
results are gone a second compaction has nothing left to remove.
5. **Cost shape.** Jev: 5–7 requests per compaction, ~125–200K input tokens total at
$0.042/M ≈ $0.005–0.008, ~1.5 s wall. Our summary: one gemini-3.8-flash call over
52–60K input tokens ≈ $0.06, 25–49 s. Both are noise against the per-turn cost of the
retained context that follows (115K vs 55K tokens on every subsequent turn).
## What this suggests for the compressor
The cheap win is not a new decision model but a retention rule the data supports: **keep
user/assistant text verbatim and drop/truncate old tool results before anything is
summarised.** We already have that layer (`_prune_old_tool_results`, phase 1); today it is
followed by a summary that also rewrites assistant text, and that rewrite is where the
recall goes. A "prune-only until the tail budget is reached, summarise only the remainder"
posture would capture most of Jev's gain at zero extra cost. If a scoring model is wanted
for the prune ranking, Jev is fast and cheap enough (1.5 s, < 1¢) but this data shows no
signal over recency at equal budget; re-test before wiring it in.
## Method notes
- Transcripts reconstructed with `scripts/reconstruct_lineage.py` from a state.db copy;
not committed. Question banks generated from the region current compaction summarises
(the most conservative boundary) and cached per transcript+cap so every arm answers the
identical exam.
- The `jev` arm counts rows (Hermes has one `role: tool` row per result), so
`preserve_recent_messages: 6` pins fewer turns than in Claude Code; `jev_tail40` widens
it to roughly lean's 25K tail and changes nothing (+1 pt).
- Eval spend for the whole run (question generation, 634 answer/judge calls): 43.5M input
tokens ≈ $33.6 at gemini-3.8-flash list, through Nous inference. Jev spend across all
arms: $0.08 via OpenRouter.