evals(compaction): matched-budget Jev/recency arms + 2026-09-19 fast-jev scorecard
Adds an eval-only budget selector to the Jev arm (rank tool pairs by Jev keep_result or by recency, keep until a token budget) so Jev's judgment can be separated from the effect of keeping text verbatim, and records the 3-transcript scorecard: default Jev +32 pts recall at 2.1x retained tokens, 1/9 the compaction cost; at matched budget Jev ranking ties recency.
This commit is contained in:
@@ -112,6 +112,13 @@ class JevOptions:
|
||||
max_state_tokens: int = 25_000
|
||||
max_request_tokens: int = 30_000
|
||||
truncate_head_chars: int = 300
|
||||
# Eval-only extension (not in the plugin): instead of thresholding, keep
|
||||
# whole call+result pairs in rank order until `result_budget_tokens` of
|
||||
# tool content is retained; everything else is dropped. `select="jev"`
|
||||
# ranks by Jev's keep_result, `select="recency"` by position and never
|
||||
# calls Jev — the control that tells whether Jev's ranking carries signal.
|
||||
select: Optional[str] = None
|
||||
result_budget_tokens: int = 0
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -475,6 +482,36 @@ class JevCompactor:
|
||||
|
||||
return {c.id: {"keep_call": noul(f"call_{c.id}"), "keep_result": noul(f"result_{c.id}")} for c in batch}
|
||||
|
||||
def _budget_decisions(self, messages, calls, candidates, answers) -> List[Decision]:
|
||||
"""Keep ranked call+result pairs until the tool-content budget is spent; drop the rest."""
|
||||
opt = self.opt
|
||||
|
||||
def pair_tokens(c: ToolCall) -> int:
|
||||
return (c.result_chars + len(_json(c.input))) // 4
|
||||
|
||||
if opt.select == "jev":
|
||||
ranked = sorted(candidates, key=lambda c: -answers.get(c.id, {}).get("keep_result", 0.0))
|
||||
else:
|
||||
ranked = sorted(candidates, key=lambda c: -c.result_index)
|
||||
kept, spent = set(), 0
|
||||
for c in ranked:
|
||||
t = pair_tokens(c)
|
||||
if spent + t > opt.result_budget_tokens:
|
||||
continue
|
||||
kept.add(c.id)
|
||||
spent += t
|
||||
out = []
|
||||
for c in calls:
|
||||
a = answers.get(c.id, {})
|
||||
base = dict(id=c.id, tool=c.tool, keep_call=a.get("keep_call", 1.0), keep_result=a.get("keep_result", 1.0))
|
||||
if c.pinned:
|
||||
out.append(Decision(**base, action="keep", reason="pinned"))
|
||||
elif c.id in kept:
|
||||
out.append(Decision(**base, action="keep", reason="kept"))
|
||||
else:
|
||||
out.append(Decision(**base, action="drop_call", reason="call_dropped"))
|
||||
return out
|
||||
|
||||
def compress(self, messages: List[Dict[str, Any]], current_tokens: int = 0, force: bool = True):
|
||||
t0 = time.time()
|
||||
opt = self.opt
|
||||
@@ -483,16 +520,19 @@ class JevCompactor:
|
||||
answers: Dict[str, Dict[str, float]] = {}
|
||||
fitted = {"tokens": 0, "stage": ""}
|
||||
batches: List[List[ToolCall]] = []
|
||||
if candidates:
|
||||
if candidates and opt.select != "recency":
|
||||
fitted = fit_state(messages, calls, opt)
|
||||
batches = batch_calls(candidates, fitted["tokens"], opt)
|
||||
with ThreadPoolExecutor(max_workers=self.concurrency) as pool:
|
||||
for result in pool.map(lambda b: self._ask_batch(fitted["state"], b), batches):
|
||||
answers.update(result)
|
||||
self.decisions = [
|
||||
decide_call(c, *(answers.get(c.id, {}).get(k, 1.0) for k in ("keep_call", "keep_result")), opt)
|
||||
for c in calls
|
||||
]
|
||||
if opt.select:
|
||||
self.decisions = self._budget_decisions(messages, calls, candidates, answers)
|
||||
else:
|
||||
self.decisions = [
|
||||
decide_call(c, *(answers.get(c.id, {}).get(k, 1.0) for k in ("keep_call", "keep_result")), opt)
|
||||
for c in calls
|
||||
]
|
||||
kept = apply_decisions(messages, self.decisions, calls, opt.truncate_head_chars)
|
||||
reasons = [d.reason for d in self.decisions]
|
||||
self.stats = {
|
||||
|
||||
@@ -65,6 +65,17 @@ POLICIES: Dict[str, Dict[str, Any]] = {
|
||||
"engine": "jev",
|
||||
"jev": {"keep_threshold": 0.15},
|
||||
},
|
||||
# Matched-budget pair (eval-only extension): keep 60K tokens of tool
|
||||
# call+result pairs ranked by Jev's keep_result vs. ranked by recency.
|
||||
# Same retained size, so the recall gap is Jev's judgment alone.
|
||||
"jev_top60k": {
|
||||
"engine": "jev",
|
||||
"jev": {"select": "jev", "result_budget_tokens": 60_000},
|
||||
},
|
||||
"recent_top60k": {
|
||||
"engine": "jev",
|
||||
"jev": {"select": "recency", "result_budget_tokens": 60_000},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
|
||||
89
evals/compaction/results/SCORECARD-2026-09-19-jev.md
Normal file
89
evals/compaction/results/SCORECARD-2026-09-19-jev.md
Normal file
@@ -0,0 +1,89 @@
|
||||
# fast-jev-compaction vs Hermes compaction — 3-transcript scorecard (2026-09-19)
|
||||
|
||||
Question asked: does https://github.com/tamaratran/fast-jev-compaction ("replace the
|
||||
compaction summary with Jev decisions: score every tool call/result, drop or truncate
|
||||
stale ones, keep everything else verbatim") beat our compressor on remaining tokens,
|
||||
compaction cost and recall accuracy?
|
||||
|
||||
Harness: `evals/compaction/runner.py` with the new `engine: jev` arm
|
||||
(`evals/compaction/jev_arm.py`, a Python port of the plugin over OpenRouter's Decisions
|
||||
API, `~typesafe/jev-latest` → served as `typesafe/jev-1.13-20260917`). Three real 500K-token
|
||||
lineage prefixes from state.db (PR review campaign, system-prompt token analysis, SIGSEGV
|
||||
fix), 15-question recall exam each, same bank for every arm, answered and judged by the
|
||||
configured `auxiliary.compression` route (gemini-3.8-flash via Nous). A fourth transcript
|
||||
(541 tool calls in 500K) could not be fitted into Jev's 25K-token state ceiling even at the
|
||||
last fitting stage — the plugin throws there and Claude Code falls back to its built-in
|
||||
summary; recorded as `jev_fallback`, not scored.
|
||||
|
||||
## Results (recall % @ retained tokens; compaction cost and wall time per event)
|
||||
|
||||
| policy | prreview | sysprompt | sigsegv | AVG | compaction $ | compaction s |
|
||||
|---|---|---|---|---|---|---|
|
||||
| current (main) | 36.7 @ 40K | 50.0 @ 64K | 43.3 @ 61K | **43.3 @ 55K** | $0.061 | 36.9 |
|
||||
| lean | 33.3 @ 40K | 53.3 @ 65K | 36.7 @ 61K | 41.1 @ 55K | $0.062 | 33.4 |
|
||||
| jev (plugin defaults) | 70.0 @ 50K | 63.3 @ 112K | 93.3 @ 181K | **75.5 @ 115K** | $0.007 | 1.4 |
|
||||
| jev_tail40 (40 pinned rows) | 73.3 @ 71K | 63.3 @ 179K | 93.3 @ 213K | 76.6 @ 154K | $0.006 | 1.4 |
|
||||
| jev_t15 (threshold 0.15) | 93.3 @ 372K | 76.7 @ 332K | 100.0 @ 484K | 90.0 @ 396K | $0.007 | 1.4 |
|
||||
| jev_top60k (Jev-ranked, 60K tool budget) | 70.0 @ 113K | 70.0 @ 173K | 93.3 @ 243K | 77.8 @ 176K | $0.007 | 1.5 |
|
||||
| recent_top60k (recency-ranked, same budget) | 76.7 @ 114K | 63.3 @ 174K | 93.3 @ 245K | 77.8 @ 177K | $0 | 0.0 |
|
||||
|
||||
`current` and `lean` are the same code path on today's main (lean tail is the default), so
|
||||
their 2–7 pt spread on identical context is the exam noise floor (15 questions ≈ ±3.3 pts).
|
||||
Per-question paired comparison, jev vs current across 45 questions: 17 wins, 1 loss, 27 ties.
|
||||
|
||||
## Findings
|
||||
|
||||
1. **Jev's default arm is +32 pts recall (75.5 vs 43.3) at 2.1× the retained tokens
|
||||
(115K vs 55K), for 1/9 the compaction cost ($0.007 vs $0.061) in 1/25 the time
|
||||
(1.4 s vs 37 s).** On the one transcript where the sizes are comparable (prreview,
|
||||
50K vs 40K) it still wins 70.0 vs 36.7.
|
||||
|
||||
2. **The recall gain is verbatim text, not Jev's judgment.** At the plugin's 0.5 threshold
|
||||
Jev's `keep_result` never exceeded 0.20 (median 0.15) and `keep_call` topped out at 0.50,
|
||||
so it dropped 100% of the 851 unpinned candidates across all three transcripts
|
||||
(kept=0, result-truncated=0). The default `jev` arm is therefore behaviourally identical
|
||||
to "delete every old tool call + result, keep every user/assistant row verbatim". The
|
||||
facts the summary loses (delegation ids, root causes, config keys, exact error strings)
|
||||
sat in assistant text the whole time.
|
||||
|
||||
3. **At a matched budget Jev's ranking ties plain recency: 77.8 vs 77.8.** Keeping 60K
|
||||
tokens of tool pairs ranked by `keep_result` (jev_top60k) vs ranked by position
|
||||
(recent_top60k) gives the same average; per transcript it is +6.7 / −6.7 / 0, inside the
|
||||
noise floor. Lowering the threshold to 0.15 (jev_t15) reaches 90% but retains 396K of
|
||||
500K — that is not compaction.
|
||||
|
||||
4. **The state ceiling does not fit Hermes scale.** Jev's 32K window forces the whole
|
||||
history into 25K tokens; at 500K every transcript needed the harshest fitting stages
|
||||
("old calls compacted/merged", "old messages collapsed") and one of four could not fit at
|
||||
all. The plugin is designed for Claude Code's ~200K compaction point; a 1M-window Hermes
|
||||
session compacting at 500K+ will fall back to the summary regularly, and once the tool
|
||||
results are gone a second compaction has nothing left to remove.
|
||||
|
||||
5. **Cost shape.** Jev: 5–7 requests per compaction, ~125–200K input tokens total at
|
||||
$0.042/M ≈ $0.005–0.008, ~1.5 s wall. Our summary: one gemini-3.8-flash call over
|
||||
52–60K input tokens ≈ $0.06, 25–49 s. Both are noise against the per-turn cost of the
|
||||
retained context that follows (115K vs 55K tokens on every subsequent turn).
|
||||
|
||||
## What this suggests for the compressor
|
||||
|
||||
The cheap win is not a new decision model but a retention rule the data supports: **keep
|
||||
user/assistant text verbatim and drop/truncate old tool results before anything is
|
||||
summarised.** We already have that layer (`_prune_old_tool_results`, phase 1); today it is
|
||||
followed by a summary that also rewrites assistant text, and that rewrite is where the
|
||||
recall goes. A "prune-only until the tail budget is reached, summarise only the remainder"
|
||||
posture would capture most of Jev's gain at zero extra cost. If a scoring model is wanted
|
||||
for the prune ranking, Jev is fast and cheap enough (1.5 s, < 1¢) but this data shows no
|
||||
signal over recency at equal budget; re-test before wiring it in.
|
||||
|
||||
## Method notes
|
||||
|
||||
- Transcripts reconstructed with `scripts/reconstruct_lineage.py` from a state.db copy;
|
||||
not committed. Question banks generated from the region current compaction summarises
|
||||
(the most conservative boundary) and cached per transcript+cap so every arm answers the
|
||||
identical exam.
|
||||
- The `jev` arm counts rows (Hermes has one `role: tool` row per result), so
|
||||
`preserve_recent_messages: 6` pins fewer turns than in Claude Code; `jev_tail40` widens
|
||||
it to roughly lean's 25K tail and changes nothing (+1 pt).
|
||||
- Eval spend for the whole run (question generation, 634 answer/judge calls): 43.5M input
|
||||
tokens ≈ $33.6 at gemini-3.8-flash list, through Nous inference. Jev spend across all
|
||||
arms: $0.08 via OpenRouter.
|
||||
Reference in New Issue
Block a user