jev_cycles_report.py renders the cycles/freed/floor/end-state table from
jev_cycles.py outputs; README carries the exact commands, thresholds and
cost so the diminishing-returns and stuck/fallback findings can be
reproduced. Fallback records now count the real tool calls.
jev_cycles.py compacts a lineage with Jev every time the estimate crosses a
threshold and records freed tokens, text floor and fitting stage per cycle.
Six runs show the floor (never-removed user/assistant text) is the binding
limit: freed-per-cycle decays to 0-8% and a 200K-window host is stuck after
0.42M tokens; >~500 tool calls between compactions cannot fit the 32K state.
Scorecard gains a decisive verdict section.
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same
transcripts: chunk-file reads until its REAL auto-compaction fires (verified
via compacted events in the rollout jsonl; peak 455-483K vs its 258K
window), then quizzes post-compaction with the identical question banks and
judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean
closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access
over its rollout history — the session_search differentiator, measured.