Files
hermes-agent/tests/evals
teknium1 139047ae2a feat(evals): fast-jev-compaction arm + compaction cost metering in the compaction eval
Adds a Python port of tamaratran/fast-jev-compaction as an eval policy
(`engine: jev`): TypeSafe's Jev decision model scores every tool call and
result over the whole history and stale ones are dropped or truncated, with
no summary and no rewriting of user/assistant text. Transport is OpenRouter's
Decisions API (~typesafe/jev-latest). A state that cannot fit the plugin's
25K-token ceiling is recorded as a fallback (the plugin's own behaviour)
rather than scored.

Every arm now reports what its compaction step cost (calls, tokens, USD —
Jev reports cost directly; summary calls are metered through the
compressor's call_llm binding and priced at OpenRouter list), and the run
writes the harness's own answer/judge token bill so the eval's spend is
visible. The question cache key includes --cap-tokens so different caps on
one transcript no longer share a bank.

Why: measure remaining tokens, compaction cost and recall accuracy of the
"decide, don't summarize" approach against current/lean on real Hermes
lineages before deciding whether any of it belongs in the compressor.
2026-09-19 11:58:37 -07:00
..