Files
hermes-agent/evals/compaction
Teknium 4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
..

Compaction Eval Harness

Measures what context compaction actually costs in recall, not just tokens.

What it does

  1. Takes a real long transcript (JSON: {"messages": [...]}, chat format).
  2. Generates a bank of factual recall questions from the region that compaction will summarize away (cached per transcript for reproducibility).
  3. Runs the transcript through ContextCompressor.compress() under each policy in the matrix (current default, aggressive tail, codex-style, ...).
  4. For each policy, asks a fresh LLM the recall questions with ONLY the post-compaction context, and judges answers against gold.
  5. Emits a scorecard: recall accuracy vs tokens retained, per policy.

Usage

# from repo root, venv active
python evals/compaction/runner.py \
    --transcript /path/to/lineage.json \
    --policies current,aggressive,floor10k \
    --questions 15 \
    --out evals/compaction/results/run1
python evals/compaction/report.py evals/compaction/results/run1

Transcripts are NOT committed (they contain real session data). Point --transcript at a local file. See fixtures.py for the expected shape and a synthetic-transcript generator used by CI smoke tests.

Building transcripts from real sessions (scripts/)

Compaction rotations mean a single active session rarely exceeds ~300K tokens, but the lineage (parent→children chain) carries the full uncompacted history. The scripts reconstruct those into eval transcripts:

# 1. ALWAYS copy the DB first — never point at the live state.db
cp ~/.hermes/state.db /tmp/state_copy.db

# 2. Find big lineages (sessions with parent_session_id form chains), then:
python evals/compaction/scripts/reconstruct_lineage.py \
    /tmp/state_copy.db <root_session_id> /tmp/lineage.json

# 3. (optional) Replay a 500K prefix through one checkout's compressor and
#    dump before/after for the HTML viewer:
python evals/compaction/scripts/replay_lineage.py <checkout> /tmp/lineage.json out.json 500000
python evals/compaction/scripts/build_html_report.py <runs_dir> report.html

reconstruct_lineage.py walks the whole descendant tree chronologically, dedupes rotation-copied rows by content hash, strips synthetic compaction artifacts (summaries, todo snapshots), and resolves the system prompt through the system_prompts dedup table (sessions only carry a hash). The HTML report renders before/after transcripts side by side with compaction artifacts color-coded.

Region-scoping tripwire

test_region_scoping.py plants sentinels in head/middle/tail and asserts the summarizer's serialized-turns input carries ONLY the middle (compacted) region in both legacy and lean modes. Run it directly or via pytest.

Policies

Defined in policies.py. Each policy maps to ContextCompressor constructor kwargs plus optional attribute overrides applied post-construction (e.g. tail_token_budget). Add new policies there — the runner picks them up by name.

Notes

  • Question generation and judging use agent.auxiliary_client.call_llm (same transport the compressor uses), so the harness needs a configured provider. Costs real tokens: ~(policies x questions) answer calls plus one generation and one judge pass.
  • Accuracy is judged 2/1/0 (correct / partial / wrong); the scorecard reports normalized percent. The judge sees gold answers, the answerer does not.
  • --also-uncompacted adds a control arm that answers from the full original transcript — the recall ceiling.