Files
hermes-agent/evals/codebase_navigability

Codebase navigability eval

Measures what a codebase costs an LLM agent to work in, as opposed to what it costs the CPU. Built for the Sep 2026 decomposition (PR #102117) and kept so future refactors are held to the same numbers. Everything here is offline and deterministic; no model calls.

The three questions it answers

  1. How much code has to be read to see one definition? (bench.py) Workload = every from <first-party module> import <Name> in tests/, resolved through re-exports to the module that defines the name. That is ~19k real "locate X" tasks nobody hand-picked. Per task: tokens of the defining file, tokens of the symbol itself, overhead = file − symbol, whether the file fits a 32k / 128k context, how many 2,000-line read_file windows it spans, how many unrelated top-level siblings sit in the same file, and the symbol's cyclomatic complexity. Tokens are real (tiktoken o200k_base); falls back to bytes/4 if tiktoken is missing and says so.

  2. What does a careful agent actually pay to look a symbol up? (lookup_sim.py) Simulates the policy a good model follows with our tools: grep -n for the definition (1 call), read_file a 200-line window around the hit (1 call), page forward in 2,000-line windows only while the definition is still running. Charges tool calls and returned tokens. Same random sample of symbols that exist on BOTH trees, so it is a paired comparison.

  3. Shape and runtime. (static_metrics.py, runtime_bench.py) LOC split (code/comment/docstring), file and function size distributions, elif-chain lengths, nesting depth, radon CC/MI, import graph (edges, fan-in/out, Tarjan SCC cycles); fresh-interpreter import time / module count / RSS for the entry points, CLI end-to-end, in-process hot paths, pytest collection, bytecode footprint. runtime_bench.py pins PYTHONPATH to the tree and asserts no module resolved from another checkout (an editable install will silently cross-contaminate otherwise).

Usage

# Independent benchmark environment only — never Hermes's selected environment.
uv venv /tmp/hermes-navigability-bench
source /tmp/hermes-navigability-bench/bin/activate
uv pip install tiktoken radon

# 1 + 2: pass two checkouts (git worktree add is the easy way to get the baseline)
git worktree add /tmp/base origin/main
python evals/codebase_navigability/bench.py /tmp/base base --out out/
python evals/codebase_navigability/bench.py .        head --out out/
python evals/codebase_navigability/lookup_sim.py /tmp/base . --sample 4000 --out out/

# 3
NAV_OUT=out/ python evals/codebase_navigability/static_metrics.py /tmp/base base
NAV_OUT=out/ python evals/codebase_navigability/static_metrics.py .        head
NAV_OUT=out/ python evals/codebase_navigability/runtime_bench.py  /tmp/base base 9
NAV_OUT=out/ python evals/codebase_navigability/runtime_bench.py  .        head 9

Use a fresh benchmark path; do not replace an existing environment. Runtime and pytest-collection measurements also require the target tree's application/test dependencies. Prepare those in a separate caller-owned output with python -m pm.build_env --source <tree> --out <fresh-output> --group dev --group test from a PM-prepared checkout, rather than injecting benchmark packages into Hermes.

bench.py and static_metrics.py take ~2 min each on a 1M-line tree; lookup_sim.py ~10 min for 4,000 symbols (it tokenizes every window it "reads"); runtime_bench.py ~4 min per tree at 9 reps.

Reading the results honestly

  • bench.py measures the naive cost (read the whole defining file). It is the number that drops ~3× when god files are split, and the one that decides whether a file fits a context window at all.
  • lookup_sim.py measures the skilled cost. It barely moves with a file split, because grep + a 200-line window already dodges file size. What moves it is (a) the definition itself getting shorter and (b) token density of the surrounding code. Stripping comments makes each line denser, so a fixed 200-line window costs more tokens after a comment-stripping refactor even though the code is smaller. Both effects are real and pull in opposite directions; report both numbers, not the flattering one.
  • Import time goes up with a split in Python (per-module overhead), and the import graph's largest cycle typically grows (intra-file coupling becomes inter-module edges). Neither is hidden by these tools.

Results for PR #102117 are in the PR description; raw JSON for that run lives in the PR thread.