AGENTS.md: Project Structure tree reflects the decomposition (run_agent 1.5k not 12k, cli 4.6k not 11k,
hermes_state facade + 21 siblings, web_routers/, evals/, test counts); new "Facade + siblings layout"
section with the sibling families table and the rules that follow (find by topic, patch where production
reads, compat pointers off limits, don't recreate god files); AIAgent/Agent Loop point at agent/turn_*.py
and conversation_loop; CLI dispatch documents _SLASH_DISPATCH + the _handle_<name>_command convention and
"Adding a Slash Command" no longer tells you to add an elif (there is no ladder to add to on either surface).
evals/codebase_navigability/: what the codebase costs an agent, not the CPU.
bench.py ~19k real "locate X" tasks from tests/ imports; tokens (tiktoken o200k) of the defining
file vs the symbol, context-window fit, read windows, siblings, symbol CC
lookup_sim.py paired grep+read simulation over 4k common symbols; tool calls + tokens returned
static_metrics.py LOC split, size distributions, elif/nesting, radon CC/MI, import graph + SCC cycles
runtime_bench.py fresh-interpreter import/CLI/hot-path/collection timings with tree-purity assertion
tests/evals/test_codebase_navigability.py pins the resolver's facade/sibling behaviour.
4.0 KiB
Codebase navigability eval
Measures what a codebase costs an LLM agent to work in, as opposed to what it costs the CPU. Built for the Sep 2026 decomposition (PR #102117) and kept so future refactors are held to the same numbers. Everything here is offline and deterministic; no model calls.
The three questions it answers
-
How much code has to be read to see one definition? (
bench.py) Workload = everyfrom <first-party module> import <Name>intests/, resolved through re-exports to the module that defines the name. That is ~19k real "locate X" tasks nobody hand-picked. Per task: tokens of the defining file, tokens of the symbol itself, overhead = file − symbol, whether the file fits a 32k / 128k context, how many 2,000-lineread_filewindows it spans, how many unrelated top-level siblings sit in the same file, and the symbol's cyclomatic complexity. Tokens are real (tiktokeno200k_base); falls back to bytes/4 if tiktoken is missing and says so. -
What does a careful agent actually pay to look a symbol up? (
lookup_sim.py) Simulates the policy a good model follows with our tools:grep -nfor the definition (1 call),read_filea 200-line window around the hit (1 call), page forward in 2,000-line windows only while the definition is still running. Charges tool calls and returned tokens. Same random sample of symbols that exist on BOTH trees, so it is a paired comparison. -
Shape and runtime. (
static_metrics.py,runtime_bench.py) LOC split (code/comment/docstring), file and function size distributions, elif-chain lengths, nesting depth, radon CC/MI, import graph (edges, fan-in/out, Tarjan SCC cycles); fresh-interpreter import time / module count / RSS for the entry points, CLI end-to-end, in-process hot paths, pytest collection, bytecode footprint.runtime_bench.pypinsPYTHONPATHto the tree and asserts no module resolved from another checkout (an editable install will silently cross-contaminate otherwise).
Usage
# deps: tiktoken + radon (bench venv or the project venv)
uv pip install tiktoken radon
# 1 + 2: pass two checkouts (git worktree add is the easy way to get the baseline)
git worktree add /tmp/base origin/main
python evals/codebase_navigability/bench.py /tmp/base base --out out/
python evals/codebase_navigability/bench.py . head --out out/
python evals/codebase_navigability/lookup_sim.py /tmp/base . --sample 4000 --out out/
# 3
NAV_OUT=out/ python evals/codebase_navigability/static_metrics.py /tmp/base base
NAV_OUT=out/ python evals/codebase_navigability/static_metrics.py . head
NAV_OUT=out/ python evals/codebase_navigability/runtime_bench.py /tmp/base base 9
NAV_OUT=out/ python evals/codebase_navigability/runtime_bench.py . head 9
bench.py and static_metrics.py take ~2 min each on a 1M-line tree; lookup_sim.py ~10 min for
4,000 symbols (it tokenizes every window it "reads"); runtime_bench.py ~4 min per tree at 9 reps.
Reading the results honestly
bench.pymeasures the naive cost (read the whole defining file). It is the number that drops ~3× when god files are split, and the one that decides whether a file fits a context window at all.lookup_sim.pymeasures the skilled cost. It barely moves with a file split, because grep + a 200-line window already dodges file size. What moves it is (a) the definition itself getting shorter and (b) token density of the surrounding code. Stripping comments makes each line denser, so a fixed 200-line window costs more tokens after a comment-stripping refactor even though the code is smaller. Both effects are real and pull in opposite directions; report both numbers, not the flattering one.- Import time goes up with a split in Python (per-module overhead), and the import graph's largest cycle typically grows (intra-file coupling becomes inter-module edges). Neither is hidden by these tools.
Results for PR #102117 are in the PR description; raw JSON for that run lives in the PR thread.