#!/usr/bin/env python3 """Agent-navigability benchmark: what does it cost an LLM agent to find and read a symbol? Workload (not hand-picked): every ``from import `` in tests/ on the tree under test, resolved to the module that DEFINES the name on that tree. Each is one "locate and read X" task. Per task, three costs are measured in real tokenizer tokens (tiktoken o200k_base, the GPT-4o/o-series tokenizer; other tokenizers differ by a roughly constant factor): file_tokens tokens to read the whole file that defines X (the naive "read_file" cost) symbol_tokens tokens of X's own definition (irreducible: you must read this) overhead file_tokens - symbol_tokens (what the file's SIZE costs you beyond the answer) fits_128k / fits_32k can the defining file be loaded whole into that context? windows_2k how many 2,000-line read_file windows the file spans (how many reads to scan it) Also: for each defining file, the number of OTHER top-level symbols it contains (how much unrelated code sits next to the answer), and the depth/complexity of the symbol itself. Usage: python evals/codebase_navigability/bench.py