Commit Graph

21 Commits

Author SHA1 Message Date
Teknium
6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Teknium
4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
Teknium
5bbb4cd6a9 lint: explicit encoding on harness file opens (ruff unspecified-encoding) 2026-08-29 19:04:29 -07:00
Teknium
037ce5cf75 eval(tool-search): check in the core-tool-deferral live A/B harness
The harness behind this PR's 288-run maintainer battery, ported from /tmp
into evals/ alongside the readtool / session_search_schema harnesses.
Arm trees parametrized via ABDEFER_BASE_TREE / ABDEFER_PR_TREE (pinned
plain checkouts), results/python roots via ABDEFER_RESULTS / ABDEFER_PYTHON.

- tasks.py: 14 tasks — per-deferred-tool coverage, multistep, long-range,
  clarify ambiguity trap, eager-only control, false-discovery distractor;
  programmatic graders with partial credit
- worker.py: isolated per-cell subprocess, temp HERMES_HOME, hermetic env,
  seeded session DB with decoys, deterministic desktop/computer_use/
  image_generate stubs, interactivity-fairness continuation, exit-3
  infra-abort (misconfig never scores)
- orchestrator.py: resume-safe battery runner, wall timeouts, retry of
  errored records, infra-abort fuse
- report.py: per-task A/B tables + mean-of-task-means
- results/SUMMARY.md: the shipped verdict; rep JSONs gitignored

Ported worker re-verified live post-port (real terra cell, score 1.0).
2026-08-29 18:42:59 -07:00
Teknium
dfff42d3dd refactor(browser_exec): schema diet 803->663 tok/call, A/B-gated; eval harness runs on Windows + Nous auth (#96300) 2026-08-27 04:39:19 -07:00
Teknium
b2bd1ac63f eval: session_search schema A/B harness + PR #95570 reference results
Live tool-use A/B for session_search schema changes: arms are git refs
(tools/session_search_tool.py extracted per ref), tasks run a minimal
agent loop over OpenRouter against a freshly seeded temp session DB with
programmatic oracles — discovery, forced forward-scroll, AND-miss
broadening, verbatim link emission, profile-link resolution, browse.

Checked-in results/pr95570/ holds the 108-run battery (3 models, 3 reps,
2 arms) that validated the PR #95570 schema diet before merge:
base 49/54 vs diet 52/54, avg tokens/task -25%.
2026-08-26 08:24:21 -07:00
Teknium
21aa327c69 style: ruff lint + format on the eval scripts (encoding args, formatting) 2026-08-17 15:29:16 -07:00
Teknium
5033aedaa3 feat(evals): add the Browser Use mode A/B benchmark from PR #81958
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI
3.0 mode) as a rerunnable eval under evals/browser-use/, following the
toolperf_abeval / evals-compaction pattern.

- tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task
  batteries (5 easy, 6 hard) exactly as run for the PR
- single_run.py: one cell = task x arm (base | pr | prns) x model x rep;
  throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate
  trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE
- orchestrate.py: resume-safe local-CDP battery driver
- orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use
  provider plugin, browserbase via REST) with per-cell session lifecycle
- report.py: scorecard aggregation with vs-base token deltas
- README.md: design, run instructions, and the recovered Aug 8-10 2026
  baseline scorecards (hard battery, backend matrix, easy round 1,
  digest ablation)

The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness
and readouts were recovered verbatim from the benchmark session's tool-call
history in state.db, with hardcoded paths parameterized. Smoke-verified
live: report.py aggregation, and single-cell runs (pr + base arms) against
a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.
2026-08-17 15:29:16 -07:00
Teknium
0ee5ae61e2 docs(evals): real Codex CLI head-to-head arm + results
scripts/codex_arm.py drives OpenAI Codex CLI end-to-end on the same
transcripts: chunk-file reads until its REAL auto-compaction fires (verified
via compacted events in the rollout jsonl; peak 455-483K vs its 258K
window), then quizzes post-compaction with the identical question banks and
judge. Results (results/codex-arm-2026-08-15/): codex 36.7% avg vs lean
closed-book 40.0% vs lean+recovery 68.3%. Codex has no runtime re-access
over its rollout history — the session_search differentiator, measured.
2026-08-15 18:41:24 -07:00
Teknium
e2990428e7 docs(evals): ship transcript-building scripts + full eval detail in-repo
- scripts/reconstruct_lineage.py: rebuild full uncompacted lineage
  transcripts from a state.db COPY (descendant-tree walk, content-hash
  dedupe, synthetic-artifact strip, system_prompts hash resolution)
- scripts/replay_lineage.py + scripts/build_html_report.py: replay a 500K
  prefix through any checkout's compressor and render before/after
  side-by-side with compaction artifacts color-coded
- README: transcript-building workflow, scoping-tripwire section
- results/SCORECARD-2026-08-15.md: full per-transcript scorecards, all 4
  exam question banks, survival analysis, methodology + caveats
2026-08-15 17:40:43 -07:00
Teknium
9146f4c851 fix(evals): explicit encoding on all file I/O (ruff PLW1514) 2026-08-15 17:07:46 -07:00
Teknium
44536bd41b docs(evals): 4-transcript compaction scorecard
lean+recovery 68.3% avg recall @ 49K retained vs current 45.8% @ 162K —
+22.5pts at 0.30x tokens. Anchor index moved GUI needle-fact recall
23.3->60.0 closed-book, 46.7->80.0 with recovery.
2026-08-15 17:01:22 -07:00
Teknium
c4bbb14e52 feat(compression): mechanical anchor index + region-scoping tripwire
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
  file paths, error strings, handles, URLs from the compacted region into a
  bounded indexed summary section. LLM-free, so needle identifiers cannot be
  paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
  golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
  summarizer input carries ONLY the compacted region (head/tail sentinels
  never reach the serialized turns body) in both legacy and lean modes.
2026-08-15 17:01:22 -07:00
Teknium
7a82457ede feat(compression): digest noise filter + FTS5 recovery sim + digest-aware query hints
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
  digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
  session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
  mine anchor identifiers
2026-08-15 17:01:22 -07:00
Teknium
8fe9025abd feat(compression): lean tail mode + recovery-aware eval arm
Lean mode (tail_mode='lean', default stays 'legacy'):
- tail budget = clamp(2.5% of window, 10K, 25K) instead of 0.20*window
- stale tail tool results demoted to session_search recovery stubs
- chunked identifier-preserving digests of the compacted region (map-reduce,
  pristine pre-prune tool contents)
- verbatim user messages embedded in summary (codex retention-by-role rule)
- deterministic session_search recovery footer

Eval: policies matrix gains lean + a '+recovery' arm giving the answerer one
simulated session_search round-trip against the archived region.
2026-08-15 17:01:22 -07:00
Teknium
33242d5ee0 feat(evals): compaction recall eval harness
Measures recall accuracy vs tokens retained across compaction policies.
Real transcripts in, LLM-generated recall exam from the summarized region,
per-policy answer+judge passes, scorecard out.
2026-08-15 17:01:21 -07:00
Teknium
fd452e26e3 feat(tools): unicode-equivalent filename retry + near-miss suggestions in read_file
NFC/NFD, narrow no-break space (U+202F), and curly quotes render
identically in a terminal — a model retyping a visually-correct path
gets 'file not found' and can never discover the byte mismatch on its
own. On not-found, canonicalize the requested name and compare against
directory entries; exactly ONE equivalent spelling reads transparently
with an explanatory note. Zero or several matches (homoglyph twins)
fall through — never guess between collisions.

Also: difflib.SequenceMatcher >=0.8 fallback in _suggest_similar_files
catches near-miss typos (AGENT.md -> AGENTS.md) that substring scoring
misses entirely.

Measured (file-only arm, 3 reps, control=guard-only vs feature):
unicode task qwen3.8-max 31k->16k tok (-48%), turns 6.7->3.7;
opus-4.8 57k->33k tok (-42%), turns 8.3->5.0; accuracy held 1.00.
near-miss: opus mildly better, qwen flat, no regressions.
2026-08-09 23:35:07 -07:00
teknium1
ddd21abfa8 chore(evals): track results/.gitignore (its own * rule excluded it from the original add) 2026-08-09 23:30:02 -07:00
Teknium
0e63ed1feb feat(tools): stat-based special-file guard for read_file + readtool eval harness
read_file on a workspace FIFO/socket blocked until the exec timeout —
the existing device guard is name-based (/dev/*, /proc/*) and cannot
see an arbitrary special file. Add _special_file_kind(): one os.stat
on the resolved path, refusing FIFO/socket/char/block devices with a
plain note ('no read was attempted') instead of hanging. Host-visible
filesystems only; regular files, dirs, and missing paths unchanged.

Also adds evals/readtool/: an A/B harness that runs the real AIAgent
against hostile-file fixtures (huge lockfile, one-line bundle, FIFO,
NFD filenames, lying extensions) and measures accuracy, turns, tool
calls, and tokens. Measured for this guard (3 reps, file-only arm):
qwen3.8-max fifo task tokens 122k -> 26k (-79%), turns 9.3 -> 5.0;
opus-4.8 tokens 40k -> 23k; accuracy held 1.00 both arms.
2026-08-09 23:30:02 -07:00
teknium
ba3fea24f1 Enhance TerminalBench 2 configuration and evaluation handling
- Added task_timeout parameter to enforce a maximum wall-clock time for each task, automatically scoring as FAIL if exceeded.
- Introduced terminal_timeout and tool_pool_size parameters to improve command execution and concurrency management.
- Updated logging to provide detailed task execution times and timeout handling, enhancing overall monitoring.
- Removed outdated evaluate_config.yaml file to streamline configuration management.
2026-02-10 22:53:24 +00:00
teknium
35ad3146a8 Add new environments and enhance tool context functionality
- Introduced new environments: Terminal Test Environment and SWE Environment, each with default configurations for testing and software engineering tasks.
- Added TerminalBench 2.0 evaluation environment with comprehensive setup for agentic LLMs, including task execution and verification.
- Enhanced ToolContext with methods for uploading and downloading files, ensuring binary-safe operations.
- Updated documentation across environments to reflect new features and usage instructions.
- Refactored existing environment configurations for consistency and clarity.
2026-02-10 19:39:05 +00:00