scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".
- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
-> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI
3.0 mode) as a rerunnable eval under evals/browser-use/, following the
toolperf_abeval / evals-compaction pattern.
- tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task
batteries (5 easy, 6 hard) exactly as run for the PR
- single_run.py: one cell = task x arm (base | pr | prns) x model x rep;
throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate
trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE
- orchestrate.py: resume-safe local-CDP battery driver
- orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use
provider plugin, browserbase via REST) with per-cell session lifecycle
- report.py: scorecard aggregation with vs-base token deltas
- README.md: design, run instructions, and the recovered Aug 8-10 2026
baseline scorecards (hard battery, backend matrix, easy round 1,
digest ablation)
The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness
and readouts were recovered verbatim from the benchmark session's tool-call
history in state.db, with hardcoded paths parameterized. Smoke-verified
live: report.py aggregation, and single-cell runs (pr + base arms) against
a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.