Commit Graph

6 Commits

Author SHA1 Message Date
teknium1
b980495847 refactor: move live-model benchmark harnesses from scripts/ into evals/
scripts/ is for repo tooling (tests runner, release, installers, CI
checks). The tool-search live tests, the toolperf A/B eval and the browser
eval benchmark are offline benchmarks that spend model budget, which is
exactly what evals/ holds; evals/browser_use already cited
scripts/toolperf_abeval as "the same pattern".

- scripts/tool_search_livetest*.py + analyze_livetest.py + LIVETEST_README
  -> evals/tool_search/ (README.md); repo-root sys.path hop adjusted for
  the extra directory level; gitignore now covers evals/tool_search/out*/
- scripts/toolperf_abeval/ -> evals/toolperf_abeval/
- scripts/benchmark_browser_eval.py -> evals/browser_use/
2026-09-13 06:06:46 -07:00
Lakshya Agarwal
428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium
d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Teknium
dfff42d3dd refactor(browser_exec): schema diet 803->663 tok/call, A/B-gated; eval harness runs on Windows + Nous auth (#96300) 2026-08-27 04:39:19 -07:00
Teknium
21aa327c69 style: ruff lint + format on the eval scripts (encoding args, formatting) 2026-08-17 15:29:16 -07:00
Teknium
5033aedaa3 feat(evals): add the Browser Use mode A/B benchmark from PR #81958
Reconstructs the 204-run benchmark battery behind #81958 (Browser Use CLI
3.0 mode) as a rerunnable eval under evals/browser-use/, following the
toolperf_abeval / evals-compaction pattern.

- tasks/easy.json + tasks/hard.json: the oracle-checked toscrape task
  batteries (5 easy, 6 hard) exactly as run for the PR
- single_run.py: one cell = task x arm (base | pr | prns) x model x rep;
  throwaway HERMES_HOME, web-fetch creds stripped, arms pinned to separate
  trees via BUBENCH_BASE_TREE / BUBENCH_PR_TREE
- orchestrate.py: resume-safe local-CDP battery driver
- orchestrate_cloud.py: backend matrix (nous-cloud via the browser_use
  provider plugin, browserbase via REST) with per-cell session lifecycle
- report.py: scorecard aggregation with vs-base token deltas
- README.md: design, run instructions, and the recovered Aug 8-10 2026
  baseline scorecards (hard battery, backend matrix, easy round 1,
  digest ablation)

The original /tmp/bu-bench workspace was lost to a tmpfs reboot; harness
and readouts were recovered verbatim from the benchmark session's tool-call
history in state.db, with hardcoded paths parameterized. Smoke-verified
live: report.py aggregation, and single-cell runs (pr + base arms) against
a real headless Chrome CDP with sonnet-5 driving browser_exec, oracle pass.
2026-08-17 15:29:16 -07:00