Activation reaches plugin discovery before the application dependencies exist. Give PM its own locked Python project and runtime so it can install or repair the application without importing that dependency tree. Keep PM outside the application workspace. A shared uv workspace resolves the application graph and cannot provide this isolation. Route mutations through an isolated worker and preserve transaction callbacks, cancellation, custom package registrations, and correlated receipts. Use the same runtime builder for source installs and packaged payloads. Keep offline wheelhouse support in that builder. Nix builds the independent PM lock as a separate derivation. Refuse lazy-disabled bootstrap before installing tools or dependencies. Move first-party YAML readers and writers to ruamel. Keep the application lock's transitive PyYAML requirements for third-party packages. Verification: - Focused canonical Python suite: 177 passed, 1 host-gated skip. - Electron backend probes: 12 passed. Electron typecheck passed. - Both uv locks, scoped lint, Bash syntax, and whitespace checks passed. - Cold activation, corrupt-app repair, offline staging, and relocation ran. - Built and exercised the Nix PM runtime and standalone YAML merge script. Six broader caller test files retain the same 24 failing test IDs as an archive of HEAD. The existing real-home guard blocks those tests before they can exercise the affected paths. No full-suite pass is claimed. Native Windows signing and full Bionic package execution remain unverified.
Browser Use Mode Benchmark
The A/B battery behind PR #81958
(Browser Use CLI 3.0 mode, salvage of #66476 by @laithrw): built-in
browser_* toolset vs the single browser_exec driver, measured as total
task tokens / tool calls / wall clock at accuracy parity on live multi-step
web tasks.
Design
- Arms differ only by tree + config.
baseruns the built-in twelvebrowser_*tools from a merge-base checkout;prrunsbrowser_exec(browser.backend: browser-use) from the branch checkout;prnsisprwith the schema's helpers digest stripped to the header (isolates the digest's value). Each cell gets a throwawayHERMES_HOME; web-fetch credentials are stripped so every arm must actually drive the browser. - Tasks are oracle-checked. toscrape-family sites (stable content, no
anti-bot), regex oracles over the final answer.
tasks/easy.json(5 tasks: price lookup, category extract, count/aggregate, login, pagination) andtasks/hard.json(6 tasks: full-category multi-page crawls, five-star rating aggregation, JS/delayed render, login chain, cross-category compare). - Resume-safe. Completed cells in
results/*.jsonlare skipped on rerun (same pattern asscripts/toolperf_abeval). - Backend matrix.
orchestrate.pydrives a local headless-Chrome CDP;orchestrate_cloud.py --backend nous-cloud|browserbaseprovisions a real cloud browser per cell through the same provider plumbing the product uses.
Run
# arms are pinned checkouts — e.g. merge-base worktree vs your branch
export BUBENCH_BASE_TREE=/path/to/merge-base-tree
export BUBENCH_PR_TREE=/path/to/branch-tree
Note: since #81958 merged (and #85170 made Browser Use the default driver),
a current-main checkout resolves to browser_exec in BOTH arms. The base
arm only measures the built-in browser_* toolset when BUBENCH_BASE_TREE
is pinned to a pre-#81958 tree (the original run used the PR's merge-base
worktree). For future A/Bs of new browser changes, pin base to the
merge-base of the change under test — the arms are generic.
google-chrome --headless=new --remote-debugging-port=9333 \
--user-data-dir=/tmp/bubench-chrome --no-first-run --disable-gpu about:blank &
python3 orchestrate.py --tasks tasks/hard.json --reps 3 # 108 cells @ 2 models x 3 arms
python3 report.py results/results.jsonl
Baseline scorecard (Aug 8-10 2026, the #81958 run — 204 cells total)
Hard-task battery, local Chrome CDP (6 tasks x 3 reps per cell; final corrected-oracle readout, nothing excluded):
model arm ok tok_mean tok_med calls wall_s vs base tok
opus4.8 base 18/18 64594 63776 4.1 25.2 —
opus4.8 pr 18/18 25934 25030 2.0 17.5 -60%
opus4.8 prns 18/18 25578 27934 3.2 23.7 -60%
kimi-k3 base 18/18 56464 53276 5.3 50.0 —
kimi-k3 pr 18/18 19230 16710 2.4 33.3 -66%
kimi-k3 prns 18/18 23099 21160 4.1 50.5 -59%
Digest ablation: pr (with helpers digest) 36/36 ok, mean 22,582 tok; prns (header-only) 36/36 ok, mean 24,339 tok — the pinned 3.4KB digest costs nothing and saves a little; the full 11KB live skill dump adds nothing.
Backend matrix (pr arm, same tasks):
model backend ok tok_mean calls wall
opus4.8 local-cdp 17/18 25934 2.0 17.5
opus4.8 nous-cloud 12/12 33330 2.8 33.8
opus4.8 browserbase 6/6 26712 2.2 23.2
kimi-k3 local-cdp 18/18 19230 2.4 33.3
kimi-k3 nous-cloud 12/12 22050 2.9 41.4
kimi-k3 browserbase 6/6 22121 2.8 35.2
Easy battery, round 1 (5 tasks x 3 reps, sonnet-5 + qwen3-coder-30b; after excluding provider-noise runs — raw chat-template XML, 0 tool calls):
model arm ok prompt compl total calls wall_s
claude-sonnet-5 base 15/15 39771 324 40095 2.7 16.5
claude-sonnet-5 pr 15/15 27482 509 27991 2.4 14.3
qwen3-coder-30b base 13/14 59509 559 60068 5.7 21.5
qwen3-coder-30b pr 10/11 57146 1616 58763 6.8 26.3
sonnet-5: −30% tokens at parity. qwen3-30b: a wash — weak coders burn the savings retrying exec code. The token win concentrates on multi-step tasks and grows with task hardness; strong models also finish in fewer tool calls.
Compatibility probes from the same run: Firecrawl cloud browsers attach fine (CDP websocket); Camofox has no CDP surface — structurally incompatible, hence the automatic fallback to the built-in toolset in #81958.
Caveats: toscrape-family sites (no anti-bot, no heavy SPA); n<=3 per cell; success-rate deltas at this n are noise — audit sub-100% cells run-by-run before calling a regression.
Provenance
The original per-run results*.jsonl files lived in /tmp/bu-bench/ (tmpfs)
and were lost in a host reboot on Aug 12 2026. The harness, task definitions,
and aggregate readouts in this directory were recovered verbatim from the
session transcripts of the benchmark run (session 20260808_050008_5f615e
tool-call history); single_run.py/orchestrate*.py are the recovered
scripts with the hardcoded /tmp/bu-bench paths parameterized. Rerunning the
battery reproduces fresh per-run data.