The native vision_analyze fast path (and the browser screenshot twins) downscaled every
embed to a fixed _EMBED_TARGET_BYTES = 256 KB. That is fine for photos but turns a
1080x2340 phone screenshot of a table into 540x1170, which the model then reads as
"unreadable" and re-requests (#112095).
The budget is now `vision.embed_target_bytes` in config.yaml (default unchanged: 256 KB,
clamped to 64 KiB..4 MiB so one setting cannot make every later request a multi-megabyte
resend), resolved in the new topical sibling tools/vision_tools_history_budget.py and read
by vision_analyze, browser_vision and browser_exec screenshots alike.
Ported from #112947 by @MohamadKanso (resolver + clamp), relocated out of the facade.
Both process-global caches were keyed by the caller's session/task id alone, so under
gateway.multiplex_profiles two profiles using the same id — a shared `browser_exec session=`
name, or two Hermes sessions whose screens report the same DISPLAY — resolved to the FIRST
profile's cloud browser / cua-driver, and a command issued in one bot's chat could act on
another bot's screen.
The key now carries the routed profile's home key whenever a served-profile scope is active
(`get_hermes_home_override()` set), the same shape `tools/approval.py::_baseline_key` and the
camofox/cloud caches already use; outside a scope every key is byte-identical to before. The
computer_use lookup, install and release paths all go through one `_scoped_sid`, so a release
under profile B never stops profile A's driver; approval-bypass state keeps the bare session id.
Fixes#110032 (report by @wolfyy970, from @vandaimer's manual test on #108914).
browser_exec with browser.cloud_provider: nous (the hermes tools picker row) fell into
the direct-API Browser Use branch and reported chrome-not-running, because
_resolve_backend_cdp gated on _use_gateway(), which only read the pre-picker
use_gateway: true flag. Recognize the picker selection too.
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.
Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.
Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.
Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.
Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.
Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.
Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.
session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.
Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.
Fixes#82903Fixes#65941Fixes#99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
Found by using the feature as a user (natural prompts, real sites, CLI PTY + native Electron), not by naming tools:
- browser_vault_save_login was registered but never offered: toolsets.py is a hand-maintained list. Added, with an
invariant test that every registered browser_vault_* tool is in the browser toolset.
- Vault tools were absent on the DEFAULT backend (Browser Use): the gate deferred to check_browser_requirements(),
which is False by design there. Gate = is_browser_use_cli_mode() or check_browser_requirements().
- The model typed a page-shown demo password with browser_type and offered to take one in chat: the vault rules
lived only on the vault tools. browser_type/browser_exec now carry a vault note when the vault tools are
present ("call browser_vault_list first … never type a password with this tool, never accept one in chat, even
if the page shows it"); the browser_exec login-wall line points at the vault instead of "ask the user".
- On Browser Use the saved item was bound to chrome://new-tab-page: the supervisor's default page session is the
daemon's blank tab. browser_vault_save_login now focuses the tab holding a password field before reading its
origin (focus_page("", accept=probe); about:/chrome: pages are never candidates). Live E2E leg added.
- Settings row: "identifier · Added <date>", origin omitted when it duplicates the label.
Live (real model): CLI on Browser Use — first visit prompts, signs in, saves; second visit fills silently; GitHub
decline (Enter or ESC) stops the agent, which refuses chat passwords. CLI on the built-in stack — same three
scenarios pass. Desktop native Electron — same three scenarios plus Settings list/remove pass. Password never in
a transcript, UI, or a file outside vault/.
On the default backend (browser.backend unset → browser_exec) the vault tools were
advertised but could never fill: the CDP supervisor that carries the secret-bearing
eval is started only by the built-in browser_* session path, so _eval_js_secret
failed closed with supervisor_required and the origin pre-check fell back to an
agent-browser CLI eval against a browser browser_exec never touched.
- browser_exec now attaches SUPERVISOR_REGISTRY to the CDP endpoint it just routed
the harness to (BU_CDP_WS/BU_CDP_URL), so the fill talks to the SAME browser over
the same secret-capable WebSocket. BU direct-cloud (BU_AUTOSPAWN) exposes no
endpoint and keeps the supervisor_required refusal.
- CDPSupervisor.focus_page(origin, accept=<js>) (used by browser_vault_fill, next commits) re-attaches the page session to the
open tab on the item's origin whose DOM holds the form being filled (browser_exec
opens its own tabs; the supervisor's initial attach picks the first page target,
which is chrome://new-tab-page). browser_vault_fill uses it before the origin
pre-check with a per-kind probe (password input / card fields / address fields).
Live: evals/vault_fill_live_e2e.py drives the real browser_exec tool against Hermes'
packaged Chromium with the login page in the third tab; A/B with the attach line
disabled fails at "did not attach a supervisor", enabled fills the password into the
/login tab and card fields into the /checkout tab with every model-facing read
scrubbed.
Also: browser_vault_list/fill described the workflow as "type the identifier with
fill_input", a helper that exists only inside browser_exec code (toolset browser-use) and is
a ghost on the built-in stack. model_tools._rewrite_browser_vault substitutes the concrete
name from the session's actual tool set (`fill_input` inside browser_exec, or browser_type),
the same dynamic cross-reference pattern browser_navigate uses for web_search.
The salvaged fix only ran the CLI in its own session on POSIX and kept
plain subprocess.run on Windows — the one platform where the wedge in
#106244 is actually reproducible: CPython's run() retries an unbounded
communicate() after kill() there, so a grandchild holding the capture
pipes blocks the worker forever. On POSIX run() wait()s the PID and
returns promptly; the grandchild merely leaks (live-reproduced on Linux).
One code path for both platforms:
- _group_popen_kwargs: start_new_session=True on POSIX,
CREATE_NEW_PROCESS_GROUP + hide flags on Windows (replaces the
hide-only _windows_popen_kwargs).
- _kill_cli_process_group: os.killpg SIGKILL on POSIX, taskkill /T /F
on Windows (same kwarg set as the sibling taskkill sites).
- The Popen decodes with encoding="utf-8", errors="replace" like every
other subprocess call in this file (windows footgun rule).
- Drain test patches the kill helper instead of os.killpg so it runs on
every host.
Windows behaviour is not live-verifiable on this Linux host.
The group kill lives behind browser_exec's os.name != "nt" branch, so
os.killpg/SIGKILL are never reached on Windows; mark the line per the
scan's suppression convention.
subprocess.run only kills the direct CLI child on TimeoutExpired; a
browser_harness daemon / Chrome helper grandchild inherits the stdout/
stderr pipes and keeps them open, so the internal communicate() blocks
on pipe EOF forever. The wedged tool call never returns, its activity
heartbeat keeps stamping last_activity_at every 30s, and the session is
pinned at "now" in the desktop sidebar indefinitely (#106244).
On POSIX, run the CLI in its own session (start_new_session=True) and
SIGKILL the whole process group on timeout, then drain the pipes under
a bounded deadline. Windows keeps subprocess.run.
Symptom: with Browser Use mode (the default) and no cloud provider / CDP
override, browser_exec left BU_CDP_* unset, so the browser-use harness ran
its own local discovery — hunting for the user's INSTALLED Chrome on its
default profile. That path needs the chrome://inspect remote-debugging
toggle plus an "Allow remote debugging?" popup per run, is blocked outright
on Chrome >=136, and on a headless host (or one without Chrome) fails with
"chrome-not-running: no supported Chromium-family browser is running". The
built-in browser_* tools never had this problem: they drive the Chromium
Hermes installs, launched through agent-browser.
Change: `_resolve_backend_cdp` step 4 becomes the local ENGINE —
lightpanda when configured, otherwise `_resolve_managed_chromium_cdp`,
which runs `agent-browser --session <key> get cdp-url` through the legacy
`_run_browser_command` (Chromium preflight/auto-install, per-key session
cache, inactivity reaper, atexit) and exports the returned ws:// endpoint
as BU_CDP_WS. Each cache key (task, or bu-named-<session>) gets its own
Chromium, so named sessions are private and skip the own-tab preamble.
Running `get cdp-url` on every call also refreshes agent-browser's idle
timer (it never sees the harness's direct CDP traffic) and follows a
relaunch after the reaper closed the browser.
Live: before, browser_exec on origin/main → exit 1 "chrome-not-running";
after → new_tab/page_info succeed on a cold start, warm reuse, post-reap
relaunch, and a named session; a missing Chromium surfaces the install
hint instead of the harness's Chrome hunt.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
Browser Use mode never read browser.engine: _resolve_backend_cdp() went
BU_CDP_* env -> CDP override -> cloud provider -> local Chrome, so
`engine: lightpanda` was a silent no-op on the default backend, and on
the built-in path it was skipped whenever a cloud provider, Camofox or a
CDP override was active without anyone saying so.
- browser_use_cli: when the engine is lightpanda and nothing with higher
precedence claimed the session, get a session from _get_session_info()
and export its endpoint as BU_CDP_URL; the browser is private to the
session key, so the own-tab preamble is skipped. The browser_exec
description gains a Lightpanda header (text-first, new_tab once then
goto_url — lightpanda-io/browser#1962).
- browser_tool: _create_local_session() spawns `lightpanda serve
--host 127.0.0.1 --port <free>` per session key (new
tools/browser_lightpanda.py), reusing the session cache, inactivity
reaper and atexit cleanup; a dead process is respawned on the next call;
orphans from a crashed Hermes are reaped through per-process records in
$HERMES_HOME/cache/browser-use/lightpanda/. New lightpanda_engine_status()
reports whether the engine is in effect or what shadows it.
- tools_config: "Lightpanda" row in the Browser Automation picker
(cloud_provider: local + engine: lightpanda; "Local Browser" resets the
engine to auto) with a binary-check post-setup.
- /browser status and hermes doctor print the engine state and, when it
is shadowed, the reason.
Copy the user's default-Chromium profile (auth state only) into a managed
snapshot, launch Hermes' packaged Chromium on it via agent-browser, and hand
the CDP endpoint to the Browser Use CLI (and built-in tools) to drive. The
snapshot is a non-default dir, so it sidesteps Chrome 136+'s default-profile
remote-debugging block and never contends with the user's running browser;
launched without mock-keychain switches so keyring-encrypted cookies decrypt.
- consent-gated browser_exec 'local' arg (schema only appears with consent)
- fail-closed on non-Chromium default / snapshot failure
- stale-session guard: reuse only when the live session is on our copy dir
- snapshot excludes extensions/service-workers (renderer wedge) + caches
Profile-spawned workers (kanban bots, cron jobs) can inherit a PATH of
only version-manager dirs — observed in the wild as one nvm node dir
repeated 7x. The uv-installed browser-use binary is a POSIX sh
trampoline that resolves dirname/realpath through PATH, so it died
with 'realpath: not found … exec: /python: not found' (exit 127)
before its own Python ever started.
_base_subprocess_env now floors the child PATH via browser_tool's
_merge_browser_path (the agent-browser backend already guards the same
hazard), degrading to appending FHS bin dirs if that import is ever
unavailable. Windows is a no-op (.cmd shims don't trampoline).
Verified: unit tests + real uvx browser-use --version under a
nvm-only-PATH worker env, rc 127 -> rc 0.
Review pass findings on the force_jpeg change:
- Broaden the JPEG mode guard from {RGBA, P} to 'not in {RGB, L}':
force_jpeg newly routes PNG inputs to the JPEG encoder, and an
LA-mode PNG (grayscale+alpha) would crash img.save() with
'cannot write mode LA as JPEG'.
- browser_use_cli's _native_screenshot_result is the THIRD native
history-embed site: it baked the data URL into a _multimodal tool
result with the 5 MB one-shot default and no dimension cap. Apply
the same 256KB/1568px/force_jpeg history-reuse policy as the two
sites already migrated.
Follow-up to #86916. That fix gave named sessions their own daemon
(socket/log/pid) and their own provider browser — but on a SHARED local
Chrome / CDP browser, a fresh named daemon still attaches to the first
existing page, the same page a sibling daemon may hold. A named session
that never calls new_tab() could still stomp another's tab.
browser_exec now prepends a small preamble to the model's code for named
sessions on shared browsers: once per daemon process (marker keyed by
uid + BU_NAME + daemon pid), it creates a fresh tab via
Target.createTarget and switch_tab()s onto it before any model code
runs. Private per-name browsers (provider-keyed bu-named-<name>, or
direct-API Browser Use cloud) skip the preamble via an internal env
sentinel popped before launch — there's nobody to collide with, and the
extra tab would leak.
Best-effort by design: if the preamble's CDP calls fail, behavior
degrades to pre-fix, never blocks the exec.
E2E against a shared headless Chrome with the STOCK harness: two named
sessions issuing bare js() writes (no new_tab) kept distinct state
(EDGE-A/EDGE-B read back intact); the sabotage run without the preamble
reproduced the clobber (both read EDGE-B). Removes the dependency on the
upstream browser-harness tab-isolation PR for correctness.
session=<name> previously set BU_NAME and then skipped backend resolution
entirely — the parameter was documented as cloud-only, so all local/CDP
work funneled through the single default daemon and one IPC socket, and
concurrent sessions (parallel subagents, simultaneous chats) clobbered
each other's browser connection. Reported by @shantanugoel on X.
Now a named session composes with whatever browser source is configured:
- BU_NAME still namespaces the harness daemon (per-name IPC socket, log,
pid — upstream already isolates these), for local Chrome and CDP.
- The /browser connect CDP override is now exported for named sessions
too; previously a named daemon ignored it and fell back to scanning
local Chrome profiles.
- On provider backends (Browserbase, Firecrawl, Nous gateway), the name
keys its own provider browser via the shared _get_session_info cache
(bu-named-<name>), so each name gets its own cloud browser, the same
name reuses one across calls and tasks, and unnamed calls keep the
per-task key.
- Direct-API Browser Use cloud configs keep the native named-daemon path
(provider resolution would double-session and double-bill).
Tool schema/description updated so models reach for session=<name> for
parallel work on any backend, not just cloud.
E2E: two named sessions against a real headless Chrome (real browser-use
CLI, BU_CDP_URL) ran concurrently, set distinct page state, and read it
back intact; sabotage run confirms the new tests fail without the fix.
The browser-use CLI runs under its own Python (uv tool / uvx), which
can differ from Hermes's venv interpreter. PYTHONPATH/PYTHONHOME
inherited from the agent process point at Hermes's venv
site-packages, and a child interpreter honors them ahead of its own —
so the CLI imported compiled C-extensions (pydantic_core) built for
the wrong interpreter and crashed with ABI mismatch /
ModuleNotFoundError (issues 83427, 84841, 86006, 86104; hits the
desktop backend on py3.14 and any shell exporting PYTHONPATH).
Strip both vars in _base_subprocess_env() — the CLI manages its own
environment and never needs Hermes's import path.
Salvaged from PR 83471 by Benjamin (@n1majne3), the earliest of two
independent fixes (also PR 84022 by @jklance16, PYTHONPATH-only);
regression test covers both vars and preserves unrelated env.
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.
- _find_cli(): probe order flipped to managed bin -> PATH ->
user-level tool dir (then uvx across the same order). A user's own
uv tool install can no longer shadow the Hermes-managed copy with a
drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
install — only the managed copy does, so selecting any backend
provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
and always delegates to install_cli(), the single owner of the
managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
HERMES_HOME/bin/browser-use counts as installed.
Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.
Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
Desktop/TUI workers can spawn with a minimal PATH that omits
~/.local/bin, the default location where uv tool install links the
browser-use binary. _find_cli() then failed to resolve an installed
CLI and Browser Use mode silently fell back to the built-in tools.
Probe the user-level tool dir (~/.local/bin on POSIX, APPDATA/uv/bin
on Windows) between PATH and the managed HERMES_HOME/bin, for both
the browser-use binary and the uvx fallback.
Salvaged from PR #83788 by @kimyxx onto current main; tests adapted
and extended with precedence and uvx coverage.
Sweep of open Windows issues affecting day-to-day agent operation
(explicitly excluding install/setup and locale classes):
- hermes_cli/_subprocess_compat.py: new split_command_line() — Windows-
safe command-line tokenizer (posix=False + quote stripping) so
backslash paths survive. POSIX behavior unchanged (plain shlex.split).
- hermes_cli/console_engine.py (#83934): console commands like
'sessions export C:\Users\me\out.jsonl' no longer silently mangle the
path into a relative filename in the cwd.
- agent/shell_hooks.py (#78293): hook commands with backslash paths now
spawn, resolve their script path, and pass hooks doctor instead of
reporting 'not executable'. All three shlex sites routed through the
shared splitter.
- agent/prompt_builder.py (#51755): system prompt now reports
Windows (11) on Windows 11 — platform.release() returns 10 for both;
distinguish via sys.getwindowsversion().build >= 22000.
- hermes_cli/commands.py (#42016): @ autocomplete no longer crashes the
prompt_toolkit event loop when rg emits a path on a different mount
(device paths \.\nul, other drive letters) — relpath ValueError is
skipped per-entry.
- tools/browser_use_cli.py (#83884): screenshot-path detection now
matches Windows drive-letter paths (C:\... and C:/...) in addition to
POSIX; Browser Use screenshots attach on Windows.
- tools/skills_hub.py + tools/skills_guard.py (#62310): the two 'MUST
stay symmetric' skill content hashes actually agree on Windows now.
Bundle keys are normalized to POSIX separators before hashing, and the
disk digest sorts by rel-posix STRING (case-sensitive) instead of Path
objects (case-insensitive on Windows). Fixes permanent false-positive
update_available for every installed skill.
Tests: tests/tools/test_windows_agent_loop_papercuts.py — 16 cases
covering each fix, including a disk-vs-bundle hash symmetry check built
with native Windows separators and a mixed-case filename.
The Browser Use CLI became the default browser backend, but nothing
provisioned it: users without uv/uvx (field report from DongyangHe on
macOS) silently fell back to the built-in browser tools with no notice.
- install_cli() in tools/browser_use_cli.py: uv tool install browser-use
via the managed uv (bootstrapped on demand), linked into
$HERMES_HOME/bin (UV_TOOL_BIN_DIR)
- _find_cli() now also probes $HERMES_HOME/bin for browser-use/uvx —
Hermes' managed uv is not on the user's PATH
- hermes tools post_setup actually installs (Camofox standard) instead
of printing instructions
- install.sh / install.ps1 provision the CLI at install time
(best-effort, non-fatal, honors --skip-browser)
- CLI startup shows a one-line notice (24h rate-limited) when the
default backend downgraded to the built-in tools
An unset browser.backend ("") now resolves to Browser Use mode whenever
the browser-use CLI is runnable (installed binary or uvx); otherwise the
built-in browser tools are kept so browsing never silently breaks.
Camofox setups always keep the built-in tools (no CDP surface), and
backend: off (including YAML 1.1 bare off -> False) forces the built-in
stack. hermes tools row highlighting follows the same effective-mode
resolution, and tests/tools/ pins CLI discovery off so host uvx installs
can't flip built-in-browser tests.
Reframe (per review): browser.backend: browser-use is now a DRIVER over
whatever browser source is configured, not a competing backend choice.
- browser_exec resolves its CDP endpoint through the same chain the
built-in tools use: BU_* env override > BROWSER_CDP_URL/browser.cdp_url
(/browser connect) > the configured cloud provider via browser_tool's
_get_session_info() — sharing the per-task session cache, expiry
replacement, inactivity reaper, and atexit cleanup instead of
duplicating them. Live-validated against Browserbase (session created,
driven, reaped) and gateway-provisioned Browser Use cloud browsers.
- Direct-API Browser Use configs skip provider resolution (the CLI talks
to their cloud natively via BU_AUTOSPAWN); the Nous-gateway variant
resolves through the provider, so subscribers get CLI mode without a
raw BROWSER_USE_API_KEY.
- Camofox: only true fallback — Firefox-based, custom HTTP API, no CDP
surface (its own health probes fail on CDP-schema calls). Active
Camofox setups keep the built-in browser tools even with
backend: browser-use set.
- hermes tools picker: provider rows and the Browser Use row are no
longer mutually exclusive; selecting a provider keeps the driver
choice, and both rows highlight when composed.
- Docs updated for driver-over-source semantics.
Camofox is selected via CAMOFOX_URL env var, not browser.cloud_provider —
so a Camofox user with a stray BROWSER_USE_API_KEY in .env matched the
legacy-migration predicate (cloud_provider unset + key present) and got
silently flipped into CLI mode, losing browser_* / Camofox entirely
(browser_exec cannot drive Camofox: its HTTP API exposes no CDP endpoint,
and the browser-use harness is CDP-only against Chromium).
is_legacy_browser_use_cloud_config() now defers to is_camofox_mode().
Follow-ups on the salvaged Browser Use CLI integration (PR #66476):
- browser_exec runs model-written Python on the host. Strip it at
tool-definition time for sessions whose resolved toolsets exclude
'terminal' so terminal-less surfaces (locked-down messaging configs)
don't silently regain host code execution through the browser toolset.
Session-level gate in model_tools, not a check_fn (check_fn results are
TTL-cached process-wide across sessions).
- Replace the live 'browser-use skill' schema fetch with a pinned helpers
digest: no third-party version-drifting text in the prompt, byte-stable
schema across machines. A/B benchmarked (108 runs, opus-4.8 + kimi-k3,
6 multi-step web tasks x 3 arms x 3 reps): pinned digest matches the
full skill dump 36/36 vs 36/36 at ~equal tokens; both cut total task
tokens ~60% vs the legacy browser_* toolset.
- Docs note for the terminal gate; contributor mapping for salvage.