* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
805 lines
41 KiB
Python
805 lines
41 KiB
Python
"""Use the Browser Use CLI 3.0 (https://browser-use.com) for browser automation
|
|
|
|
When browser.backend is "browser-use", the model gets ``browser_exec`` tool
|
|
instead of default browser tools
|
|
"""
|
|
|
|
import contextlib
|
|
import importlib
|
|
import importlib.util
|
|
import json
|
|
import logging
|
|
import os
|
|
import re
|
|
import signal
|
|
import subprocess
|
|
import sys
|
|
import time
|
|
from pathlib import Path
|
|
from typing import Any, Callable, Dict, List, Optional
|
|
|
|
from hermes_constants import get_hermes_home
|
|
from utils import is_truthy_value
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
_BACKEND_KEY = "browser-use"
|
|
BACKEND_DISABLED = "off"
|
|
|
|
# Cloud daemon names become the BU_NAME env var
|
|
_SESSION_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9_-]{0,63}$")
|
|
|
|
# Set on the env dict by the CDP resolvers when the resolved browser is EXCLUSIVE to this named session
|
|
# (per-name provider / named BU cloud / Lightpanda). Popped before the subprocess launches — never exported.
|
|
_PRIVATE_BROWSER_SENTINEL = "_HERMES_BU_PRIVATE_BROWSER"
|
|
# Internal route provenance: this exec resolved to a browser on the Bot Desktop display and must use
|
|
# the same human-control lease fence as the built-in browser tools. Popped before launching the CLI.
|
|
_BOT_DESKTOP_BROWSER_SENTINEL = "_HERMES_BU_BOT_DESKTOP_BROWSER"
|
|
|
|
# Prepended to the model's code for named sessions on SHARED browsers (a /browser connect CDP override): the
|
|
# harness daemon attaches to the first existing page at startup, so two fresh named daemons can land on the
|
|
# SAME tab. Steering each onto a tab it created prevents clobbering. Runs once per daemon (marker keyed by
|
|
# BU_NAME + daemon pid).
|
|
_OWN_TAB_PREAMBLE = """\
|
|
# hermes: pin this named session to its own tab (once per daemon process)
|
|
def _hermes_ensure_own_tab():
|
|
import os as _os, tempfile as _tf
|
|
_name = _os.environ.get("BU_NAME", "default")
|
|
try:
|
|
# Key the marker by the daemon's pid so a daemon restart (which
|
|
# re-attaches to the first shared page) re-pins automatically,
|
|
# while agent-driven tab switches mid-session are left alone.
|
|
from browser_harness import _ipc as _bipc
|
|
_dpid = _bipc.pid_path(_name).read_text().strip() or "0"
|
|
except Exception:
|
|
_dpid = "0"
|
|
_uid = _os.getuid() if hasattr(_os, "getuid") else 0
|
|
_marker = _os.path.join(
|
|
_tf.gettempdir(), "hermes-bu-owntab-%s-%s-%s" % (_uid, _name, _dpid)
|
|
)
|
|
if _os.path.exists(_marker):
|
|
return
|
|
try:
|
|
# Force a fresh target: new_tab() would REUSE a blank current tab,
|
|
# which is exactly the tab a sibling daemon may also hold.
|
|
_tid = cdp("Target.createTarget", url="about:blank").get("targetId")
|
|
if _tid:
|
|
switch_tab(_tid)
|
|
except Exception:
|
|
pass # best-effort: worst case is pre-fix behavior
|
|
try:
|
|
open(_marker, "w").close()
|
|
except OSError:
|
|
pass
|
|
_hermes_ensure_own_tab()
|
|
del _hermes_ensure_own_tab
|
|
"""
|
|
|
|
_DEFAULT_TIMEOUT_S = 300
|
|
_MIN_TIMEOUT_S = 5
|
|
_MAX_TIMEOUT_S = 1800
|
|
_STDERR_CAP_CHARS = 4000
|
|
|
|
_TASK_ID_SAFE_RE = re.compile(r"[^A-Za-z0-9._-]+") # filesystem-safe task ids
|
|
# Screenshot paths printed by capture_screenshot(): POSIX or Windows drive-letter absolute.
|
|
_IMAGE_PATH_RE = re.compile(r"((?:[A-Za-z]:[\\/]|/)[^\s\"']+?\.(?:png|jpe?g|webp))", re.IGNORECASE)
|
|
# http(s) URL literals in exec code checked against browser_navigate's policy
|
|
_URL_RE = re.compile(r"https?://[^\s'\"\\)]+", re.IGNORECASE)
|
|
_FHS_BIN_DIRS = ("/usr/local/sbin", "/usr/local/bin", "/usr/sbin", "/usr/bin", "/sbin", "/bin")
|
|
|
|
|
|
def _quiet(fn: Callable[[], Any], default: Any, log_prefix: str = "") -> Any:
|
|
"""``fn()``, or ``default`` on any exception (debug-logged when ``log_prefix`` is set)."""
|
|
try:
|
|
return fn()
|
|
except Exception as e:
|
|
if log_prefix:
|
|
logger.debug("%s: %s", log_prefix, e)
|
|
return default
|
|
|
|
|
|
def _lazy_call(module: str, name: str, default: Any, log_prefix: str) -> Any:
|
|
"""Call ``module.name()`` resolved at call time (tests stub the module / patch the
|
|
attribute); on any failure log ``log_prefix`` and return ``default``."""
|
|
return _quiet(lambda: getattr(importlib.import_module(module), name)(), default, log_prefix)
|
|
|
|
|
|
def _camofox_active(context: str = "") -> bool:
|
|
return _lazy_call("tools.browser_camofox", "is_camofox_mode", False, f"Camofox activity check failed{context}")
|
|
|
|
|
|
def _real_profile_consented() -> bool:
|
|
return _lazy_call("tools.browser_tool_cloud", "_use_real_profile", False, "real-profile consent lookup failed")
|
|
|
|
|
|
def _set_cdp_env(env: dict, cdp: str) -> None:
|
|
"""Export a CDP endpoint under the BU_CDP_* contract (http(s) → URL, else WS)."""
|
|
env["BU_CDP_URL" if cdp.startswith(("http://", "https://")) else "BU_CDP_WS"] = cdp
|
|
|
|
|
|
def _has_cdp_env(env: dict) -> bool:
|
|
return bool(env.get("BU_CDP_WS") or env.get("BU_CDP_URL"))
|
|
|
|
|
|
def _export_session_cdp(env: dict, get_session_info: Callable[[str], Any], cache_key: str,
|
|
fail_msg: Callable[[Exception], str], no_cdp_msg: str) -> Optional[str]:
|
|
"""Export the CDP endpoint from ``get_session_info(cache_key)``; error string on failure / no CDP."""
|
|
try:
|
|
cdp = str((get_session_info(cache_key) or {}).get("cdp_url") or "")
|
|
except Exception as e:
|
|
return fail_msg(e)
|
|
if not cdp:
|
|
return no_cdp_msg
|
|
_set_cdp_env(env, cdp)
|
|
return None
|
|
|
|
|
|
def _blocked_url_in_code(code: str) -> Optional[str]:
|
|
"""Return an error if a URL literal fails the built-in navigation checks."""
|
|
from tools.browser_tool import evaluate_url_safety
|
|
return next((err.get("error", "Blocked: unsafe URL") for err in map(evaluate_url_safety, _URL_RE.findall(code or "")) if err), None)
|
|
|
|
|
|
def _base_subprocess_env() -> dict:
|
|
from tools.browser_tool import _build_browser_env
|
|
env = _build_browser_env()
|
|
# The harness runs on Hermes's own interpreter, but a bundled Desktop install boots that
|
|
# interpreter with its site dir on PYTHONPATH (no venv to activate), and the harness's daemon
|
|
# re-runs sys.executable. Point PYTHONPATH at the harness's site dir, replacing whatever the
|
|
# agent process inherited, so both the CLI and its daemon import the same packages.
|
|
env.pop("PYTHONHOME", None)
|
|
site_dir = _harness_site_dir()
|
|
if site_dir:
|
|
env["PYTHONPATH"] = site_dir
|
|
else:
|
|
env.pop("PYTHONPATH", None)
|
|
env["PATH"] = _floor_subprocess_path(env.get("PATH", ""))
|
|
env.setdefault("ANONYMIZED_TELEMETRY", "false")
|
|
return env
|
|
|
|
|
|
def _floor_subprocess_path(path: str) -> str:
|
|
"""Keep system commands reachable from profile workers with a minimal PATH.
|
|
|
|
Browser subprocesses may launch shell helpers; version-manager-only paths
|
|
omit /usr/bin. Reuse the shared browser PATH floor on POSIX.
|
|
"""
|
|
if os.name == "nt":
|
|
return path
|
|
with contextlib.suppress(Exception):
|
|
from tools.browser_tool_install import _merge_browser_path
|
|
return _merge_browser_path(path or "")
|
|
parts = [p for p in (path or "").split(os.pathsep) if p]
|
|
return os.pathsep.join(parts + [d for d in _FHS_BIN_DIRS if d not in set(parts) and os.path.isdir(d)])
|
|
|
|
|
|
def _read_browser_cfg() -> dict:
|
|
"""Return the ``browser:`` config section, or {} on any failure."""
|
|
try:
|
|
from hermes_cli.config import cfg_get, read_raw_config
|
|
cfg = cfg_get(read_raw_config(), "browser", default={})
|
|
return cfg if isinstance(cfg, dict) else {}
|
|
except Exception as e:
|
|
logger.debug("Could not read browser config section: %s", e)
|
|
return {}
|
|
|
|
|
|
def _use_gateway(browser_cfg: dict) -> bool:
|
|
"""True when the browser section selects the Nous Tool Gateway — by the current ``hermes tools``
|
|
picker row (``cloud_provider: nous``) or the pre-picker ``use_gateway: true`` flag. Reading only
|
|
the legacy flag missed every picker-configured gateway, and the direct-API branch it fell into
|
|
holds no credentials in managed mode (#108310)."""
|
|
if is_truthy_value(browser_cfg.get("use_gateway"), default=False):
|
|
return True
|
|
try:
|
|
from tools.tool_backend_helpers import NOUS_MANAGED_PROVIDER
|
|
except Exception: # pragma: no cover — helper ships with the package
|
|
return False
|
|
return str(browser_cfg.get("cloud_provider") or "").strip().lower() == NOUS_MANAGED_PROVIDER
|
|
|
|
|
|
def get_browser_backend() -> str:
|
|
"""Configured browser backend key ("" = unset → default). YAML 1.1 parses an
|
|
unquoted ``off`` as False — that must mean BACKEND_DISABLED, not "unset"."""
|
|
raw = _read_browser_cfg().get("backend")
|
|
return (BACKEND_DISABLED if raw is False else "") if isinstance(raw, bool) else str(raw or "").strip().lower()
|
|
|
|
|
|
def is_legacy_browser_use_cloud_config(browser_cfg: dict) -> bool:
|
|
"""True for pre-CLI direct-API Browser Use cloud configs. An explicit backend or
|
|
a non-Browser-Use cloud_provider wins; Camofox is selected via env var, not
|
|
cloud_provider, so a Camofox user with a stray BROWSER_USE_API_KEY keeps it."""
|
|
if not isinstance(browser_cfg, dict) or browser_cfg.get("backend"):
|
|
return False
|
|
provider = str(browser_cfg.get("cloud_provider") or "").strip().lower()
|
|
if provider not in {"browser-use", ""} or _use_gateway(browser_cfg) or _camofox_active(" during migration"):
|
|
return False
|
|
# Profile credential: a multiplexed secondary must not inherit the default's cloud mode.
|
|
from agent.secret_scope import get_secret
|
|
return bool(get_secret("BROWSER_USE_API_KEY", ""))
|
|
|
|
|
|
def is_browser_use_cli_mode() -> bool:
|
|
"""True when the Browser Use CLI replaces the built-in browser stack. Browser Use mode is the DEFAULT:
|
|
unset ``browser.backend`` ("") enables it whenever browser-harness is importable (a core dependency);
|
|
``browser.backend: off`` keeps the built-in browser_* tools. Camofox always falls back to the built-in
|
|
tools (Firefox, custom HTTP API, no CDP surface for the harness)."""
|
|
if _camofox_active():
|
|
return False
|
|
backend = get_browser_backend()
|
|
return backend == _BACKEND_KEY if backend else (is_legacy_browser_use_cloud_config(_read_browser_cfg()) or _find_cli() is not None)
|
|
|
|
|
|
def default_downgrade_notice() -> Optional[str]:
|
|
"""One-line notice when ``browser.backend`` is unset but the CLI is not runnable, so
|
|
the session fell back to the built-in tools. Rate-limited to once per 24h via a stamp file."""
|
|
try:
|
|
if get_browser_backend() or _camofox_active() or _find_cli() is not None:
|
|
return None # explicit choice / Camofox / CLI present — nothing downgraded
|
|
stamp = Path(get_hermes_home()) / "cache" / ".browser_use_default_notice"
|
|
now = time.time()
|
|
with contextlib.suppress(OSError):
|
|
if 0 <= now - stamp.stat().st_mtime < 24 * 3600:
|
|
return None
|
|
with contextlib.suppress(OSError):
|
|
stamp.parent.mkdir(parents=True, exist_ok=True)
|
|
stamp.touch()
|
|
os.utime(stamp, (now, now))
|
|
return ("browser-harness is missing from Hermes's Python environment — using the built-in browser tools. "
|
|
"Run `hermes update` to re-sync it, or set `browser.backend: off` in config.yaml to silence this.")
|
|
except Exception as e: # pragma: no cover — a notice must never break startup
|
|
logger.debug("browser-use downgrade notice failed: %s", e)
|
|
return None
|
|
|
|
|
|
def _harness_site_dir() -> Optional[str]:
|
|
"""The site dir Hermes's interpreter imports ``browser_harness`` from, or None."""
|
|
spec = importlib.util.find_spec("browser_harness")
|
|
if spec is None or not spec.origin:
|
|
return None
|
|
return str(Path(spec.origin).resolve().parent.parent)
|
|
|
|
|
|
def _find_cli() -> Optional[List[str]]:
|
|
"""The Browser Use CLI's engine (browser-harness) is a core dependency of Hermes's own venv,
|
|
so every install, the Desktop bundle included, runs it on the current interpreter."""
|
|
if _harness_site_dir() is None:
|
|
return None
|
|
return [sys.executable, "-m", "browser_harness.run"]
|
|
|
|
|
|
def _workspace_dir(task_id: Optional[str]) -> Optional[str]:
|
|
"""Stable per-task scratch dir that persists across browser_exec calls"""
|
|
if os.environ.get("BH_AGENT_WORKSPACE"):
|
|
return os.environ["BH_AGENT_WORKSPACE"]
|
|
try:
|
|
safe = _TASK_ID_SAFE_RE.sub("_", str(task_id or "default"))[:80] or "default"
|
|
path = Path(get_hermes_home()) / "cache" / "browser-use" / "workspace" / safe
|
|
path.mkdir(parents=True, exist_ok=True)
|
|
return str(path)
|
|
except Exception as e:
|
|
logger.debug("browser_exec workspace unavailable: %s", e)
|
|
return None
|
|
|
|
|
|
def _find_screenshot(stdout: str, since: float) -> Optional[str]:
|
|
"""Last screenshot path printed during this exec that exists and was written after
|
|
the exec started, or None."""
|
|
for path in reversed(_IMAGE_PATH_RE.findall(stdout or "")):
|
|
try:
|
|
if os.path.isfile(path) and os.path.getmtime(path) >= since - 1:
|
|
return path
|
|
except OSError:
|
|
continue
|
|
return None
|
|
|
|
|
|
def _native_screenshot_result(result: Dict[str, Any], path: str) -> Optional[Dict[str, Any]]:
|
|
"""Build a multimodal tool result attaching path for vision models"""
|
|
try:
|
|
from tools.vision_tools import (_EMBED_MAX_DIMENSION,
|
|
_resize_image_for_vision, _should_use_native_vision_fast_path)
|
|
from tools.vision_tools_history_budget import resolve_embed_target_bytes
|
|
if not _should_use_native_vision_fast_path():
|
|
return None
|
|
# History-reuse cap: this data URL bakes into the tool result and is re-sent every later turn —
|
|
# same policy as the vision_analyze / browser_vision native embeds.
|
|
data_url = _resize_image_for_vision(Path(path), mime_type="image/png",
|
|
max_base64_bytes=resolve_embed_target_bytes(),
|
|
max_dimension=_EMBED_MAX_DIMENSION, force_jpeg=True)
|
|
text = json.dumps(result, ensure_ascii=False)
|
|
attached = text + "\n\nThe screenshot from this call is attached — inspect it with your native vision."
|
|
return {"_multimodal": True, "text_summary": text, "meta": {"screenshot_path": path, "native_vision": True},
|
|
"content": [{"type": "text", "text": attached}, {"type": "image_url", "image_url": {"url": data_url}}]}
|
|
except Exception as e:
|
|
logger.debug("Native screenshot attach failed (falling back to text): %s", e)
|
|
return None
|
|
|
|
|
|
def _served_profile_tag() -> str:
|
|
"""``""`` outside a served-profile scope (every legacy key stays byte-identical); under a
|
|
multiplexed turn, the routed profile's home key — one profile's browser must never be handed
|
|
to another that happens to use the same session name or task id (#110032)."""
|
|
from hermes_constants import get_hermes_home_override, hermes_home_key
|
|
return "" if get_hermes_home_override() is None else hermes_home_key()
|
|
|
|
|
|
def _backend_cache_key(task_id: Optional[str], session_name: str = "") -> str:
|
|
"""Session-cache key for a backend browser: named sessions get their own; served profiles get their own."""
|
|
key = f"bu-named-{session_name}" if session_name else (task_id or "browser-exec-default")
|
|
tag = _served_profile_tag()
|
|
return f"{key}@{tag}" if tag else key
|
|
|
|
|
|
def _resolve_lightpanda_cdp(env: dict, task_id: Optional[str], session_name: str = "") -> Optional[str]:
|
|
"""Point the harness at a Hermes-spawned ``lightpanda serve`` (``browser.engine: lightpanda`` and
|
|
nothing of higher precedence claimed the session). Each cache key gets its own process via the
|
|
legacy ``_get_session_info()`` (cache, reaper, atexit): private browser, own-tab preamble skipped."""
|
|
try:
|
|
from tools.browser_tool_session import _get_session_info
|
|
from tools.browser_tool_lightpanda_fallback import _using_lightpanda_engine
|
|
if not _using_lightpanda_engine():
|
|
return None
|
|
except Exception as e: # stubbed browser_tool in tests / engine lookup failure
|
|
logger.debug("browser engine lookup failed: %s", e)
|
|
return None
|
|
err = _export_session_cdp(
|
|
env, _get_session_info, _backend_cache_key(task_id, session_name),
|
|
lambda e: (f"Lightpanda could not be started: {e} Set browser.engine to auto "
|
|
"to use local Chrome, or switch backends via `hermes tools` → Browser Automation."),
|
|
"Lightpanda session returned no CDP endpoint. Set browser.engine to auto to use local Chrome.",
|
|
)
|
|
if err is None:
|
|
env[_PRIVATE_BROWSER_SENTINEL] = "1"
|
|
env[_BOT_DESKTOP_BROWSER_SENTINEL] = "1"
|
|
return err
|
|
|
|
|
|
def _reach_sandbox_cdp(cdp: str) -> str:
|
|
"""A CDP endpoint agent-browser reported from INSIDE the terminal backend's sandbox is that sandbox's
|
|
loopback: unreachable from this host (Docker bridge / ssh remote). The harness, the vault supervisor
|
|
and ``browser_exec`` all connect from here, so forward the port over the sandbox's exec stream and hand
|
|
them the local end. Chromium's DevTools accepts any loopback ``Host`` header, port included."""
|
|
try:
|
|
from tools.browser_tool_session import _browser_in_sandbox
|
|
if not _browser_in_sandbox():
|
|
return cdp
|
|
from urllib.parse import urlsplit, urlunsplit
|
|
from tools.bot_desktop import runtime as _bd_runtime, sandbox_host
|
|
from tools.environments import streams
|
|
parts = urlsplit(cdp)
|
|
if parts.hostname not in ("127.0.0.1", "localhost") or not parts.port:
|
|
return cdp
|
|
sandbox = _bd_runtime._sandbox_env(create=True)
|
|
if sandbox is None:
|
|
return cdp
|
|
local = streams.forward_port(sandbox, parts.port, user=sandbox_host._user_for(sandbox))
|
|
return urlunsplit(parts._replace(netloc=f"127.0.0.1:{local}"))
|
|
except Exception as e: # a failed forward degrades to the unreachable endpoint's own error
|
|
logger.debug("sandbox CDP forward unavailable: %s", e)
|
|
return cdp
|
|
|
|
|
|
def _resolve_managed_chromium_cdp(env: dict, task_id: Optional[str], session_name: str = "") -> Optional[str]:
|
|
"""Point the harness at Hermes' packaged Chromium, launched through agent-browser for this cache key —
|
|
the same browser the built-in tools drive. Left alone, the harness discovers the user's INSTALLED
|
|
Chrome on its default profile, which needs the chrome://inspect toggle + an Allow popup per run and
|
|
is blocked outright on Chrome >=136; on a headless box it just reports ``chrome-not-running``.
|
|
``get cdp-url`` runs through ``_run_browser_command`` (legacy cache, inactivity reaper, atexit, Chromium
|
|
preflight/auto-install) on EVERY call: it launches the browser cold, follows a relaunch, and refreshes
|
|
the agent-browser daemon's idle timer, which never sees the harness's direct CDP traffic."""
|
|
try:
|
|
from tools.browser_tool_session import _run_browser_command
|
|
from tools.browser_tool import _get_open_command_timeout
|
|
except Exception as e: # pragma: no cover — stubbed browser_tool in tests
|
|
logger.debug("managed chromium resolution unavailable: %s", e)
|
|
return None
|
|
res = _run_browser_command(_backend_cache_key(task_id, session_name), "get", ["cdp-url"],
|
|
timeout=_get_open_command_timeout(first_open=True))
|
|
cdp = str(((res or {}).get("data") or {}).get("cdpUrl") or "") if (res or {}).get("success") else ""
|
|
if not cdp:
|
|
return (f"The local browser could not be started: {(res or {}).get('error') or 'agent-browser returned no CDP endpoint'} "
|
|
"Run `hermes tools` → Browser Automation to (re)install Chromium, or switch backends.")
|
|
cdp = _reach_sandbox_cdp(cdp)
|
|
_set_cdp_env(env, cdp)
|
|
env[_PRIVATE_BROWSER_SENTINEL] = "1" # one Chromium per cache key: nothing to share a tab with
|
|
env[_BOT_DESKTOP_BROWSER_SENTINEL] = "1"
|
|
return None
|
|
|
|
|
|
def _resolve_local_engine_cdp(env: dict, task_id: Optional[str], session_name: str = "") -> Optional[str]:
|
|
"""Local engine (no provider / override): ``browser.engine: lightpanda`` or the packaged Chromium."""
|
|
err = _resolve_lightpanda_cdp(env, task_id, session_name)
|
|
if err or _has_cdp_env(env):
|
|
return err
|
|
return _resolve_managed_chromium_cdp(env, task_id, session_name)
|
|
|
|
|
|
def _resolve_backend_cdp(env: dict, task_id: Optional[str], session_name: str = "") -> Optional[str]:
|
|
"""Point the harness at the configured backend's CDP endpoint; error string on failure.
|
|
|
|
Precedence: (1) ``BU_CDP_WS``/``BU_CDP_URL`` already in env (operator override); (2) ``BROWSER_CDP_URL``
|
|
env / ``browser.cdp_url`` (``/browser connect``); (3) a cloud provider via the legacy ``_get_session_info()``
|
|
so browser_exec shares the SAME session machinery (per-task cache, expiry, reaper, atexit);
|
|
(4) the local engine — ``browser.engine: lightpanda`` or Hermes' packaged Chromium via agent-browser
|
|
(never the harness's own discovery of the user's installed Chrome); (5) BU direct-API configs → None:
|
|
the CLI reaches BU cloud natively (BU_AUTOSPAWN). ``session_name`` (BU_NAME) keys the session cache so
|
|
each name gets its OWN browser — what makes named sessions concurrent-safe.
|
|
"""
|
|
if _has_cdp_env(env):
|
|
return None
|
|
try:
|
|
from tools.browser_tool_cloud import _get_cloud_provider
|
|
from tools.browser_tool_session import _get_session_info
|
|
from tools.browser_tool_cdp import _get_cdp_override
|
|
except Exception as e: # pragma: no cover — stubbed browser_tool in tests
|
|
logger.debug("browser_tool backend resolution unavailable: %s", e)
|
|
return None
|
|
override = _quiet(_get_cdp_override, "")
|
|
if override:
|
|
_set_cdp_env(env, override)
|
|
return None
|
|
provider = _quiet(_get_cloud_provider, None, "Cloud provider lookup failed")
|
|
if provider is None:
|
|
return _resolve_local_engine_cdp(env, task_id, session_name)
|
|
|
|
# Browser Use direct-API configs: the CLI talks to BU cloud natively (BU_AUTOSPAWN / auth login) — the
|
|
# legacy provider would create a second, redundant session. Nous-gateway configs (cloud_provider: nous
|
|
# from the picker, or the pre-picker use_gateway: true) DO resolve through the provider: the gateway
|
|
# provisions the browser server-side and returns its CDP URL.
|
|
provider_key = str(getattr(provider, "name", "") or "").strip().lower()
|
|
if provider_key == _BACKEND_KEY and not _use_gateway(_read_browser_cfg()):
|
|
env[_PRIVATE_BROWSER_SENTINEL] = "1" # named BU cloud browsers are exclusive to their daemon
|
|
return None
|
|
|
|
provider_name = type(provider).__name__
|
|
err = _export_session_cdp(
|
|
env, _get_session_info, _backend_cache_key(task_id, session_name),
|
|
lambda e: (f"Cloud browser provider {provider_name} failed to provide a session: {e}. "
|
|
"Fix the provider configuration or switch backends via `hermes tools` → Browser Automation."),
|
|
f"Cloud browser provider {provider_name} returned no CDP endpoint, so Browser Use mode "
|
|
"cannot drive it. Switch to the built-in browser tools for this provider.",
|
|
)
|
|
# A provider browser keyed bu-named-<name> is exclusive to this session — the
|
|
# own-tab preamble would just leak a blank tab into it.
|
|
if err is None and session_name:
|
|
env[_PRIVATE_BROWSER_SENTINEL] = "1"
|
|
return err
|
|
|
|
|
|
def _resolve_real_profile_cdp(env: dict, force_local: bool) -> Optional[str]:
|
|
"""Point the harness at the user's real-profile copy-browser (a SNAPSHOT of their default Chromium
|
|
profile, hermes_cli.browser_connect) when consented. Two ways in: the effective backend is already local
|
|
(no provider, CDP override, or legacy BU cloud config) → silent upgrade; or ``force_local`` (consent-gated
|
|
``local`` arg) → the user's browser even under a cloud backend. Operator overrides (BU_CDP_* env,
|
|
/browser connect, ``browser.cdp_url``) own the session either way. Fail closed: a launch error is
|
|
returned so a consented user is never silently downgraded."""
|
|
if not _real_profile_consented() or _has_cdp_env(env):
|
|
return None
|
|
try:
|
|
from tools.browser_tool_cdp import _get_cdp_override_raw
|
|
from tools.browser_tool_cloud import _get_cloud_provider
|
|
from tools.browser_tool_real_profile import _real_profile_cdp
|
|
except Exception as e: # pragma: no cover — stubbed browser_tool in tests
|
|
logger.debug("real-profile backend resolution unavailable: %s", e)
|
|
return None
|
|
if _quiet(_get_cdp_override_raw, ""):
|
|
return None
|
|
# Only auto-upgrade genuinely-local attaches; any cloud path (provider, provider lookup failure, or
|
|
# legacy BU cloud config) stays on its backend unless the model passes local=true.
|
|
if not force_local and (_quiet(_get_cloud_provider, object()) is not None
|
|
or is_legacy_browser_use_cloud_config(_read_browser_cfg())):
|
|
return None
|
|
cdp, err = _real_profile_cdp()
|
|
if cdp and not err:
|
|
_set_cdp_env(env, cdp)
|
|
env[_BOT_DESKTOP_BROWSER_SENTINEL] = "1"
|
|
return err or None
|
|
|
|
|
|
def _attach_vault_supervisor(env: dict, task_id: Optional[str]) -> None:
|
|
"""Attach the per-task CDP supervisor to the browser this exec drives so ``browser_vault_fill`` has
|
|
a secret-capable WebSocket (never argv) into the SAME browser. Only CDP-routed backends expose an
|
|
endpoint; BU direct-cloud (BU_AUTOSPAWN) does not, and the vault tools report ``supervisor_required``."""
|
|
cdp = env.get("BU_CDP_WS") or env.get("BU_CDP_URL")
|
|
if not cdp:
|
|
return
|
|
try:
|
|
from tools.browser_supervisor import SUPERVISOR_REGISTRY
|
|
from tools.browser_tool_cdp import _get_dialog_policy_config, _resolve_cdp_override
|
|
policy, timeout_s = _get_dialog_policy_config()
|
|
SUPERVISOR_REGISTRY.get_or_start(task_id=task_id or "default", cdp_url=_resolve_cdp_override(cdp),
|
|
dialog_policy=policy, dialog_timeout_s=timeout_s)
|
|
except Exception as exc:
|
|
logger.debug("browser_exec: CDP supervisor attach failed (non-fatal): %s", exc)
|
|
|
|
|
|
def _route_backend(env: dict, session: str, task_id: Optional[str], local: bool) -> Optional[str]:
|
|
"""Resolve where the harness connects; returns an error string or None. Real-profile consent runs
|
|
BEFORE provider resolution so a hit short-circuits the cloud path via the BU_CDP_* env contract. Named
|
|
sessions compose with the backend: BU_NAME namespaces the harness daemon (IPC socket, log, pid) and on
|
|
provider backends additionally keys its own cloud browser."""
|
|
rp_err = _resolve_real_profile_cdp(env, force_local=local)
|
|
if rp_err:
|
|
return rp_err
|
|
# local=True is only served by the real-profile route; consent off must not pretend.
|
|
if local and not _has_cdp_env(env) and not _real_profile_consented():
|
|
return ("local=true was requested but browser.use_real_profile is off. Enable it in config.yaml "
|
|
"(browser.use_real_profile: true) or the desktop Settings → Browser section, then retry.")
|
|
return _resolve_backend_cdp(env, task_id, session_name=session)
|
|
|
|
|
|
def _group_popen_kwargs() -> dict:
|
|
"""Popen kwargs starting the CLI in its own process group (a new session on POSIX) so a
|
|
timeout can take down every process that inherited the capture pipes, not just the CLI
|
|
child. Windows also hides the console the .cmd shim would flash (as browser_tool does)."""
|
|
def _flags() -> dict:
|
|
from hermes_cli._subprocess_compat import windows_hide_flags
|
|
si = subprocess.STARTUPINFO()
|
|
si.dwFlags |= subprocess.STARTF_USESHOWWINDOW
|
|
return {"creationflags": windows_hide_flags() | getattr(subprocess, "CREATE_NEW_PROCESS_GROUP", 0),
|
|
"startupinfo": si}
|
|
return _quiet(_flags, {}, "Windows hide-flags unavailable") if os.name == "nt" else {"start_new_session": True}
|
|
|
|
|
|
def _clamp_timeout(timeout_s: Any) -> int:
|
|
try:
|
|
return max(_MIN_TIMEOUT_S, min(int(timeout_s), _MAX_TIMEOUT_S))
|
|
except (TypeError, ValueError):
|
|
return _DEFAULT_TIMEOUT_S
|
|
|
|
|
|
# After a whole-group SIGKILL, every pipe holder is dead, so the drain below is normally
|
|
# instant; the deadline only guards against a process outside the group still holding a pipe.
|
|
_POST_KILL_DRAIN_S = 10.0
|
|
|
|
|
|
def _kill_cli_process_group(proc) -> None:
|
|
"""SIGKILL the CLI's whole process group (POSIX; ``start_new_session`` made pgid == pid) or,
|
|
on Windows, its process tree via ``taskkill /T /F`` — the only group-wide kill it offers."""
|
|
if os.name == "nt":
|
|
from hermes_cli._subprocess_compat import windows_hide_flags
|
|
with contextlib.suppress(OSError, subprocess.SubprocessError):
|
|
subprocess.run(["taskkill", "/T", "/F", "/PID", str(proc.pid)], stdin=subprocess.DEVNULL,
|
|
capture_output=True, text=True, encoding="utf-8", errors="replace", timeout=10,
|
|
check=False, creationflags=windows_hide_flags())
|
|
return
|
|
with contextlib.suppress(ProcessLookupError, PermissionError):
|
|
os.killpg(proc.pid, signal.SIGKILL) # windows-footgun: ok — POSIX only, the nt branch returned above
|
|
|
|
|
|
def _run_cli_killing_process_group(cmd, code, env, timeout):
|
|
"""Run the CLI in its own process group and kill the whole group on timeout.
|
|
|
|
``subprocess.run`` only kills the direct child on ``TimeoutExpired``; a grandchild that
|
|
inherited the stdout/stderr pipes (browser_harness daemon / Chrome helper) is orphaned
|
|
still holding them, and on Windows ``run()``'s unbounded post-kill ``communicate()`` then
|
|
blocks on pipe EOF forever — so the tool call, plus its activity heartbeat, wedges (#106244).
|
|
"""
|
|
proc = subprocess.Popen(
|
|
cmd, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE,
|
|
text=True, encoding="utf-8", errors="replace", env=env, **_group_popen_kwargs(),
|
|
)
|
|
try:
|
|
stdout, stderr = proc.communicate(input=code, timeout=timeout)
|
|
except subprocess.TimeoutExpired:
|
|
_kill_cli_process_group(proc)
|
|
with contextlib.suppress(subprocess.TimeoutExpired):
|
|
proc.communicate(timeout=_POST_KILL_DRAIN_S)
|
|
raise
|
|
return subprocess.CompletedProcess(cmd, proc.returncode, stdout, stderr)
|
|
|
|
|
|
def browser_exec(code: str, session: str = "", timeout_s: int = _DEFAULT_TIMEOUT_S,
|
|
task_id: Optional[str] = None, local: bool = False):
|
|
"""Run Python code through the browser-use CLI, and return its output"""
|
|
from agent.redact import redact_sensitive_text
|
|
from tools.registry import tool_error, tool_result
|
|
if not code or not code.strip():
|
|
return tool_error("No code provided. Pass Python that uses the pre-imported helpers, e.g. new_tab(\"https://example.com\") then print(page_info()).")
|
|
|
|
blocked = _blocked_url_in_code(code)
|
|
if blocked:
|
|
return tool_error(blocked)
|
|
|
|
cmd = _find_cli()
|
|
if not cmd:
|
|
return tool_error("browser-harness is missing from Hermes's Python environment. "
|
|
"Run `hermes update` to re-sync it.")
|
|
|
|
env = _base_subprocess_env()
|
|
if session:
|
|
if not _SESSION_RE.match(session):
|
|
return tool_error(f"Invalid session name {session!r}: use 1-64 letters, digits, "
|
|
"dashes, or underscores (e.g. 'r7k2').")
|
|
env["BU_NAME"] = session
|
|
route_err = _route_backend(env, session, task_id, bool(local))
|
|
if route_err:
|
|
return tool_error(route_err)
|
|
bot_desktop_browser = bool(env.pop(_BOT_DESKTOP_BROWSER_SENTINEL, None))
|
|
|
|
# SHARED browser (/browser connect CDP override): pin each named session to its own tab (see
|
|
# _OWN_TAB_PREAMBLE). Private per-name browsers skip this — nothing to collide with.
|
|
private_browser = env.pop(_PRIVATE_BROWSER_SENTINEL, None) # always pop: never exported to the CLI
|
|
if session and not private_browser:
|
|
code = _OWN_TAB_PREAMBLE + code
|
|
|
|
workspace = _workspace_dir(task_id)
|
|
if workspace:
|
|
env["BH_AGENT_WORKSPACE"] = workspace
|
|
|
|
# BU_AUTOSPAWN makes the CLI start a Browser Use cloud browser when no local
|
|
# Chrome/CDP endpoint is reachable (their API key authenticates it)
|
|
if "BU_AUTOSPAWN" not in env and is_legacy_browser_use_cloud_config(_read_browser_cfg()):
|
|
env["BU_AUTOSPAWN"] = "1"
|
|
|
|
timeout = _clamp_timeout(timeout_s)
|
|
started = time.time()
|
|
|
|
def dispatch() -> Dict[str, Any]:
|
|
_attach_vault_supervisor(env, task_id)
|
|
try:
|
|
return {"proc": _run_cli_killing_process_group(cmd, code, env, timeout)}
|
|
except subprocess.TimeoutExpired:
|
|
return {"error_result": tool_error(
|
|
f"browser-use exec timed out after {timeout}s. The daemon may still be working; retry "
|
|
f"with a larger timeout_s (max {_MAX_TIMEOUT_S}), or split the work into several calls that "
|
|
"append to workspace files — anything already written to the workspace is preserved."
|
|
)}
|
|
except OSError as e:
|
|
return {"error_result": tool_error(f"Failed to launch browser-use CLI: {e}")}
|
|
|
|
if bot_desktop_browser:
|
|
from tools.browser_tool_session import run_fenced
|
|
dispatched = run_fenced({"features": {"local": True}}, dispatch)
|
|
else:
|
|
dispatched = dispatch()
|
|
if "proc" not in dispatched:
|
|
if "error_result" in dispatched:
|
|
return dispatched["error_result"]
|
|
return tool_result(dispatched)
|
|
proc = dispatched["proc"]
|
|
|
|
# browser_vault_fill registers injected values with this forced model-egress
|
|
# boundary. Preserve raw stdout only for screenshot-path detection below.
|
|
result = {
|
|
"success": proc.returncode == 0,
|
|
"exit_code": proc.returncode,
|
|
"output": redact_sensitive_text(proc.stdout, force=True),
|
|
}
|
|
if workspace:
|
|
result["workspace"] = workspace
|
|
if session:
|
|
result["session"] = session
|
|
stderr = redact_sensitive_text((proc.stderr or "").strip(), force=True)
|
|
if len(stderr) > _STDERR_CAP_CHARS:
|
|
stderr = stderr[:_STDERR_CAP_CHARS] + "\n… (stderr truncated)"
|
|
if stderr:
|
|
result["stderr"] = stderr
|
|
screenshot = _find_screenshot(proc.stdout, started)
|
|
if screenshot:
|
|
result["screenshot_path"] = screenshot
|
|
native = _native_screenshot_result(result, screenshot)
|
|
if native is not None:
|
|
return native
|
|
return tool_result(result)
|
|
|
|
|
|
_HEADER_BASE = (
|
|
"Drive a real web browser via the Browser Use CLI: `code` runs as full Python (stdlib available) "
|
|
"with pre-imported browser helpers; stdout comes back in the result. Start `code` with a one-line "
|
|
"comment describing the step for the user in plain language, max 60 chars "
|
|
"(e.g. `# Searching Amazon for paper towels`) — the UI shows it as the step label.\n\n"
|
|
"STATE: the browser session and workspace persist across calls; Python variables do NOT (fresh "
|
|
"interpreter each call). The workspace dir is $BH_AGENT_WORKSPACE (also `workspace` in every result); "
|
|
"functions defined in agent_helpers.py there are auto-imported into every call. For multi-item tasks "
|
|
"('all N products / every entry'), append each batch to a JSON/CSV file in the workspace, then read it "
|
|
"back and aggregate in code — dedupe/count/sort with Python, not in your head — and verify the "
|
|
"collected count against what was asked before answering.\n\n"
|
|
"Batch each sub-procedure (navigate, wait, extract, act) into one call — do not spend a call per "
|
|
"action — but for long extractions prefer several medium calls that append to workspace files over "
|
|
"one giant call, so progress survives timeouts."
|
|
)
|
|
|
|
_HEADER_VISION = (
|
|
" Screenshots are attached to your context automatically: when the exec output contains a "
|
|
"capture_screenshot() path, the image arrives with this tool's result and you inspect it directly "
|
|
"with your own vision — never send browser screenshots to a separate vision tool."
|
|
)
|
|
|
|
_HEADER_TEXT_ONLY = (
|
|
" Your model cannot view images, so work text-first: page_info() for state, js() for "
|
|
"reading/extracting DOM text, fill_input(selector, text) for inputs, and "
|
|
"js(\"document.querySelector('…').click()\") for clicks — skip the screenshot-driven workflow described below."
|
|
)
|
|
|
|
# Appended when the local engine is Lightpanda: no graphical renderer, and one CDP
|
|
# connection holds one page — a second Target.createTarget fails with
|
|
# TargetAlreadyLoaded (drop the new_tab() sentence once lightpanda-io/browser#1962 lands).
|
|
_HEADER_LIGHTPANDA = (
|
|
" The local engine is Lightpanda (no graphical renderer, one page per session): capture_screenshot() "
|
|
"is unavailable, so work text-first; navigate with new_tab(url) exactly once, then goto_url(url) for "
|
|
"every later navigation — a second new_tab() fails with TargetAlreadyLoaded."
|
|
)
|
|
|
|
# Pinned quick-reference for the CLI's pre-imported helpers, replacing the live
|
|
# ``browser-use skill`` fetch (uncontrolled third-party text in every schema: version
|
|
# drift, supply-chain exposure, byte-unstable prompt). A/B benchmarked ~equal.
|
|
_HELPERS_DIGEST = (
|
|
"\n\nHELPERS (pre-imported): new_tab(url) opens/navigates (use for the FIRST navigation), goto_url(url) "
|
|
"navigates the current tab, wait_for_load() after navigation, page_info() summarizes the current page "
|
|
"state, js(expr) evaluates a JS expression and returns its value (js('document.title'); wrap function "
|
|
"bodies as js('(() => {...})()') — a bare '() => {...}' returns the function itself, uncalled), "
|
|
"fill_input(selector, text) types into inputs, click_at_xy(x, y) clicks viewport coordinates, "
|
|
"capture_screenshot() saves and prints a screenshot path, cdp('Domain.method', **kwargs) is raw CDP — "
|
|
"cdp('Accessibility.getFullAXTree')['nodes'] lists every element's role/name/backendDOMNodeId (filter "
|
|
"in Python before printing; it is thousands of nodes), then cdp('DOM.getBoxModel', backendNodeId=n) "
|
|
"gives click coordinates. ensure_real_tab() recovers from a stale/internal tab. Login walls: never guess "
|
|
"credentials; see the vault note below if present, otherwise stop and ask the user."
|
|
)
|
|
|
|
|
|
def _description_header() -> str:
|
|
"""Header tailored to whether the active model can see images natively"""
|
|
if _lazy_call("tools.browser_tool_lightpanda_fallback", "lightpanda_engine_status", (False, ""),
|
|
"lightpanda engine status unavailable")[0]: # no screenshots, whatever the model sees
|
|
return _HEADER_BASE + _HEADER_TEXT_ONLY + _HEADER_LIGHTPANDA
|
|
vision = _lazy_call("tools.vision_tools", "_should_use_native_vision_fast_path", False, "")
|
|
return _HEADER_BASE + (_HEADER_VISION if vision else _HEADER_TEXT_ONLY)
|
|
|
|
|
|
def _dynamic_schema_overrides() -> dict:
|
|
overrides: dict = {"description": _description_header() + _HELPERS_DIGEST}
|
|
# ``local`` exists ONLY when the user consented to real-profile browsing — everyone
|
|
# else's schema carries zero extra surface. The caller memoizes on config.yaml mtime,
|
|
# so toggling consent applies next session, not mid-chat.
|
|
if _real_profile_consented():
|
|
props = dict(BROWSER_EXEC_SCHEMA["parameters"]["properties"])
|
|
props["local"] = {
|
|
"type": "boolean", "default": False,
|
|
"description": ("Drive the user's own local browser (a Hermes-managed copy of their real "
|
|
"default-Chromium profile, logins/cookies included) instead of the configured "
|
|
"cloud browser backend. Use when the user asks to act as themselves — their "
|
|
"accounts, their sessions. No-op when the backend is already local. Default false."),
|
|
}
|
|
overrides["parameters"] = {**BROWSER_EXEC_SCHEMA["parameters"], "properties": props}
|
|
return overrides
|
|
|
|
|
|
BROWSER_EXEC_SCHEMA = {
|
|
"name": "browser_exec",
|
|
# Static fallback description, used only when the managed CLI is unavailable
|
|
"description": (_HEADER_BASE + _HELPERS_DIGEST
|
|
+ "\n\n(The browser-use CLI is not installed yet. Install it with `hermes tools` (Browser Automation → Browser Use).)"),
|
|
"parameters": {
|
|
"type": "object",
|
|
"properties": {
|
|
"code": {"type": "string", "description": "Python code to execute using the pre-imported browser helpers. Use print(...) for any data you need back."},
|
|
"session": {"type": "string", "description": "Named isolated browser session — its own daemon and (on cloud backends) own browser, so concurrent tasks don't share tabs. Reuse the same name on every related call; omit for the shared default session."},
|
|
"timeout_s": {"type": "integer", "default": _DEFAULT_TIMEOUT_S,
|
|
"description": f"Max seconds to wait for the code to finish (default {_DEFAULT_TIMEOUT_S}, max {_MAX_TIMEOUT_S})."},
|
|
},
|
|
"required": ["code"],
|
|
},
|
|
}
|
|
|
|
|
|
# browser_exec is additionally gated at tool-definition time — sessions whose toolsets
|
|
# lack ``terminal`` never see it (model_tools._compute_tool_definitions). check_fn only
|
|
# answers "is Browser Use mode configured"; surface policy lives with the session.
|
|
from tools.registry import registry
|
|
|
|
registry.register(
|
|
name="browser_exec",
|
|
toolset="browser-use",
|
|
schema=BROWSER_EXEC_SCHEMA,
|
|
handler=lambda args, **kw: browser_exec(
|
|
code=args.get("code", ""), session=args.get("session", "") or "",
|
|
timeout_s=args.get("timeout_s", _DEFAULT_TIMEOUT_S), task_id=kw.get("task_id"),
|
|
local=bool(args.get("local", False)),
|
|
),
|
|
check_fn=is_browser_use_cli_mode,
|
|
dynamic_schema_overrides=_dynamic_schema_overrides,
|
|
emoji="🌐",
|
|
)
|