Files
Teknium 399d956903 feat(bot_desktop): Bot Screen, computer_use and the browser run inside the terminal backend (#121169)
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends

The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.

docker/sandbox-desktop.Dockerfile is that base plus:
  - the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
    zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
  - the exact package set the Hermes -desktop image installs (TigerVNC,
    Xfce components, dbus, xauth, fonts)
  - Playwright's headed Chromium (same build as the -desktop image)
  - cua-driver 0.28.2 from its pinned release tarball

No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.

docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.

* feat(docker): sandbox desktop base on python3.13-nodejs26

Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.

* feat(docker): bake agent-browser into hermes-sandbox:desktop

The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.

* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend

A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.

Now the desktop lives where the terminal lives:

- tools/environments/streams.py: one primitive per spawn-per-call backend, the
  local argv prefix that runs its remainder inside the sandbox with stdio open
  (`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
  vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
  the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
  prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
  gateway). `auto` follows the backend; a sandbox that cannot host a screen
  REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
  lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
  there; the host driver's runtime contract is irrelevant then; check_fn is
  true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
  with the daemon, socket dir and profile inside the sandbox; screenshots are
  fetched back so MEDIA: paths keep working; recycle closes the sandbox
  daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
  in a sandbox, a unix socket otherwise.

Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.

* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only

DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).

* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read

Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.

* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order

Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host

sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.

* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox

Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).

* chore: retrigger CI (zero-job dispatch failure)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore(config): template stamps v47 and shows the new default sandbox image

The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.

* feat(sandbox): the default image change is a decision, not a surprise

A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.

Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.

Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.

Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.

* fix(config): both plain defaults that preceded the desktop sandbox image are template copies

main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.

* ci(docker): build the sandbox image on release/dispatch, not every main push

Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.

* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta

The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.

* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)

* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)

* fix(bot_desktop): placement is the authority; sandbox screen survives restarts

Review findings on the sandbox-hosted Bot Screen, each reproduced live first.

Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.

Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.

SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.

CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.

pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.

Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.

Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.

Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.

* docs(bot-screen): no literal tmp path in the profile-location note

* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)

* fix(bot_desktop): adopting a screen the sandbox kept records the marker

Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.

Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).

* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)

* docs(bot-screen): what the Apptainer path inherits from the image and what it does not

* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 03:34:07 -07:00

25 KiB
Raw Permalink Blame History

title, sidebar_position
title sidebar_position
Bot Screen 17

Bot Screen

On a headless Linux gateway host (a server, a cloud VM, Hermes Cloud) each bot gets its own desktop: an Xfce screen the bot's computer_use and headed browser act on, streamed live into Hermes Desktop. Watch what the bot does, take over when it hits a login, 2FA prompt, CAPTCHA or payment step, then hand control back and let it continue with the session you just signed in to. The bot keeps working after you close the app or turn off your laptop; the screen lives on the gateway host, not on your machine. If the gateway runs its terminal in a sandbox (terminal.backend: docker, ssh or singularity), the screen lives inside that sandbox instead, alongside the shell, so the bot's computer_use and browser never act outside the boundary you drew (see Where the screen runs).

Every Hermes profile ("bot") has its own screen, its own browser profile and its own cookies. Screens are work surfaces, not security boundaries: the bots share the host's user account, files and network (the same model as other hosted-agent products).

Threat model. The screen's RFB socket, the X display, the browser profile and the control-lease file all belong to the gateway's OS user. Any process running as that user — another bot on the same host, and the bot's own terminal tool included — can reach them directly, bypassing the pane and the lease. The lease is a tool-level fence on computer_use and the browser tools, not an OS one. One boundary is wider than the OS user: Chromium's DevTools port (the dock's Browser and every agent-browser launch advertise one on loopback so the agent can attach) is reachable by any local user on the host, and Chromium offers no per-user restriction for it. Running each bot as its own OS user is out of scope; if that isolation matters to you, or the host has untrusted local users, put the bots on separate hosts. Two timing details worth knowing: the WebSocket bridge caches its lease decision for up to 250 ms between re-reads of the lease file, so a takeover made by another process is enforced within that window (the bot's tool results are voided by the lease epoch regardless of the window). And the viewer's single-use, 30-second display_ticket travels as a URL query parameter on purpose — noVNC cannot negotiate WebSocket subprotocols, so a header is not an option — which means a reverse proxy's access log may record an already-spent ticket.

Requirements

  • The gateway host runs Linux. macOS and Windows hosts already have a real display; the pane is not offered there.

  • TigerVNC's Xvnc and the Xfce core components are installed on the host. Nothing installs them silently: hermes update and fresh installs leave every machine as it is. When they are missing the Screen pane in Hermes Desktop shows Install on host — one click runs the package manager on the gateway host (it asks for that host's sudo password in a masked card; the password goes to that host only and is never stored) and streams the log. When Hermes itself runs as root — the usual case in a container — the installer runs the package manager directly, with no sudo and no password card. When it is not root and the host has no sudo at all, the pane and the CLI print the exact install command for you to run on the host instead of showing a card. The official Docker image (nousresearch/hermes-agent, which also powers Hermes Cloud) is that second case: the gateway runs as an unprivileged user and the image has no sudo, so the pane shows the apt-get line and an operator runs it once as root in the container (docker exec -u 0 <container> apt-get install -y …). Add chromium to that line if you want the dock's Browser icon; see Browser sessions below. From a shell, hermes computer-use screen status prints the exact line and hermes computer-use screen install runs it:

    Distro Packages
    Debian / Ubuntu tigervnc-standalone-server xfce4-panel xfwm4 xfdesktop4 xfce4-settings xfce4-terminal dbus-x11 x11-xserver-utils x11-utils xauth fonts-dejavu-core
    Fedora tigervnc-x11-server xfce4-panel xfwm4 xfdesktop xfce4-settings xfce4-terminal dbus-daemon xsetroot xset xdpyinfo xprop xorg-x11-xauth setxkbmap dejavu-sans-fonts
    Arch tigervnc xfce4-panel xfwm4 xfdesktop xfce4-settings xfce4-terminal xorg-xsetroot xorg-xset xorg-xdpyinfo xorg-xprop xorg-xauth xorg-setxkbmap ttf-dejavu

    On Fedora the Xvnc binary is in tigervnc-x11-server (not tigervnc-server-minimal) and dbus-run-session comes from dbus-daemon. Deliberately not the xfce4 metapackage: it pulls in the screensaver, power manager and polkit agent that lock or prompt a headless desktop.

  • Computer Use enabled for the bot (cua-driver installed).

  • Memory. Measured in the official image: the gateway idles at ~300 MB, Xvnc + Xfce add ~220 MB, and the headed Chromium a human opens during a takeover adds 0.5–1 GB (one page: ~550 MB). Plan on ~1.1–1.5 GB per open screen with a browser; the desktop alone is cheap, the browser is the cost. CPU is not a constraint (idle desktop ≈ 0.01 core, live streaming ≈ 0.03 core). The packages take ~930 MB of disk on Debian 13.

    Before starting a screen, Hermes checks that the host — or its container cgroup, whichever is tighter — has bot_desktop.min_free_memory_mb free (default 1536; 0 disables the check). Below that the pane shows why in place of Start screen and hermes computer-use screen start refuses; a screen already running is never taken down by this check. A screen nobody uses is stopped after bot_desktop.idle_stop_minutes (default 30) and comes back on the next use, so an instance pays for a desktop only while something is on it. Practical guidance for small instances: 4 GB runs the desktop, 8 GB is where a takeover with a browser is comfortable.

Baking the packages into a container image

An image for a hosted or unprivileged deployment cannot install anything at run time, so the packages have to be built in. CI publishes two variants of every version: the unsuffixed tags (:latest, :v*) without them, and the -desktop tags (:latest-desktop, :v*-desktop) with them. A hosted deployment (Fly Machines, Azure container instances) gets Bot Screen by pulling the suffixed tag; a build argument could not reach it anyway, since it never runs a build. Nothing in the provisioner selects -desktop yet, so a hosted instance still comes up slim; pulling the suffixed tag yourself works today.

Build your own only if you want the packages in a custom image. The official Dockerfile has an opt-in build argument, off by default so a plain docker build . stays lean:

docker build --build-arg HERMES_BOT_DESKTOP=1 -t hermes-agent:screen .

It adds TigerVNC, the Xfce components and the distro chromium (the sandbox fallback described under Browser sessions) — the apt layer measured ~930 MB on Debian 13. It adds no second Playwright browser: every image, slim or -desktop, already carries PM's pinned full Chromium, which can open a window. Nothing starts at boot; an image built this way costs no memory until a screen is started.

Using it

Every bot's computer is one click away in three places of Hermes Desktop:

  • Bots → a bot → Scheduled Jobs: the bot's screen is the hero at the very top of the pane, above the title and the routines: a live preview of the desktop (refreshed every few seconds while the pane is visible) with who holds control; click the picture to expand into live access. While the screen is off or not installed the same box says so and offers Start / Install.
  • Bots → right-click a bot → Open Screen. The same menu has Open Screen when the bot uses it: with it checked, the Screen tab comes forward on the bot's first computer_use or browser call of a run, so you watch it work instead of finding out afterwards. Off by default, per bot. It raises the tab without taking your keyboard focus, never fires for replayed history, at most once every 30 seconds, and if you close the tab mid-run it stays closed until the bot's next run.
  • Sessions sidebar, grouped by gateway / profile: the same Screen box sits under each profile's header, so a profile's machine is reachable from its conversations too.
  1. Open the Screen with any of the entries above. The screen is off by default and nothing starts it for you: click Start screen in the pane, run hermes computer-use screen start on the host, or set bot_desktop.auto_start: true if you want a headless host to start the screen by itself on the bot's first computer_use call or first headed browser use (browser.headed: true) — off so that installing TigerVNC never yields a screen nobody asked for. A headed browser opens on the screen once it is running.
  2. The pane streams the bot's desktop. The chip in the header says who is in control: Bot is in control by default.
  3. Click Take over. The border turns red, your keyboard and mouse now drive the bot's screen. Sign in, solve the CAPTCHA, approve the payment.
  4. Click Hand back. The bot regains control and re-captures the screen before continuing. Closing the pane also hands control back. A dropped connection is different: if your laptop lid closes or Wi-Fi drops while you hold control, you keep it — the bot stays locked out of a screen you may be mid-login on — until you reconnect and hand back. If you come back after a reload and the pane still says a human holds control, a Hand back (force) button appears to clear it.

While you hold control, the bot's computer_use and browser tools are refused with human_has_control, captures included. This is a tool-level fence, not an OS one: the bot runs as the same user as its screen. Don't type secrets into a bot you wouldn't trust with them.

When the bot hits a step it should not do itself (a login, 2FA, a CAPTCHA, a payment) it says so in its reply and ends its turn; the ask reaches you in whatever chat you are on. Take over when you are ready, do the step, hand back, and tell the bot to continue. Nothing blocks on the bot's side while it waits: taking over is always yours to start, and a bot never holds a tool call open waiting for you.

Two viewers on one screen: the most recent Take over wins; the previous controller drops back to watching.

Browser sessions that survive the handoff

While the screen runs, the bot's browser tool and the dock's Browser icon are the same browser: the Chromium agent-browser drives, with one persistent user-data-dir per bot (<HERMES_HOME>/bot-desktop/browser-profile; set AGENT_BROWSER_PROFILE to pin your own — ~ expands, and a relative path such as pin resolves against that bot's HERMES_HOME, i.e. <HERMES_HOME>/pin). Click Browser during a takeover and you are in the bot's own windows and cookie jar; what you sign in to is what the bot uses afterwards and in every later session, until the site expires the login. Set browser.headed: true so the bot's own browsing is visible on the screen too.

The dock is seeded once, the first time the screen starts for a profile. The guard is the panel layout file <HERMES_HOME>/bot-desktop/xdg/xfce4/xfconf/xfce-perchannel-xml/xfce4-panel.xml: while it exists the launcher leaves the panel alone, so changing AGENT_BROWSER_EXECUTABLE_PATH or AGENT_BROWSER_PROFILE and restarting the screen does not re-pin the Browser icon. Delete that file and the dock is rebuilt on the next screen start from whatever is installed then.

Which Chromium the dock and the bot use: an explicit AGENT_BROWSER_EXECUTABLE_PATH wins; otherwise Hermes uses the PM-managed Chromium and falls back to a system chromium / google-chrome. A non-root user on a host with kernel.apparmor_restrict_unprivileged_userns=1 (Ubuntu 23.10 and later) gets the reverse order, because there the managed build cannot set up its sandbox and exits with FATAL: No usable sandbox!, while the distro's Chromium ships with an AppArmor profile that allows it. If the pick is wrong for your host, set AGENT_BROWSER_EXECUTABLE_PATH=/usr/bin/chromium (or your Chrome path) in the gateway's environment. The official Docker image points AGENT_BROWSER_EXECUTABLE_PATH at PM's pinned full Chromium, which can draw a window, so the dock's Browser icon uses it; the -desktop tags also carry the distro chromium — set AGENT_BROWSER_EXECUTABLE_PATH=/usr/bin/chromium on hosts that refuse the pinned build's sandbox. A Playwright headless shell is never used for the icon: when it is the only browser, the pane / screen status report no headed browser until you install a headed one (apt-get install chromium). The dock icon starts the browser with the same sandbox settings agent-browser uses in that container, so the human's Browser and the bot's browser are one and the same.

CLI

hermes computer-use screen status          # installed? running? who holds control?
hermes computer-use screen start           # start this profile's screen
hermes computer-use screen stop            # stop it; refuses while a human holds control
hermes computer-use screen stop --force    # ...unless you say so (also frees a stuck lease)
hermes computer-use screen install [-y]    # apt/dnf/pacman the packages
hermes -p research computer-use screen start   # another bot's screen

Where the screen runs

computer_use, the bot's browser and the screen they act on always run in the same place. bot_desktop.placement decides where:

terminal.backend placement: auto (default) What that means
local gateway host The terminal, the screen, the browser and computer_use all share the machine running the gateway.
docker, ssh, singularity inside the sandbox Xvnc + Xfce, Chromium and cua-driver run in the container / on the SSH host, spawned through the same docker exec / ssh channel the terminal uses. The pane streams the sandbox's screen; nothing of the host desktop is reachable.
modal, daytona, vercel_sandbox refused These backends cannot host a display yet. Rather than quietly running the screen on the host beside the sandbox you chose for the agent, Start explains and points at placement: gateway.

placement: gateway forces the pre-existing behaviour (screen on the gateway host even with a sandboxed terminal) as an explicit opt-in; placement: terminal forces the sandbox and errors when it cannot host one (with a local backend the terminal is the gateway host, so it resolves there).

Placement is policy, not a snapshot of what happens to be running. When the screen is placed in the sandbox, the first browser or computer_use call brings it up there on demand (no auto_start opt-in needed: the sandbox is the boundary you chose, and a screen inside it touches nothing outside it), and when it cannot come up the call fails with the reason. The host is never the fallback for a sandbox whose screen is down. A gateway restart does not lose the screen either: the host-side marker records which container owns it, so the restarted gateway re-attaches to a still-running sandbox, and Stop takes down the screen where it actually runs even if you changed placement in the meantime.

The sandbox image

The sandbox needs the desktop stack. nousresearch/hermes-sandbox:desktop is the default image for every container backend (Docker, Modal, Daytona, Singularity): the nikolaik/python-nodejs base (Python 3.13 / Node 26) plus TigerVNC, the Xfce components, a headed Chromium, agent-browser, cua-driver and the everyday tools that base lacked (jq, ripgrep, fd, tmux, rsync, sudo for the image's pn user). Its default user is root, like the old default, so shell workflows do not change. An image you pinned yourself is left alone, and the screen then tells you it needs this image or bot_desktop.placement: gateway:

terminal:
  backend: docker
  docker_image: nousresearch/hermes-sandbox:desktop

With a plain image the Screen pane reports the missing binaries and names this tag.

Under Singularity/Apptainer the same image is converted to a SIF (docker://nousresearch/hermes-sandbox:desktop); Dockerfile ENV survives the conversion, the image's USER does not: everything runs as you, so the browser profile lands in your $HOME inside the container, which is the persistent overlay by default. The instance runs --containall, so its temp dir (where the screen's runtime state lives) is Apptainer's session tmpfs, 64 MiB unless your admin raised sessiondir max size. This path is verified against the Apptainer documentation, not exercised live.

An SSH host is whatever you point the backend at, so it carries the stack itself: the same binaries (TigerVNC, Xfce, cua-driver, agent-browser with a Chromium it can find), reachable from a non-interactive login session. That last part is where a host built from the desktop image differs from docker exec: a Dockerfile ENV never reaches an ssh session, so the image also writes PLAYWRIGHT_BROWSERS_PATH to /etc/environment for PAM to apply. A host of your own needs the equivalent, or agent-browser reports "Chrome not found" over ssh while working in a local shell.

Upgrading from the previous default. A Docker sandbox you already have is kept, not replaced: when docker_image is unset and a persisted container runs another image (the old default, nikolaik/python-nodejs:python3.11-nodejs20), the terminal keeps using that container and you decide the switch. The interactive CLI asks once at startup; the Screen pane shows the same choice with Switch image / Keep current image; hermes config set terminal.docker_image nousresearch/hermes-sandbox:desktop is the same answer from any shell. Either answer writes terminal.docker_image, and a written image is a decision: the container is recreated on the next terminal call only when you chose the new image, and only once the new image has been pulled (a private or misspelled tag, or a registry outage, keeps your current container running instead of leaving you with nothing). What a switch means: files under /root and /workspace stay (they are host directories under ~/.hermes/sandboxes/), packages installed inside the container with apt/pip/npm -g are reinstalled on demand, and Python 3.11 virtualenvs need a rebuild on 3.13. Gateways and cron never decide; they keep the sandbox and log the notice. Configs that literally held the old default were unset on upgrade (that value was the template copied, not a pin). Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of the configured image, so an existing sandbox there is untouched and only a fresh one gets the new image. Desktop processes run as the image's unprivileged pn (uid 1000); Chromium gets --no-sandbox inside containers (Docker's seccomp profile denies the user namespaces its own sandbox needs; the container is the sandbox).

Runtime state inside the sandbox (X socket, cookie, launcher log) lives under <sandbox tmp>/hermes-bot-desktop/<profile>/; the host keeps only a marker under <HERMES_HOME>/bot-desktop/. The browser profile (logins, cookies) lives in the desktop user's home inside the sandbox, ~/.hermes/bot-desktop/browser-profile, shared by the agent's browser and the dock's Browser icon. It follows the container's own persistence: kept across stops and restarts of a persisted container, gone with an ephemeral one or when you approve an image switch (the container's writable layer is what a switch replaces). It is deliberately not under the container's temp dir, which Docker mounts as a small tmpfs that is emptied on every stop. Screenshots the browser tools take are copied back to the host so MEDIA: paths keep working, the pane's thumbnail is grabbed inside the sandbox, and browser_exec / the vault autofill reach the sandbox's Chromium through a port forwarded over the same docker exec / ssh channel.

Configuration

bot_desktop:
  geometry: "1440x900"      # screen size; the viewer scales to fit the pane
  auto_start: false         # set true to start on the first computer_use call or headed browser use
  min_free_memory_mb: 1536  # refuse to start below this much free memory (0 = never check)
  idle_stop_minutes: 30     # stop a screen nobody used for this long (0 = keep it up)
  placement: auto           # auto | terminal | gateway — see "Where the screen runs"

auto_start is off by default. Start the screen from the Desktop's Screen pane (Start screen), from hermes computer-use screen start, or set the flag to true for a headless host that should bring its screen up the first time the bot calls computer_use or opens a headed browser (browser.headed: true) and no display is available.

State lives under <HERMES_HOME>/bot-desktop/ per profile (RFB Unix socket, Xauthority, launcher log, per-profile xfconf).

How it works

  • TigerVNC Xvnc is the X server and the RFB server in one process, per profile, listening only on a 0600 Unix socket. No TCP port, no VNC password: only processes running as the gateway's user can reach it (see the threat model above), and the gateway's WebSocket bridge is the authenticated way in.
  • Xfce starts component-wise (xfsettingsd, xfwm4 --compositor=off, xfdesktop, xfce4-panel) under a private D-Bus session, without xfce4-session, so nothing tries to lock the screen or reach logind.
  • Hermes Desktop bundles noVNC. It asks the gateway for a single-use ticket (display.observe) over its normal authenticated connection and opens a sibling WebSocket to /api/display/ws; the gateway splices the RFB stream through. Nothing new is exposed; the pane works over local, SSH, URL+token and Hermes Cloud connections alike.
  • Control lease. The gateway drops keyboard, pointer and clipboard messages from any viewer that does not hold the lease, at the RFB byte level; noVNC's view-only flag is only the UI hint. The same lease gates computer_use and the browser tools. It is a file under <HERMES_HOME>/bot-desktop/: no file means the bot holds control (a fresh profile); a file that exists but cannot be read or parsed fails closed — the bot is treated as locked out until the next successful hand-off rewrites it. Xvnc never pushes the screen's clipboard to viewers (-SendCutText=0), so watchers do not receive what the person in control copies; pasting into the screen still works.
  • Display binding. The launcher publishes DISPLAY, XAUTHORITY and the D-Bus address; every cua-driver and headed-browser spawn for that profile inherits them, so the bot never acts on a display a human is sitting at.

Troubleshooting

  • "Screen packages missing" — click Install on host in the pane, or run the printed install line on the gateway host (not on the machine running Hermes Desktop). The pane refuses a second install while one is running.
  • Screen starts then stops — read <HERMES_HOME>/bot-desktop/launcher.log.
  • Typing produces wrong characters during a takeover — the screen runs a US keymap so RFB keysyms and cua-driver agree, and noVNC sends raw keycodes (QEMU extended key events) once Xvnc offers them, so on a non-US physical keyboard layout-dependent keys (Y/Z, symbols) land as their US counterparts while you hold control. Type passwords with that in mind, or change the layout with setxkbmap on that DISPLAY.
  • Bot says human_has_control after you left — click Hand back in the pane (or Hand back (force) after a reload). From a shell, hermes computer-use screen stop --force releases the lease and stops the screen (without --force the command refuses while a human holds control, so a runbook can never yank a live takeover); hermes computer-use screen start brings it back with the bot in control.

Testing under WSL

WSL2 counts as a supported Linux host: screen status reports it as such and the pane is offered. One WSLg quirk gets in the way of the first start: WSLg

mounts /tmp/.X11-unix read-only, so Xvnc cannot create its display socket and dies with Cannot establish any listening sockets in launcher.log. Replace the mount with a writable directory before starting the screen:

sudo umount /tmp/.X11-unix  # no-tmp: ok — X11 socket directory, fixed by the protocol
sudo mkdir -p /tmp/.X11-unix && sudo chmod 1777 /tmp/.X11-unix  # no-tmp: ok — same

The mount comes back on the next WSL restart; repeat the two commands then.