* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
33 KiB
title, sidebar_position
| title | sidebar_position |
|---|---|
| Computer Use | 16 |
Computer Use
Hermes Agent can drive your desktop — clicking, typing, scrolling, dragging — in the background on macOS, Windows, and Linux. Your cursor doesn't move, keyboard focus doesn't change, and your virtual desktops / Spaces don't switch on you. You and the agent co-work on the same machine.
Unlike most computer-use integrations, this works with any tool-capable model — Claude, GPT, Gemini, or an open model on a local OpenAI-compatible endpoint. There's no Anthropic-native schema to worry about.
How it works
The built-in computer_use toolset is the recommended Hermes integration. It
speaks MCP over stdio to
cua-driver, an open-source background
computer-use driver. Each platform uses the appropriate accessibility +
input stack under the hood:
| Platform | Accessibility tree | Input dispatch |
|---|---|---|
| macOS | AX (private SkyLight SPIs) | SLPSPostEventRecordTo — pid-scoped, no cursor warp |
| Windows | UIAutomation | SendInput + PostMessage — no focus steal |
| Linux | AT-SPI (X11 + Wayland) | XTest (X11) / virtual-keyboard (Wayland) |
The result is the same on every platform: the agent can read the accessibility tree of any visible window AND post synthesized events without bringing it to front, switching virtual desktops, or moving the real OS cursor.
For the underlying contract — why background mode matters, the no-foreground invariant, click-dispatch internals — see cua.ai/docs/explanation/the-no-foreground-contract.
Which machine it drives
computer_use acts on the same machine the bot's screen lives on, never on
the machine running Hermes Desktop. On a gateway with terminal.backend: local that is the gateway host. With a sandboxed terminal (docker, ssh,
singularity) the driver runs inside the sandbox on the sandbox's own
display, so it can only ever touch what the terminal can; the sandbox image
must carry cua-driver (nousresearch/hermes-sandbox:desktop does). Modal,
Daytona and Vercel sandboxes cannot host a display yet, so with those
backends computer_use refuses unless bot_desktop.placement: gateway opts
into driving the host. Details: Bot Screen → Where the screen
runs.
Enabling
The driver ships with Hermes. cua-driver is pinned in pm/lock.json
and is a default PM package: the installers, a bare hermes pm install, and
hermes update install it on every macOS, Windows, and glibc Linux target
(cua-driver publishes no musl or Android build). The desktop app's bundle
carries it too. To leave it out, pass --skip-computer-use on POSIX or
-SkipComputerUse on Windows (or run hermes pm install --without cua-driver);
Hermes remembers the choice, and hermes pm install cua-driver undoes it.
If the download failed or you opted out earlier, any of these installs it:
hermes tools→ pick🖱️ Computer Use— installs the driver automatically if it's still missing.- Dashboard / desktop app → toggle the Computer Use toolset — if the driver is missing, the toggle kicks off the install in the background automatically (watch progress in the toolset panel).
Manual install / repair:
hermes computer-use install
This asks PM to prepare the pinned cua-driver package (verified against
pm/lock.json) — it does not run the upstream installer. Use
hermes computer-use status to verify the install.
Already have cua-driver? Hermes reuses it when it supports the 0.20 runtime
contract. During setup, toolset enablement, hermes update, and the first
computer_use call of a session, Hermes checks the local version and
manifest. It repairs an old or incomplete standard installation through
PM (at most once per session at runtime). A binary
selected with HERMES_CUA_DRIVER_CMD stays
under your control, so Hermes reports the incompatibility and leaves it
unchanged.
If you install Cua Driver first, cua-driver skills install installs Cua's
skill pack under ~/.cua-driver/skills/cua-driver. Hermes autodetection is a
planned cua-driver follow-up, so currently point Hermes at that directory or
symlink it into your skill space. You can also register raw Cua MCP tools as a
custom MCP server, but that is an alternative for users who need the low-level
interface. The built-in toolset provides Hermes actions, configuration,
approvals, and diagnostics.
After installing, regardless of which path you took, grant the platform-appropriate prereqs:
| Platform | Prereqs |
|---|---|
| macOS | System Settings → Privacy & Security → Accessibility + Screen Recording. Grant the identity named by hermes computer-use doctor (CuaDriver, com.trycua.driver, in every permission mode — the driver daemon always launches through CuaDriver.app). |
| Windows | None at install time. If you're driving over SSH (not RDP / console), you need the autostart pattern — see cua.ai/docs/how-to-guides/driver/windows-ssh for the Session 0 ↔ Session 1+ proxy. |
| Linux | A reachable display server: DISPLAY set for X11, or XDG_SESSION_TYPE=wayland. Wayland sessions need an XWayland bridge for capture. AT-SPI must be on (default on GNOME/KDE/Xfce). |
Then start a session with the toolset enabled:
hermes -t computer_use chat
or add computer_use to your enabled toolsets in ~/.hermes/config.yaml.
Permission modes and logged-in browser profiles
Hermes maps its existing approval UX onto cua-driver's immutable runtime modes. Permission mode and capability manifest approval are launch settings. They cannot change after the runtime starts:
| Hermes session | cua-driver mode | Human intervention |
|---|---|---|
| Manual or smart approvals (default) | standard |
Normal Hermes approvals; Cua stops at its protected boundary |
computer_use.permission_mode: bounded + reviewed manifest |
private bounded daemon |
You review and approve the capability manifest once, at launch |
--yolo, /yolo, or approvals.mode: off |
private unrestricted daemon |
One explicit Hermes risk acceptance; no runtime Cua prompts |
Browser work — including pages in a signed-in profile — goes through the
browser toolset (browser_exec), not computer_use. The former
computer_use.grant_existing_profile opt-in was removed along with the typed
browser route; a leftover key in config.yaml is ignored.
Bounded mode for repeatable automation
For recurring browser automation (cron jobs, scheduled research against an
authenticated app), bounded mode uses a capability manifest you review once:
# config.yaml
computer_use:
permission_mode: bounded
capability_manifest: ~/.hermes/cua-manifest.yaml
The manifest names the apps, browser profile kinds, allowed origins, and
typed tools the session may use (see the
cua-driver permission modes reference
for the format). Hermes launches a private runtime with
--capability-manifest ... --approve-capability-manifest; anything outside
the manifest fails closed inside cua-driver. A missing or unreadable manifest
fails loudly at session start rather than silently downgrading. Session YOLO
still overrides bounded for that one session.
On macOS, private-session daemons launch through the installed
CuaDriver.app bundle (so permission grants attribute to the driver's own
identity instead of resetting with every Hermes build), and Hermes verifies
the bundle's code signature — exact com.trycua.driver identifier and the
official signing team — before launching it. If you build cua-driver from
source (unsigned), opt in explicitly:
# config.yaml
computer_use:
allow_unsigned_driver: true # local driver development only
Each MCP transport owns a private lifecycle session inside its runtime. A
public session name is only a label for cursor identity and session-scoped
state. It does not select, share, or keep a runtime alive. Turning /yolo off,
resetting or closing the Hermes session, cancellation cleanup, or process exit
closes that transport session. Hermes also stops private runtimes that it
launched for bounded or unrestricted access. One Hermes
conversation cannot change another runtime's mode or grants. Bounded and
unrestricted modes use a private embedded daemon, launched through
CuaDriver.app on macOS (see above).
smart approval remains standard: an LLM classification cannot stand in for
a reviewed manifest.
YOLO/unrestricted mode does not protect against prompt injection or unintended input. Use it only in a disposable VM or with accounts and data whose full compromise you accept.
hermes computer-use doctor — your first triage stop
hermes computer-use doctor runs cua-driver's structured
health_report MCP tool and prints a per-check matrix. It's the single
fastest way to find out why an action isn't working.
$ hermes computer-use doctor
⚠️ cua-driver VERSION on darwin: degraded
✅ binary_version: cua-driver VERSION
✅ platform_supported: macOS 26.4.1 (arm64)
✅ session_active: MCP session is active.
❌ bundle_identity: Process has no CFBundleIdentifier.
→ Run the binary inside CuaDriver.app so TCC grants attribute correctly.
✅ tcc_accessibility: Accessibility is granted.
✅ tcc_screen_recording: Screen Recording is granted.
✅ ax_capability: AX is trusted and reachable.
✅ screen_capture_capability: ScreenCaptureKit reachable; 1 display(s) shareable.
- Exit code 0 when overall is
ok— everything's wired up. - Exit code 1 when
degradedorfailed— at least one check failed; the hint on each failure tells you what to fix. - Exit code 2 when the cua-driver binary itself isn't reachable.
Useful flags:
--include CHECK— run only the listed checks (repeat for multiple)--skip CHECK— skip a check (wins over--include)--json— emit the raw structured payload, same shape as thetools/call health_reportMCP response
The check matrix is platform-aware: bundle_identity / tcc_* are
skip on Windows + Linux because those concepts don't apply.
ax_capability checks AX on macOS, UIA on Windows, AT-SPI on Linux —
each with the right diagnostic hint when it can't reach.
On Linux, where the daemon is a hand-written systemd user unit or XDG
autostart entry rather than a managed autostart, doctor also reads those
units: a cua-driver ExecStart pointing at a pruned
packages/releases/<version>/ directory is reported as a failing
daemon unit (...) check (point it at ~/.cua-driver/packages/current/cua-driver),
and a unit that runs cua-driver serve gets a daemon (...) check that
connects to its socket — fail when nothing is listening (crash loop,
stopped, never started), pass when the daemon answers. Reinstalling the
driver does not start a daemon; systemctl --user status <unit> does.
hermes computer-use status prints the same dead-daemon line and exits 1.
The agent cursor and sessions
When the agent acts, you'll see a tinted overlay cursor glide
across the screen to where each click / type / scroll lands. The real
OS cursor never moves. The overlay shows where the agent is acting. Each
Hermes run declares a public cua-driver session name (something like
hermes-3a7b9c14d2e8). The name labels cursor identity and related state, so
concurrent runs and subagents get distinct cursors. The MCP transport owns the
private lifecycle session inside the runtime; the public name does not.
The overlay cursor is cosmetic — captures, clicks, and typing all work
without it. Hermes disables it automatically where it is a known failure
mode: macOS (idle CPU burn), headless Linux / WSL2 / containers, and
Linux X11 desktops (the overlay is a fullscreen always-on-top window
that can get stuck over every workspace after an unclean session end,
wedging desktop input). Linux Wayland and Windows keep the overlay. Set
computer_use.no_overlay: false in config.yaml to force the cursor on
(or true to force it off) on any platform.
Tune the cursor with cua-driver's CLI flags or the runtime
set_agent_cursor_style MCP tool — see
cua.ai/docs/how-to-guides/driver/personalize-cursor
for the full menu (built-in arrow vs teardrop silhouette, custom
SVG / PNG / ICO via --cursor-icon, runtime gradient colors, bloom
halo).
Going deeper — the cua-driver skill pack
Hermes keeps its wrapper skill (skills/autonomous-ai-agents/computer-use/SKILL.md)
focused on the Hermes-side computer_use workflow and action vocabulary. For
platform details, recording semantics, browser page interaction, and other
deep Cua behavior, install the skill pack that the cua-driver team ships and
maintains directly:
cua-driver skills install
The command links the pack into ~/.hermes/skills/cua-driver (Hermes is one
of the agents cua-driver skills status reports). The wrapper remains the
workflow layer: the pack documents the driver's own MCP vocabulary
(get_window_state, element_token, snapshot_id), which the computer_use
wrapper translates to for you — keep calling computer_use(action=...). The
pack contains:
| File | Topic |
|---|---|
SKILL.md |
The cross-platform core (snapshot invariant, no-foreground contract, click dispatch, AX-tree mechanics) |
MACOS.md |
macOS specifics: no-foreground contract, AXMenuBar navigation, SkyLight click dispatch, Apple Events JS bridge |
WINDOWS.md |
Windows specifics: UIA tree, UWP / ApplicationFrameHost hosting, Session 0 isolation, autostart pattern |
LINUX.md |
Linux specifics: AT-SPI tree, X11 / Wayland, terminal-emulator detection |
RECORDING.md |
Trajectory + video recording semantics |
WEB_APPS.md |
Browser-page interaction tips |
TESTS.md |
Replay-by-trajectory workflow |
These are platform deep dives, not duplicates of the Hermes skill —
when an agent reports "on Windows, my click landed on the wrong
element," it reads WINDOWS.md for the UIA / UWP context that
explains why and what to do differently.
cua-driver skills status shows what's installed and which agent
harnesses it's linked into. Today the autodetect list covers Claude
Code, Codex, OpenCode, OpenClaw, and Antigravity; Hermes
autodetection is planned as a follow-up in trycua/cua — until
then, run cua-driver skills install once and point your harness at
the resulting ~/.cua-driver/skills/cua-driver directory (or symlink
it into your usual skill space).
Quick example
User prompt: "Find my latest email from Stripe and summarise what they want me to do."
The agent's plan (this is the same shape on macOS / Windows / Linux — the model substitutes the platform's idiomatic shortcut and app name):
computer_use(action="capture", mode="som", app="Mail")— gets a screenshot of the email app with every sidebar item, toolbar button, and message row numbered.computer_use(action="click", element=14)— clicks the search field.computer_use(action="type", text="from:stripe")computer_use(action="key", keys="return", capture_after=True)— submit and get the new screenshot.- Click the top result, read the body, summarise.
During all of this, your cursor stays wherever you left it and the email app never comes to front.
Receiving the actual screenshot
Screenshots taken during computer control are normally internal — they exist so the model can see the screen, and the agent replies in text. But every image capture also saves a bounded, shareable copy under Hermes' image cache and reports its path, so on attachment-capable surfaces (Telegram, Discord, Desktop, and other gateway platforms) you can simply ask:
"Send me a screenshot of my screen."
and the agent delivers the real image as a native attachment, not just a description. On the CLI there is no attachment channel, so the agent gives you the saved file's path instead.
Only the 20 most recent capture files are kept, and screenshots are never sent automatically — only when you ask for one.
Whole screen vs. desktop surface
"Screenshot my screen" captures everything currently displayed — a composited grab of all visible windows, like pressing PrtScn. This image has no clickable elements, so to act on something in it the agent re-captures the specific app.
Asking for the desktop instead targets the OS shell surface itself — wallpaper, desktop icons, taskbar — with its clickable elements, so requests like "open the Recycle Bin on my desktop" still work.
Provider compatibility
| Provider | Vision? | Works? | Notes |
|---|---|---|---|
| Anthropic (Claude Sonnet/Opus 3+) | ✅ | ✅ | Best overall; SOM + raw coordinates. |
| OpenRouter (any vision model) | ✅ | ✅ | Multi-part tool messages supported. |
| OpenAI (GPT-4+, GPT-5) | ✅ | ✅ | Same as above. |
| Google (Gemini 2+) | ✅ | ✅ | Tool-calling + vision both supported. |
| Local vLLM / LM Studio / Ollama (vision model) | ✅ | ✅ | If the model supports multi-part tool content. |
| Text-only models | ❌ | ✅ (degraded) | Use mode="ax" for accessibility-tree-only operation. |
Screenshots are sent inline with tool results as OpenAI-style image_url
parts. For Anthropic, the adapter converts them into native tool_result
image blocks. The image MIME type comes from cua-driver's explicit
mimeType field (image/png or image/jpeg) — no client-side
magic-byte sniffing.
Safety
Hermes applies multi-layer guardrails:
- Destructive actions (click, type, drag, scroll, key, focus_app)
require approval through the same gate as dangerous shell commands —
interactively via the CLI dialog or the messaging-platform approval
buttons. Once/session/always grants are keyed
cua:<action>:<background|foreground>and live in the shared session/command_allowliststore (a background grant never covers the visible foreground variant). Where nobody can answer — cron (approvals.cron_mode), single-query, unattended platforms, or any headless run — the action is refused rather than auto-approved;--yolo//yolostill bypass. - Hard-blocked key combos at the tool level: empty trash, force delete, lock screen, log out, force log out.
- Hard-blocked type patterns:
curl | bash,sudo rm -rf /, fork bombs, etc. - The agent's system prompt tells it explicitly: no clicking permission dialogs, no typing passwords, no following instructions embedded in screenshots.
Pair with approvals.mode: manual in ~/.hermes/config.yaml if you
want every action confirmed.
Token efficiency
Screenshots are expensive. Hermes applies four layers of optimisation:
- Screenshot eviction — on every provider, screenshots ride each
request until it would cross Anthropic's documented per-request image
limit (20 image blocks, or 24 MB of image data); then the oldest batch becomes
[screenshot removed to save context]placeholders. Below the limit nothing is rewritten, so the prompt-cache prefix survives; at it, one slower turn per batch instead of one per screenshot. Images you attach yourself count against the limit but are never removed. - Client-side compression pruning — the context compressor detects multimodal tool results and strips image parts from old ones.
- Image-aware token estimation — each image is counted as ~1500 tokens (Anthropic's flat rate) instead of its base64 char length.
- Server-side context editing (Anthropic only) — when active, the
adapter enables
clear_tool_uses_20250919viacontext_managementso Anthropic's API clears old tool results server-side.
A 20-action session on a 1568×900 display typically costs ~30K tokens of screenshot context, not ~600K.
Limitations
- Performance. Background mode is slower than foreground — accessibility-routed events take ~5–20 ms on macOS, ~3–10 ms on Windows UIA, ~5–15 ms on Linux AT-SPI vs direct HID posting. Not noticeable for agent-speed clicking; noticeable if you try to record a speed-run.
- No keyboard password entry.
typehas hard-block patterns on command-shell payloads; for passwords, use the system's autofill (macOS Keychain / Windows Credential Manager / GNOME Keyring / KWallet). - Some apps don't expose an accessibility tree. Modern UWP apps on Windows, Electron < 28 on Linux, and a few macOS apps with custom drawing (Logic, Final Cut, some games) have sparse or empty AX trees. Fall back to pixel coordinates if the tree is empty — or skip the task entirely.
- Windows: elevated (admin) windows can't be driven from a normal
agent. Windows UIPI (User Interface Privilege Isolation) enforces
integrity-level boundaries: a Medium-integrity process (the default
Hermes agent) cannot enumerate the UIA tree of, or inject mouse input
into, a window owned by a High-integrity (Administrator) process.
Symptom:
capture(mode='som')returns 0 elements andclick(...)reports success while doing nothing, even though the screenshot renders fine (GDI capture sits below the integrity check). Keyboard events partially bypass UIPI, so Tab / Enter can still navigate an elevated dialog. This is an OS constraint, not a cua-driver bug — it affects every Windows automation stack. To drive elevated windows, run the Hermes agent itself at High integrity (launch from an elevated terminal); otherwise target non-elevated windows. - Windows:
hermes computer-use doctorfails with "Access is denied" while the tool works. A cua-driver installed underC:\Program Files\WindowsAppscannot be executed by the Hermes venv interpreter (WinError 5 fromCreateProcess), even though the shell resolves the same binary fine. The doctor now reports this as a diagnosis instead of a traceback. Fix once: reinstall with the upstream installer (lands under your user profile) or setHERMES_CUA_DRIVER_CMDto a copy outsideWindowsApps. - Platform-specific deployment gotchas:
- macOS uses private SkyLight SPIs. Apple can change them in any OS update. Hermes warns when the installed cua-driver is older than the version it was tested against.
- Windows SSH sessions run in Session 0, which has no interactive desktop. Drive Hermes from inside the RDP / console session, or set up cua-driver's autostart Scheduled Task — windows-ssh has the recipe.
- Linux requires a reachable display server. Headless servers
get one from Bot Screen: a per-profile Xfce
desktop over TigerVNC, streamed into Hermes Desktop, where you can
take over for logins and 2FA. You start it from the Desktop's
Screen pane or
hermes computer-use screen start; it starts on first use (the firstcomputer_usecall or headed browser use) only whenbot_desktop.auto_start: trueis set (off by default). Pure Wayland sessions need an XWayland bridge for screen capture (cua-driver's Wayland inject path handles input independently).
For cross-platform GUI automation without the desktop overhead (and
without TCC / Session 0 / X11 setup), the browser toolset uses a
real headless Chromium and is the right answer for web-only tasks.
Configuration
Permission mode and manifest (see Permission modes above):
computer_use:
permission_mode: standard # standard (default) | bounded
capability_manifest: "" # capability manifest path, required for bounded
On Linux, native Wayland support remains an explicit opt-in. Hermes passes the
opt-in to every cua-driver process, including gateway sessions, only when that
process also has WAYLAND_DISPLAY:
computer_use:
native_wayland: true
Restart a running gateway after changing this setting.
Override the driver binary path (tests / CI / local builds):
HERMES_CUA_DRIVER_CMD=/path/to/your/cua-driver
Windows auto-start (opt-in)
On Windows, cua-driver can run from a per-boot Scheduled Task
(cua-driver-serve) so it is already listening when Hermes needs it. This
task is opt-in: by default Computer Use starts the driver on demand,
per session — exactly as on macOS and Linux — and no scheduled task is
registered when you install or enable the toolset (#97389).
Set this in config.yaml to opt in (the task is registered — or repaired —
the next time the driver is installed or the toolset is enabled):
computer_use:
autostart: true # default: false (on-demand; no scheduled task)
You need this when driving Windows over SSH: Session 0 has no interactive
desktop, so an on-demand driver cannot reach one
(windows-ssh has the
recipe). If the task exists but you want it gone, remove it with
cua-driver autostart disable (or schtasks /Delete /TN cua-driver-serve)
from an elevated shell — Hermes does not re-register it once
computer_use.autostart is false.
Swap the backend entirely (for testing):
HERMES_COMPUTER_USE_BACKEND=noop # records calls, no side effects
Telemetry
cua-driver ships with anonymous usage telemetry (PostHog) enabled by default
upstream. Hermes disables it for you — on every cua-driver invocation
(the MCP backend, status, doctor, and install) Hermes sets
CUA_DRIVER_RS_TELEMETRY_ENABLED=0 in the driver's environment.
To opt back in (let cua-driver use its own default and send telemetry), set
this in config.yaml:
computer_use:
cua_telemetry: true # default: false (telemetry off)
When it's on, hermes computer-use doctor reports telemetry: enabled;
when off (the default), it reports telemetry: disabled via CUA_DRIVER_RS_TELEMETRY_ENABLED.
Testing against a local cua-driver build
When you're developing cua-driver itself — or want to test an
unreleased fix — point Hermes at a binary you built from source instead
of the published release. Hermes resolves the driver with
shutil.which("cua-driver") and does not enforce
HERMES_CUA_DRIVER_VERSION, so a local build (reported as
0.0.0-local-*) is accepted as-is. Two approaches:
Option A — install-local (build + put it on PATH)
From your trycua/cua checkout, run the upstream local installer. It
builds the Rust backend in release mode and drops cua-driver into the
same install layout the production installer uses, adding its bin dir
to your PATH:
# Windows (PowerShell), from the cua repo root
./libs/cua-driver/scripts/install-local.ps1 -NoAutoStart
# macOS / Linux, from the cua repo root (defaults to a debug build without --release)
./libs/cua-driver/scripts/install-local.sh --release
- Windows stages the build under
%USERPROFILE%\.cua-driver\packages\…and junctions%LOCALAPPDATA%\Programs\Cua\cua-driver\bin(added to your User PATH) to it. macOS/Linux symlinkscua-driverinto~/.local/bin(override with--bin-dir <path>). -NoAutoStartskips registering thecua-driver-servelogon daemon — you don't need it for Hermes testing (see notes).
Then open a fresh shell (so the PATH change is visible) and confirm:
cua-driver --version # local builds report 0.0.0-local-release
# Windows: (Get-Command cua-driver).Source
# macOS/Linux: which cua-driver
Option B — point Hermes straight at the built binary (fastest loop)
Skip the install ceremony entirely: cargo build and set
HERMES_CUA_DRIVER_CMD to the resulting binary. Best for rapid
edit/build/test.
cargo build -p cua-driver # add --release for a release build; run from libs/cua-driver/rust
# Windows (.env)
HERMES_CUA_DRIVER_CMD=C:\path\to\cua\libs\cua-driver\rust\target\debug\cua-driver.exe
# macOS / Linux (.env)
HERMES_CUA_DRIVER_CMD=/path/to/cua/libs/cua-driver/rust/target/debug/cua-driver
Confirm Hermes is using your build
hermes computer-use statusprints the resolved binary path and version.hermes computer-use doctorconfirms the binary is reachable and exercises the full MCP path end-to-end.- In a session,
computer_use(action="capture")exercises the spawnedcua-driver mcpchild process.
Notes & gotchas
- Hermes spawns a
cua-driver mcpstdio proxy. In a normal session the proxy connects to (and may start) the standard machine daemon. In explicit Hermes YOLO, Hermes instead owns a privatecua-driver serve --embeddedchild and points the proxy at its private socket or named pipe. The Windows autostart/UIAccess pattern still matters for interactive Session 1+ input from SSH — see the Limitations section. - Locked binary on Windows. A running
cua-driver-servedaemon can holdcua-driver.exeand block an overwrite on rebuild.install-local.ps1renames the locked binary out of the way automatically; if youcargo buildmanually (Option B), stop it first withcua-driver autostart disable(orschtasks /End /TN cua-driver-serve). - Rebuild loop. After editing cua-driver source, re-run
install-local(rebuilds, restages, flips thecurrentjunction) for Option A, or just re-cargo buildfor Option B — no Hermes change needed either way. - Local builds skip the version check. Hermes warns when the
installed cua-driver is older than its per-OS tested baseline, but
exempts
0.0.0-local-*dev builds — so your local build never triggers that warning.
Troubleshooting
First action when anything's off: run hermes computer-use doctor.
The structured per-check matrix tells you (and any agent helping you
debug) exactly what's wrong.
Specific failure modes the doctor doesn't catch:
computer_use backend unavailable: cua-driver is not installed —
Run hermes computer-use install to fetch the cua-driver binary, or
run hermes tools and enable the Computer Use toolset.
Clicks seem to have no effect — Capture and verify. A modal you
didn't see may be blocking input. Dismiss it with escape or the close
button.
macOS: System Settings shows CuaDriver ON, but hermes computer-use permissions status / doctor report Accessibility or Screen Recording as
not granted — the stored grant is stale. macOS keys each permission row to
the app's code-signing requirement; a row written for an earlier CuaDriver
build stops matching after a driver update, and flipping the toggle does not
rewrite it. Reset the affected rows and re-grant:
tccutil reset Accessibility com.trycua.driver
tccutil reset ScreenCapture com.trycua.driver
hermes computer-use permissions grant
Element indices are stale — SOM indices are only valid until the
next capture. Re-capture after any state-changing action. The
wrapper carries opaque element_tokens for stale detection — you'll
see an explicit error rather than a wrong click.
"blocked pattern in type text" — The text you tried to type
matches the dangerous-shell-pattern list. Break the command up or
reconsider.
Empty captures on Linux — DISPLAY not set, or you're on pure
Wayland without an XWayland bridge. hermes computer-use doctor will
flag this as ax_capability: fail with a Set DISPLAY (X11)… hint.
Empty captures on Windows over SSH — You're in Session 0 (the services session). Drive from RDP / console directly, or set up the autostart pattern — see cua.ai/docs/how-to-guides/driver/windows-ssh.
See also
- Hermes-side skill —
skills/autonomous-ai-agents/computer-use/SKILL.md— teaches the Hermescomputer_useaction vocabulary; this is what the agent loads. - cua-driver skill pack — for platform-specific deep dives
(macOS no-foreground contract, Windows UIA + Session 0, Linux AT-SPI
- X11/Wayland, recording, browser pages), run
cua-driver skills installand readMACOS.md/WINDOWS.md/LINUX.md/RECORDING.md/WEB_APPS.md. Hermes autodetection is a planned follow-up; currently point Hermes at the installed pack directory or symlink it into your skill space.
- X11/Wayland, recording, browser pages), run
- cua.ai/docs — the cua-driver project's documentation:
- What is computer use? — concept intro
- The no-foreground contract — why background mode matters
- Install reference — cross-platform install details
- Personalize the agent cursor — built-in shapes, custom assets, runtime overrides
- Drive Windows over SSH — the Session 0 → Session 1+ autostart pattern
- Keep cua-driver running — autostart / daemon lifecycle
- Connect your agent — register cua-driver with various harnesses (Hermes among them)
- cua-driver source (trycua/cua)
- Browser automation for cross-platform web tasks where you don't need to drive native apps.