- tools/browser_tool_install.py: keep pm-clean's frozen old-updater stub; main's
UTF-8 decode fix touched only the npx prefetch body it replaces.
- tests/hermes_cli/test_update_scoped_reconciliation.py: keep pm-clean's test
subset (catch-up rides the PM completion owner) and take main's gateway-less
host evidence (#120740): the updated seed that holds the host at a running
gateway, and the two gateway-less matrices for the source change that merged
cleanly into update_cmd_fleet.py.
The Desktop serve backend polls /api/fs/default-cwd, whose git-branch probe
captured git output with text=True and no encoding. subprocess then decodes
with locale.getencoding() - cp936 on zh-CN Windows - so a UTF-8 branch name or
git's localized stderr raised UnicodeDecodeError inside communicate()'s
_readerthread on every poll ([gateway-crash] ... 'gbk' codec can't decode).
Pin encoding='utf-8', errors='replace' on the remaining in-process captures
whose children emit UTF-8: the git branch probe, the self-repo guard's git
alias read, the Bot Mode DM transport (a Hermes CLI child, stdio forced to
UTF-8; builds on the salvaged errors='replace'), the agent-browser npx probe,
the lightpanda help probe, and the cua-driver stderr drain thread.
Fixes#83851
b5a9f44829 switched the shared computer_use config reader to
load_config_readonly. That reader also serves the daemon-launch paths
(overlay flag, Wayland child env), and tests/computer_use pins
hermes_cli.config.load_config as the seam at 13 sites; the deepcopy it
saves is noise next to an 80-500 ms capture. The ledger half of that
commit stays.
PROOF: tests/computer_use/test_cua_wayland_env.py and
test_cua_no_overlay.py::test_linux_x11_explicit_false_overrides_auto_detect
red on the previous head, green now (macOS-only
test_serve_process_disables_overlay fails identically on origin/main).
Drop the `_ax_max_elements_sent` side channel, its `getattr` default and the
`contextlib.suppress(Exception)` around the lazy `cua_backend` import:
`_cua_configured_ax_max_elements()` already swallows config errors, and the
same lazy import is bare in `_resolve_capture_windows`. `capture()` now reads
`_gws_args().get("max_elements", 0)` directly (one more cached
`load_config_readonly` hit); the vision lane still reports 0.
PROOF: tests/tools/test_computer_use_ax_walk_bound.py 5 passed; hardcoding
`ax_max_elements=0` in capture() turns `stub.capture("ax").ax_max_elements == 350`
red (observed), restored green. ruff clean; check-windows-footguns clean; real import ok.
tool.py inferred 'walk capped' by re-reading cua's private config through
_cua_configured_ax_max_elements() — a leaky abstraction that also cost a
config load per capture. The backend now records the max_elements it put on
the get_window_state call and returns it on CaptureResult.ax_max_elements
(defaulted field, 0 = unbounded / vision); _capture_view compares
len(cap.elements) >= cap.ax_max_elements > 0 and the getattr guards go
since the field is always set. Hint text and line order unchanged.
PROOF: tests/tools/test_computer_use_ax_walk_bound.py 5 passed (hint tests
now build CaptureResult(ax_max_elements=bound)); the new
_ax_max_elements_sent assertion is red with cua_backend_capture.py
reverted to HEAD~1 ('AttributeError: ... no attribute _ax_max_elements_sent',
1 failed 4 passed). Probe: 200/200 -> capped hint + 'full ' dropped;
199/200 and 8/0 -> full tree promised. check-windows-footguns --all clean.
_max_ledger_bytes() runs on every ledger append (right after ledger_enabled()
already deep-copied the config) and _computer_use_cfg() now runs twice per
computer-use capture; neither mutates the result, so use
load_config_readonly() as hermes_cli/config.py documents for read-only hot
paths (skips the deepcopy, ~half the cache-hit cost). Tests that patched
load_config now patch load_config_readonly as well (test-only edit).
PROOF: real import under PYTHONSAFEPATH shows load_config() ==
load_config_readonly() (same dict shape; defaults 5242880 / 200 resolve);
tests/tools/test_skill_ledger.py + test_computer_use_ax_walk_bound.py +
test_skill_ledger_delta.py -> 39 passed, 0 failed.
`_capture_view`'s hint promised a "full element tree saved to elements_file",
but once `computer_use.ax_max_elements` bounds the driver's walk the spill
holds only the first N nodes. When the capture returned at least the bound
(bound > 0), drop "full" and add "accessibility walk capped at N elements;
pass app= to narrow" so the model knows how to reach the rest.
`capture()` sent `get_window_state` with only `pid`, `window_id` and `session`, so the driver
walked the target's entire accessibility tree before Hermes trimmed the surfaced element list
to `_DEFAULT_MAX_ELEMENTS` (100) and spilled the rest to a cache file. Every node past the
first ~100 was paid for and discarded.
Measured on macOS (cua-driver 0.28.2, M-series), same window, bound fixed at 200:
| Target | Unbounded walk | Bounded |
|---|---|---|
| 1,444-node Chrome window | 540 ms | 83 ms |
| 456-node Finder window (pathologically slow AX surface) | 6.9 s | 0.6 s |
The bound is lossless for the response: the bounded element list is a *prefix* of the
unbounded walk (checked by role/label/depth at 200/400/600/1000), so nothing the model sees
changes.
The bound is internal and config-driven — `computer_use.ax_max_elements`, default 200,
`0` disables (driver default) — and deliberately NOT a model-facing schema parameter:
`max_elements` was removed from the tool schema on purpose, so a capture's surfaced window
stays at its fixed default.
(cherry picked from commit 9f07ff428bd1fcfbcc0e5089d834e1f33d8d1b76)
The first cut appended the recovery only in `_tcc_row`, which is built by the
0.10 fallback probes. Drivers that serve health_report (0.22+, the ones a
stale row actually bites) render their own tcc_* rows untouched, so `doctor`
never showed it. Apply the hint at the report seam like the display-count
guard so both paths get it.
Also: one CUA_DRIVER_BUNDLE_ID (permissions.py is the leaf; the daemon
imports it), the CLI derives the field set from the same table, and the
unverified "--upgrade repairs stale rows" claim is dropped from hint + docs
(the user hitting this is already on a current driver).
`hermes computer-use permissions status` and the doctor's fallback `tcc_*`
rows told a user whose System Settings toggle already showed CuaDriver ON to
"grant it in System Settings" — the one step that cannot help. macOS keys the
TCC row to the app's code-signing requirement; a row written for an earlier
CuaDriver build stops matching after a driver update and flipping the toggle
does not rewrite it (trycua/cua#3170, repaired by the cua-driver >= 0.22
installer on update). The daemon then reports Accessibility / Screen Recording
false while the pane shows ON, and every click silently no-ops (#99732).
`permissions.stale_tcc_grant_hint(*missing)` renders the reset for exactly
the missing services (`tccutil reset Accessibility|ScreenCapture
com.trycua.driver`, then `hermes computer-use permissions grant`); both
surfaces append it only when a grant is actually reported False. Docs: the
computer-use page still said bounded/unrestricted daemons run "under the
Hermes host identity" — they launch through CuaDriver.app since #95381 — and
gains the stale-row troubleshooting entry.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
A hand-written Linux daemon unit (systemd user unit / XDG autostart entry
running `cua-driver serve`) can be dead for days — crash loop, stopped, never
started — while `hermes computer-use doctor` and `status` report a healthy
binary: the runtime contract only checks the binary (`manifest`), and nothing
ever connected to the daemon socket (#114748).
- tools/computer_use/cua_backend.py::cua_daemon_listening — socket-level
liveness via `cua-driver status [--socket PATH]` (rc 0 = a daemon answered,
"not running" = dead, anything else = unknown). Never raises.
- tools/computer_use/doctor.py::cua_daemon_units — one scan of the units that
run cua-driver (kind, unit, exec target, runs `serve`, `--socket` path with
`%h` expanded); the pruned-Exec guard now filters that list instead of
re-scanning.
- doctor.py::_apply_daemon_liveness_guard — per `serve` unit: `pass` when its
socket answers, `fail` (degrading `ok`) when not, with the hint that a driver
reinstall does not start the daemon. Unconfigured daemons are never probed:
on Linux the MCP runtime needs none, so a silent default socket is normal.
- `hermes computer-use status` prints the dead-daemon line and exits 1.
- docs: the Linux daemon-unit paragraph under the doctor section.
Not changed: _maybe_repair_runtime_contract. The repair is gated on the
binary-level contract only; a dead daemon never enters it, so a reinstall was
never triggered by the daemon (the `.release_installed/<version>` marker the
reporter saw is written by cua-driver itself on any first run of the binary —
observed via strace of `cua-driver status` under a fresh HOME).
Supersedes #114928 (@Finn763): same finding, ~600-line implementation with a
new module and a repair-gate rewrite; this is the ~90-line version on the
existing doctor seams.
systemd --user and XDG autostart honour $XDG_CONFIG_HOME; scanning only
~/.config left hosts that set it with the same silent green doctor the PR
removes. One invariant test, red before.
Linux has no managed cua-driver autostart, so daemon units are hand-written
against a concrete release directory. The installer prunes all but the last
five release dirs, so a versioned Exec reference crash-loops with 203/EXEC
after every upgrade while every binary-level check stays green (#114748).
doctor now scans systemd user units and XDG autostart entries for cua-driver
Exec targets under packages/releases/<version>/ that no longer exist and
appends a fail check with the packages/current recovery hint, downgrading an
ok overall to degraded (#114748).
After the shared _accepts_tool_result_images gate the capture route still
demanded _lookup_supports_vision(...) is True, so a whitelisted provider with a
catalog-unknown model (proxy alias) went native in vision_analyze and aux in
computer_use. The route is now the negation of the shared gate; the redundant
catalog lookup helpers are gone.
The capture route (`tools/computer_use/vision_routing.py`) opened the native
multimodal lane only for providers on `_TOOL_RESULT_MEDIA_PROVIDERS` (or a
profile `supports_vision` flag), while the `vision_analyze` fast path also
accepted a catalog vision hit. For deepseek/deepseek-flash the two gates
returned opposite verdicts, so a screenshot paid a same-model auxiliary round
trip and the main model acted on prose instead of pixels, while vision_analyze
embedded the image natively (#115248).
`tools.vision_tools._accepts_tool_result_images(provider, model, cfg)` is now
the single predicate (profile veto first, then provider tool-result media OR
capability lookup) and both lanes call it. The lookup keeps the full
`_lookup_supports_vision` order (config override -> catalog -> local probes ->
profile), so Ollama/local VLMs are unaffected.
Supersedes #115276 (@Tranquil-Flow), whose direction this follows; its
models.dev-only lookup in the fast path would have dropped the local-runtime
and Ollama probes.
hermes_cli.runtime_paths (venv generations, selection, activation) moves to
pm.environments, and gains venv_bin_dir / venv_python / project_python. Every
in-tree caller asks pm for an interpreter now; pm no longer reaches back into
hermes_cli for its own environment layout (pm.packages, pm.extras, pm.ensure,
pm.paths imported hermes_cli.runtime_paths). The three open-coded
"Scripts/python.exe or bin/python" ladders in pm collapse onto venv_python.
hermes_constants.venv_python_path / venv_bin_dir and hermes_cli.runtime_paths
stay as frozen-updater-surface shims only (tests/compat/old_updater_surface.json).
To keep the boot path light, pm/__init__ resolves its facade lazily (PEP 562)
and pm.registry loads the built-in package definitions on first read instead of
at import: `import hermes_bootstrap` now loads pm + pm.environments only (25ms,
was 37ms with the eager facade dragging in the downloader). The stripped-payload
fixtures that ship only pre-import files keep working for the same reason.
Also restores two frozen-surface re-exports the F401 sweep dropped
(banner._github_compare_behind, cua_backend.resolve_cua_driver_cmd).
- 15 `MERGE-CHECK:` conflict-resolution comments removed from prod code (two were
TODOs already done: the utf-8-sig sessions.json read lives in session_persistence,
the pm-aware cron script helpers in scheduler_script).
- 49 imports the branch left unused (ruff F401, none present at the merge base,
none inside PLUGIN-COMPAT blocks). update_cmd's frozen-surface re-exports are
trimmed to the names tests/compat/old_updater_surface.json actually lists under
hermes_cli.update_cmd; the rest resolve through hermes_cli.main.__getattr__.
- tools/environments/local_gitbash_probe.py: nothing imported it once _find_bash
delegated to pm.shell().
- Three try/except wrappers around calls that cannot raise (install_truststore,
get_hermes_home, and a duplicated except clause in supermemory).
Extract only backend admission/lifetime protection from #114565 (484b5d7), adapted to the existing DISPLAY identity and human-lease fences in #108914. No WSL selection, installer, config or UI changes.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Conflicts (all keep-both): tools/browser_tool.py takes main's scope-bound
passthrough + loopback NO_PROXY and still routes the env through the Bot
Desktop's desktop_env; i18n gains main's sudoDesc/sudoCommandUnavailable
beside our sudoInstallDesc; test_profiles keeps both sides' new tests.
Main replaced every *.request/*.respond event pair with real JSON-RPC server→client requests and
made Python the single source of the wire (tui_gateway/contracts → generated TS). Bot Screen now
speaks that dialect:
- tui_gateway/contracts/display.py declares the eight display.* methods, the display.install.sudo
server request and the four display.* events; the generated TS/OpenRPC is regenerated.
- The install sudo card is a `display.install.sudo` server request (app-level, empty session);
the desktop answers it through respondToServerRequest, so the respondMethod/origin plumbing
that pinned a reply to its socket is gone: a JSON-RPC response cannot land elsewhere.
- screen-connection.ts consumes the generated DisplayStatus/DisplayLease instead of hand-typed
copies (holder named by viewer_hash only; epoch always present).
- computer_use: the lease fence and main's screenshot-dedup session key ride the same read
handlers; the fence fires before a frame reaches the dedup cache.
- server.py keeps the display module registration a naive --theirs would have dropped.
On Windows the Hermes venv interpreter cannot CreateProcess a binary under
C:\Program Files\WindowsApps (WinError 5) even though the shell resolves it,
so `hermes computer-use doctor` died with a raw PermissionError traceback
from _open_mcp. Catch the spawn OSError and print what failed, why the tool
may still work (PATH resolves another copy), and the fix (reinstall outside
WindowsApps or HERMES_CUA_DRIVER_CMD), exit 2.
cua-driver 0.21 refuses a bare element_index:
click: bare element_index is not accepted; pass element_token,
or snapshot_id together with element_index
_maybe_attach_element_token gated solely on the trycua/cua#1961 capability
vocabulary. 0.21 stopped publishing per-tool capability sets — every tool
reports an empty set — while still accepting element_token in its input
schema. The gate therefore fails closed on 0.21.x and we send the bare
index, so the driver rejects the call.
The effect is total: every element-targeted click is refused, leaving
agents with only blind pixel coordinates. Observed against cua-driver
0.21.0 on X11, where a capture returned 762 elements with 762 tokens
cached and every subsequent click still failed with snapshot_id_required.
Check the live input schema first — supports_input_property() already
exists for exactly this, and its docstring notes it "deliberately inspects
tools/list rather than ... requiring a capability token the driver never
shipped". The capability check is retained as a fallback so drivers that
did ship the vocabulary are unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.
Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
Both process-global caches were keyed by the caller's session/task id alone, so under
gateway.multiplex_profiles two profiles using the same id — a shared `browser_exec session=`
name, or two Hermes sessions whose screens report the same DISPLAY — resolved to the FIRST
profile's cloud browser / cua-driver, and a command issued in one bot's chat could act on
another bot's screen.
The key now carries the routed profile's home key whenever a served-profile scope is active
(`get_hermes_home_override()` set), the same shape `tools/approval.py::_baseline_key` and the
camofox/cloud caches already use; outside a scope every key is byte-identical to before. The
computer_use lookup, install and release paths all go through one `_scoped_sid`, so a release
under profile B never stops profile A's driver; approval-bypass state keeps the bare session id.
Fixes#110032 (report by @wolfyy970, from @vandaimer's manual test on #108914).
`computer_use` carried two Bot Screen actions, `request_handoff` and `wait_for_human`, that let
the model ask for the screen and then block a tool call until the human handed it back. Both only
make sense when a person is guaranteed to be watching the Desktop pane; from Telegram, the CLI or
a cron worker the request lands nowhere and the wait burns minutes before returning. The blocking
wait also fought the sequential tool deadline (600 s default vs 420 s), so the model saw a timeout
before the wait returned while the thread stayed parked.
The agent now simply says what it needs in its reply and ends the turn; the user takes over from
the pane, does the step, hands back and tells it to continue. Take over / hand back and every fence
(actions refused with human_has_control, epoch-voided results, suppressed thumbnails) are
unchanged. Removes the handoff module, the schema entries and their host-conditional rewriter,
the lease's pending_handoff field and wait helpers, and the "Bot needs you" badge in the pane.
The frozen computer_use schema shipped `request_handoff` / `wait_for_human`
and the 'take over this screen from the Hermes Desktop app' copy to every host,
including macOS and Windows where there is no Bot Desktop: the model learned
two actions that can never succeed and a hand-off story the user cannot follow.
`tools.computer_use.schema.schema_for_host(supported=...)` is a pure function
returning the frozen schema when supported and the same schema minus the
handoff actions, their only parameters (`reason`, `grace`) and their copy
otherwise; one `_DYNAMIC_SCHEMA_REWRITERS` row applies it when
`tools.bot_desktop.runtime.is_supported_host()` is False (browser_navigate's
web-hint row is the precedent). Linux keeps the identical object, so
prompt-cache parity is untouched there.
Test: tests/tools/test_computer_use_schema_host.py — supported=False has no
handoff vocabulary and keeps every other action; supported=True is the frozen
schema (import error on bc36ddb5f969; the host is never faked).
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
Input actions never receive fence=, so a take-over then hand-back during
approval still clicks: assert_agent_may_act only asks if the holder is
human now. Compare admitted.epoch immediately before _dispatch.
Test uses a recording backend and does not patch _dispatch.
Addresses kvnloo on #108914.
(cherry picked from commit 97ce5d5a579ccc797c7ef50b6bb92465955a3277)
_get_backend keys its cache on permission mode only, so a cua-driver
cached before the Bot Desktop started (or restarted on another display
number) keeps acting on the old seat, or on no display. desktop_identity()
returns the DISPLAY a fresh spawn would get and backend_display_stale()
compares it with the identity recorded at cache time; tool.py (owned by
another lane) records the identity next to _backend_permission_modes[sid]
and detaches+stops a stale backend the way the /yolo mode swap does.
(cherry picked from commit 896304bcbb98b12f821c6e19f10b9f174fa08889)
wait_for_release polled until the full timeout (10 min by default) even
when the holder never left AGENT, so an unseen request_handoff blocked the
turn instead of letting the model chase the user in chat. After `grace`
seconds (param, default 60, capped at the timeout) with the handoff still
pending and no human holding, return {ok: false, code: no_takeover}. Once
a human holds the screen the full timeout still applies to the hand-back.
(cherry picked from commit 1402af036044a7191ec885fc8fa5c24206880909)
The epoch check ran only after _dispatch() returned, by which point the
capture path had already written the PNG to the media cache, spilled the
element tree to disk and routed the frame through auxiliary vision — the
human's screen had left the process before being "discarded". A fence
callable is threaded into _dispatch; capture and capture_after call it
the moment backend.capture() returns, and the post-dispatch check stays
for every other action.
(cherry picked from commit 70c09e5fad89d2817f13aea772bd845ab7cc4c8a)
Independent review of the PR head found the control boundary only held inside
one process and several claims the code did not back. Each item below was
reproduced, fixed, covered by an invariant test proven red without the fix, and
re-verified live on a real Xvnc/Xfce screen.
- Lease authority on disk. `lease.json` under an fcntl lock in the profile's
bot-desktop dir; every read goes to the file. `hermes serve` (viewer bridge),
the messaging gateway, a CLI turn and isolated workers now agree. Live: a
takeover in process A made `computer_use capture` in process B return
human_has_control; release in C made B work again.
- Takeover fences admitted actions. `handle_computer_use` re-checks the lease
under the dispatch lock and discards a result produced after the lease epoch
changed, so an action admitted before a takeover cannot picture what the human
typed during approval / backend start-up waits.
- Sudo reply pinned to its origin. `SudoRequest.origin` records the
(connection, profile) the card came from; SudoDialog answers through
`requestGatewayForAgent` on that socket, never the foreground gateway. A
password typed for host A can no longer reach host B. `sudo.expire` and
`display.install.sudo.expire` now tear the card down (the Desktop never
handled sudo.expire).
- Dock Browser IS the bot's browser. `tools/bot_desktop/browser.py` resolves one
identity — the Chromium agent-browser drives + a persistent per-profile
user-data-dir (`bot-desktop/browser-profile`) — and both sides use it: the
agent env gets AGENT_BROWSER_EXECUTABLE_PATH / AGENT_BROWSER_PROFILE, the
dock launcher gets the same exe + --user-data-dir. Live: the bot wrote
localStorage on http://127.0.0.1:8765 through agent-browser; a dock click and
a typed URL on the screen showed BOT-WROTE-THIS in Chrome for Testing.
- Safe display reuse. Allocation under a host-wide lock; a recorded number is
reused only when no live server holds it; the launcher never unlinks a lock
whose pid is alive. Live: A stopped, B took :20, A restarted on :21, B kept
running.
- Install worker keeps the caller's profile scope (copy_context carries the
HERMES_HOME override and the transport); the done event carries the requested
profile's status.
- Honest scope: `bot_desktop.auto_start` defaults to false (opt-in; Start lives
in the Screen pane); request_handoff no longer claims a Telegram/Discord
message was sent — the model relays the ask in its reply; docs match.
A bot running on a headless Linux gateway now gets its own desktop (TigerVNC
Xvnc + a minimal Xfce session, one per profile) that Hermes Desktop streams
live. The user can watch the bot work, take over to type a login / 2FA code /
CAPTCHA, and hand control back; the bot refuses every computer_use action
(screenshots included) while a human holds the screen, then resumes with the
session cookies the human just created.
Why this shape:
- The screen lives on the machine Hermes runs on, not in a vendor cloud browser,
so it works for any app the bot drives and keeps the session on the user's host.
- Xfce components are launched individually (xfwm4, xfce4-panel, xfdesktop,
xfsettingsd) under a private dbus session rather than xfce4-session/the
metapackage: no screensaver, power manager or polkit agent to lock or prompt
a headless desktop.
- Transport is raw RFB over a WebSocket beside /api/ws, authenticated with a
one-shot ticket minted through the already-authenticated RPC channel; noVNC
runs in the Electron renderer. The bridge parses the RFB client stream and
drops keyboard/pointer/clipboard (incl. QEMU Extended KeyEvent, which noVNC
switches to once Xvnc advertises it) from any viewer that does not hold the
lease, so viewOnly is enforced server-side, not by the client.
- One lease per profile (agent | human viewer) is the single truth for the RFB
bridge, the computer_use tool and the Desktop UI; taking control evicts other
viewers' input with close code 4000 control-taken.
- Auto-start happens only at the computer_use tool boundary (headless host,
packages present, bot_desktop.auto_start=true); env builders stay pure so
status probes and tests never spawn X servers. tests/tools/conftest.py pins
the binaries to "missing" for the same reason the browser-use fixture does.
Surfaces: Desktop (Bots → right-click → Open Screen; Take over / Hand back),
CLI (`hermes computer-use screen status|start|stop|install-deps`), tool
(`computer_use` actions request_handoff / wait_for_human), gateway RPCs
(display.status/start/stop/observe/lease.acquire/lease.release + display.lease
event), docs page user-guide/features/bot-screen.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.