Commit Graph

30 Commits

Author SHA1 Message Date
ethernet
3f4b8840f1 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	pyproject.toml
#	tests/fixtures/resolution_allowlist.json
2026-09-23 07:36:31 -04:00
teknium1
88bd29213b refactor(bot-screen): drop the agent-side handoff actions; takeover is always human-initiated
`computer_use` carried two Bot Screen actions, `request_handoff` and `wait_for_human`, that let
the model ask for the screen and then block a tool call until the human handed it back. Both only
make sense when a person is guaranteed to be watching the Desktop pane; from Telegram, the CLI or
a cron worker the request lands nowhere and the wait burns minutes before returning. The blocking
wait also fought the sequential tool deadline (600 s default vs 420 s), so the model saw a timeout
before the wait returned while the thread stayed parked.

The agent now simply says what it needs in its reply and ends the turn; the user takes over from
the pane, does the step, hands back and tells it to continue. Take over / hand back and every fence
(actions refused with human_has_control, epoch-voided results, suppressed thumbnails) are
unchanged. Removes the handoff module, the schema entries and their host-conditional rewriter,
the lease's pending_handoff field and wait helpers, and the "Bot needs you" badge in the pane.
2026-09-13 09:30:43 -07:00
teknium1
1fcb15f206 fix(bot-screen): computer_use only advertises the screen handoff where a Bot Desktop can exist
The frozen computer_use schema shipped `request_handoff` / `wait_for_human`
and the 'take over this screen from the Hermes Desktop app' copy to every host,
including macOS and Windows where there is no Bot Desktop: the model learned
two actions that can never succeed and a hand-off story the user cannot follow.

`tools.computer_use.schema.schema_for_host(supported=...)` is a pure function
returning the frozen schema when supported and the same schema minus the
handoff actions, their only parameters (`reason`, `grace`) and their copy
otherwise; one `_DYNAMIC_SCHEMA_REWRITERS` row applies it when
`tools.bot_desktop.runtime.is_supported_host()` is False (browser_navigate's
web-hint row is the precedent). Linux keeps the identical object, so
prompt-cache parity is untouched there.

Test: tests/tools/test_computer_use_schema_host.py — supported=False has no
handoff vocabulary and keeps every other action; supported=True is the frozen
schema (import error on bc36ddb5f969; the host is never faked).
2026-09-13 05:59:34 -07:00
teknium1
16289e22f7 fix(bot-screen): wire lane call sites: atomic install claim, redacted lease views everywhere, backend rebind on display change, footgun annotations, grace in schema 2026-09-12 18:59:56 -07:00
Teknium
b41c8ddfc4 feat: Bot Screen — per-bot Xfce desktop streamed into Hermes Desktop with human take-over
A bot running on a headless Linux gateway now gets its own desktop (TigerVNC
Xvnc + a minimal Xfce session, one per profile) that Hermes Desktop streams
live. The user can watch the bot work, take over to type a login / 2FA code /
CAPTCHA, and hand control back; the bot refuses every computer_use action
(screenshots included) while a human holds the screen, then resumes with the
session cookies the human just created.

Why this shape:
- The screen lives on the machine Hermes runs on, not in a vendor cloud browser,
  so it works for any app the bot drives and keeps the session on the user's host.
- Xfce components are launched individually (xfwm4, xfce4-panel, xfdesktop,
  xfsettingsd) under a private dbus session rather than xfce4-session/the
  metapackage: no screensaver, power manager or polkit agent to lock or prompt
  a headless desktop.
- Transport is raw RFB over a WebSocket beside /api/ws, authenticated with a
  one-shot ticket minted through the already-authenticated RPC channel; noVNC
  runs in the Electron renderer. The bridge parses the RFB client stream and
  drops keyboard/pointer/clipboard (incl. QEMU Extended KeyEvent, which noVNC
  switches to once Xvnc advertises it) from any viewer that does not hold the
  lease, so viewOnly is enforced server-side, not by the client.
- One lease per profile (agent | human viewer) is the single truth for the RFB
  bridge, the computer_use tool and the Desktop UI; taking control evicts other
  viewers' input with close code 4000 control-taken.
- Auto-start happens only at the computer_use tool boundary (headless host,
  packages present, bot_desktop.auto_start=true); env builders stay pure so
  status probes and tests never spawn X servers. tests/tools/conftest.py pins
  the binaries to "missing" for the same reason the browser-use fixture does.

Surfaces: Desktop (Bots → right-click → Open Screen; Take over / Hand back),
CLI (`hermes computer-use screen status|start|stop|install-deps`), tool
(`computer_use` actions request_handoff / wait_for_human), gateway RPCs
(display.status/start/stop/observe/lease.acquire/lease.release + display.lease
event), docs page user-guide/features/bot-screen.
2026-09-12 18:57:40 -07:00
Teknium
facb9eb4f1 fix: trim computer use tool schema guidance 2026-09-08 17:03:01 -07:00
Teknium
8f877d40f9 refactor(computer_use): single-blank layout between top-level defs (AST-identical) 2026-09-02 16:35:33 -07:00
Teknium
04403a5632 refactor(computer_use): compact schema.py literal (value byte-identical; verified via tool-schema dump) 2026-09-02 16:12:51 -07:00
Teknium
f780cb36d8 refactor(computer_use): drop the cua_browser_* route — browser work goes through browser_exec
Real-profile browsing routes all in-page browser work through the Browser Use
CLI (browser_exec), which obsoletes the cua-driver typed-browser surface baked
into computer_use. Remove it so computer_use is a pure DESKTOP-control tool
(screenshots / mouse / keyboard / window management) and every call's schema
drops ~24 browser-only params + 9 actions.

- schema.py: 9 cua_browser_* actions and the typed-browser param block removed;
  14 desktop actions + shared params kept; description drops the browser rung.
- tool.py: cua_browser entries out of _SAFE/_DESTRUCTIVE_ACTIONS; the whole
  cua_browser dispatch block deleted; {"type","cua_browser_type"} → "type"
  (desktop typing untouched); _config_preauthorized (browser-prepare-only, a
  no-op for every desktop action) and the browser-page escalation hint removed.
- browser_route.py deleted (no importers outside the package); cua_backend.py
  drops the import + typed_browser_* methods; backend.py drops the non-abstract
  defaults.
- tests: browser-route/contract suites removed; browser assertions trimmed.

Desktop control unchanged. 233 computer_use tests pass; the 1 remaining failure
(test_gateway_session_key_yolo_maps_to_unrestricted_mode) is a pre-existing
cross-test state leak — fails identically on origin/main, passes in isolation.
2026-08-26 19:25:33 -07:00
Teknium
3da5897c39 refactor(computer_use): diet schema + delete prompt block (~1.4K tok/call); remove max_elements, ladder moves to response verdicts 2026-08-26 04:50:13 -07:00
Teknium
65605d4a7a fix(computer_use): pitch background-FIRST (not background-only) in schema, prompt block, and skill 2026-08-26 04:50:13 -07:00
Teknium
ce9b9a6351 feat(computer_use): guide models from full-screen grabs to interactive lanes
Full-screen captures carry no element tree, so the CaptureResult now has a
'note' field surfaced in the tool summary telling the model to call
capture(app='<AppName>') or capture(app='desktop') when it needs to act on
what it sees. Schema description updated to distinguish app='screen'
(composited full-screen image) from app='desktop' (shell surface with
clickable elements); docs + regression tests (14, sabotage-verified) added.
2026-08-25 02:31:34 -07:00
Gille
188d47919a feat(computer-use): expose screenshots for chat delivery 2026-08-19 22:58:02 -07:00
Francesco Bonacci
5f049b517b fix(computer-use): make the typed-browser bind/snapshot split discoverable
`cua_browser_state` has two branches, chosen implicitly: any call carrying
pid or window_id is a *binding* (browser_route.py:252), anything else is a
*snapshot*. A binding clears session state, mints fresh tab_ids, returns
binding metadata with no page content, and sets verification_required.

Nothing in the response says that. A caller that keeps passing pid/window_id
- the natural reading of "bind to this window, then read it" - re-binds
forever: the tab_id it just received is unbound by the next bind, so every
cua_browser_navigate comes back browser_verification_required, and the
refusal ("take a fresh snapshot") points at the same call that just re-bound.
Observed live as 11 consecutive refused navigates before the model gave up
and fell back to foreground SendInput on the address bar.

The same confusion silently swallowed include_screenshot: both calls that
requested one were bindings, which carry no page content, so the flag had
nothing to attach to and was dropped without comment.

A binding response now reports snapshot_required, next_step
(fresh_browser_state, matching the existing token convention) and a hint
naming the exact next call; requesting a screenshot on a binding reports
screenshot_deferred instead of dropping it. The verification refusal now
says to call cua_browser_state WITHOUT pid/window_id and why re-sending them
does not help. The schema documents that include_screenshot applies to
snapshots.

Behavior of the bind and snapshot branches themselves is unchanged - this is
purely about making the split legible to the caller.

Unit-tested. Not verified end to end on the reporting host: the driver
refuses the bind upstream there (`browser_requires_setup: no owned DevTools
endpoint`, and it does not accept a user-launched --remote-debugging-port),
so the typed route never reaches this branch. That attach failure is a
separate cua-driver issue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 11:34:40 -07:00
Francesco Bonacci
a403fe6f92 feat(computer-use): support Cua Driver 0.20 runtime contracts 2026-08-16 11:34:40 -07:00
Teknium
20cf326bd1 fix(computer-use): align browser authorization with live-verified cua-driver 0.19.3 contract
Live-tested against the real cua-driver 0.19.3 binary (Linux x86_64):

- bounded serve flags corrected: the daemon accepts
  --session-policy/--approve-session-policy, not the docs'
  --capability-manifest names (which it rejects). Verified end-to-end:
  a bounded daemon with a real policy file starts and reports running.
- browser-approve verified real but interactive-only (refuses without a
  TTY) and its token is a legacy compatibility path disabled by default
  on current drivers (per the live browser_prepare schema). Kept as a
  passthrough; no longer presented as the primary route.
- NEW primary standard-mode route, verified live: launch the runtime
  with cua-driver's trusted-launcher grant. config opt-in
  computer_use.grant_existing_profile: true appends
  --grant existing-profile to the standard-mode MCP spawn (MCP
  initialize verified accepting the flag). Default false = attachment
  keeps failing closed. Never applied to bounded/unrestricted daemons.
- Skill, system prompt, tool schema, and docs updated to the verified
  ladder: config grant > bounded manifest > YOLO; token = legacy.
2026-08-15 15:04:32 -07:00
Teknium
48dd9c87cf feat(computer-use): user-facing authorization for cua-driver browser attachment
Completes the typed cua_browser_* route (PR #74166 lineage) with the
authorization surface that makes existing-profile attachment and
repeatable bounded automation reachable by real users:

- hermes computer-use browser-approve: CLI passthrough that mints
  cua-driver's five-minute single-use attachment token for one exact
  (pid, window_id). The user, never the model, is the token source.
- approval_token passthrough on cua_browser_prepare (schema + dispatch +
  browser_route), forwarded only for existing_profile and only as a
  non-empty string.
- computer_use.permission_mode: bounded + capability_manifest config:
  private per-session embedded daemon launched with
  --capability-manifest/--approve-capability-manifest; missing manifest
  fails loudly. 'unrestricted' is deliberately NOT a config value —
  it stays bound to the explicit per-session YOLO toggle.
- Skill + system-prompt + docs guidance for the three authorization
  rungs and the isolated-profile-first default.

E2E-verified against a temp HERMES_HOME: real config resolution to
bounded, loud failure without a manifest, real argparse path driving a
fake cua-driver binary, standard default preserved.
2026-08-15 15:04:32 -07:00
Francesco Bonacci
c268397752 feat(computer_use): align cua-driver 0.10 permission modes 2026-07-29 12:19:37 -07:00
Francesco Bonacci
847e401b74 feat(computer_use): align cua-driver 0.9 contracts
Salvaged from PR #67807 by @f-trycua onto current main.

- Foreground gate: discover delivery_mode support from the live tools/list
  inputSchema.properties (fail closed), not the never-shipped
  input.delivery_mode capability token
- bring_to_front: standalone strict-schema MCP tool (inject_session=False),
  separate approval scope, requires foreground
- Verdict precedence: confirmed > unverifiable (verify before retry) >
  suspected_noop/refusal (escalate); surfaced as explicit verdict field
- Typed cua_browser_* route inside computer_use (browser_route.py) with
  exact-binding, adapter-injected session, snapshot-scoped refs
- Per-Hermes-session backend isolation + release_computer_use_session seam
  wired into AIAgent.close()
- Recorded 0.9 tools/list fixture replaces fabricated capability tokens
2026-07-29 12:19:37 -07:00
Teknium
9d6d772837 feat(computer_use): follow cua-driver's verify → escalate ladder (#67123)
Hermes' computer_use wrapper dropped cua-driver's structured action verdicts,
exposed no delivery_mode, and injected background-only guidance — so the agent
reported unverified no-ops as success and concluded cua-driver 'cannot drive'
Electron/Chromium surfaces (observed live on tldraw offline). Fixes #67052.

Phase A — preserve the result contract:
- ActionResult carries verified/effect/escalation/path/degraded/code/delivery_mode
- CuaDriverBackend._action() reads structuredContent (was data-only); a helper
  normalizes it, additive and None-safe on old drivers
- _text_response surfaces the fields additively (ok stays transport-only)

Phase B — bounded, model-reachable foreground:
- delivery_mode (background|foreground) + bring_to_front on the schema, dispatcher,
  ABC, and all input methods
- foreground is capability-gated (input.delivery_mode); old drivers get a
  structured foreground_unsupported refusal, never a silent background downgrade
- no automatic/hidden foreground retry — the model selects it from the signal

Phase C — guidance + isolation:
- system prompt (prompt_builder) and bundled skills/computer-use/SKILL.md go from
  background-ONLY to background-FIRST, teaching the AX→PX→foreground ladder driven
  by returned effect/escalation, not predicted from the app being Electron
- foreground approval scoped by (action, delivery_mode): a background approval
  never silently authorizes foreground
- approval state keyed per session_id so concurrent gateway runs don't leak unlocks

Tests: tests/tools/test_computer_use_delivery_ladder.py (15) cover confirmed/
unverifiable/suspected_noop/degraded/old-driver verdicts, delivery_mode gating +
foreground_unsupported, and session-scoped foreground approval. Existing 265
computer_use tests still green.

Live E2E (real cua-driver 0.8.3 + tldraw offline on Linux/X11): a background click
returned effect='unverifiable'/path='ax' (no fabricated success), and a foreground
request returned code='foreground_unsupported' — correct on a driver that predates
the input.delivery_mode capability.
2026-07-18 13:59:35 -07:00
kshitij
2b6897f982 fix(computer-use): target Linux app windows reliably (#63725)
* fix(computer-use): target Linux app windows reliably

Resolve app filters through the canonical cua-driver MCP app metadata and join running app PIDs back to windows. Preserve an exact selected window across capture_after, support direct capture by pid/window_id, and send the active window ID for coordinate pointer actions on Linux.

Co-authored-by: annguyenNous <annguyenNous@users.noreply.github.com>
Co-authored-by: grimmjoww578 <willies578@gmail.com>
Co-authored-by: ai-ag2026 <261867348+ai-ag2026@users.noreply.github.com>

* fix(computer-use): address review on PR #63725

Address three review comments from @f-trycua:

1. type_text, press_key, and hotkey now carry _active_window_id and
   fail closed when it is missing. Previously they sent only the PID,
   so CUA Driver fell back to the first window for that PID — input
   could reach the wrong window in multi-window apps.

2. Coordinate scroll x/y are now capability-gated behind
   input.scroll.coordinates. CUA Driver 0.7.1 Linux schema rejects
   x/y on scroll; omitting them when the driver doesn't advertise
   support avoids the schema rejection while still routing via
   window_id.

3. Windows are sorted by z_index descending (higher = front, per CUA
   Driver semantics) instead of ascending. Null z_index (Wayland) is
   coerced to 0 in _ingest_windows so it doesn't crash the sort and
   sorts to the back instead of being selected as the capture target.

---------

Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>
Co-authored-by: annguyenNous <annguyenNous@users.noreply.github.com>
Co-authored-by: grimmjoww578 <willies578@gmail.com>
Co-authored-by: ai-ag2026 <261867348+ai-ag2026@users.noreply.github.com>
2026-07-16 00:34:20 +05:30
Teknium
30e5d0092d feat(computer-use): add whole-screen/desktop capture target
capture(app='screen'|'desktop') now resolves to the OS shell/desktop
window (Windows Progman/WorkerW desktop or Shell_TrayWnd taskbar, macOS
Finder/Dock) so 'show me my screen' and 'click the taskbar' work.
Previously capture() only matched application windows, and the schema
advertised 'or the whole screen' without any code path delivering it.

cua-driver is window-oriented (no virtual-desktop or per-monitor MCP
tool), so a single image still cannot span multiple monitors — the
schema now states this and the no-desktop-window path returns a clear
message instead of silently grabbing the frontmost app.
2026-06-22 12:21:58 -07:00
Francesco Bonacci
f2e37549c6 feat(computer_use): cross-platform cua-driver (macOS/Windows/Linux)
Make the computer_use toolset platform-agnostic by driving cua-driver on
macOS, Windows, and Linux. Consumes the 8 cua-driver decoupling surfaces
(capability discovery, structuredContent AX tree, opaque element_token,
click button enum, explicit mimeType, machine-readable manifest,
structured list_windows, structured health_report), each degrading
gracefully on older drivers.

Adds `hermes computer-use doctor` (drives cua-driver health_report with a
per-OS check matrix and an exit 0/1/2 ok/degraded/blocked contract), full
typed wrappers for the previously-uncovered cua-driver tools plus a generic
call_tool escape hatch, per-session agent-cursor lifecycle, platform-aware
system-prompt guidance (host-deterministic, cache-safe), and honors
HERMES_CUA_DRIVER_CMD end-to-end.

Replaces the macOS-only skills/apple/macos-computer-use skill with a
cross-platform skills/computer-use skill, and refreshes the EN + zh-Hans
docs.

Supersedes #44221 (Windows-enablement salvage of #30660).

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-06-22 06:42:30 -07:00
briandevans
280dd4513a fix(computer-use): address Copilot review on max_elements cap
Four findings from Copilot's review on PR #22891, all in the AX
elements-array cap added by 22fa1ed:

1. The truncation note ("response truncated to N of M elements") was
   appended unconditionally — including in the som/vision multimodal
   path, whose response carries a screenshot rather than an `elements`
   array. The note described a payload field that wasn't present.
   Moved the note into the AX-text branch where the array actually
   appears.

2. `_format_elements(cap.elements)` ran on the full untrimmed list with
   its own `max_lines=40` cap, so a caller passing `max_elements=10`
   would see summary lines referencing `#11..#40` even though the JSON
   `elements` array only held #1..#10. Format on `visible_elements`
   instead so the summary indices always exist in the response.

3. `_coerce_max_elements` enforced a lower bound but no upper bound,
   so `max_elements=10_000_000` silently disabled the safeguard and
   reintroduced the original context-blow-up. Added a hard cap
   (`_MAX_ALLOWED_MAX_ELEMENTS = 1000`) that clamps oversized values.

4. The schema string said "Default 100" but the property carried no
   `default` field, and claimed `max_elements` had no effect on som/
   vision while the image-missing fallback path can still return an
   elements array. Added `"default": 100`, `"maximum": 1000`, and
   clarified the fallback-path wording.

Each finding gets a regression test:

- test_capture_ax_clamps_oversized_max_elements_to_hard_cap
- test_capture_ax_summary_indices_match_returned_elements
- test_capture_multimodal_summary_omits_truncation_note
- test_schema_max_elements_documents_default_and_upper_bound

Verified with `pytest tests/tools/test_computer_use.py` (53 passed,
including the 5 new cases). Confirmed each new test fails on the
pre-fix code path before applying the production change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 19:07:32 -07:00
briandevans
bb694bad42 fix(computer-use): cap AX elements array to prevent context blowup (#22865)
`computer_use(action='capture', mode='ax')` returned the full AX element
list verbatim in the JSON response. Dense Electron / Obsidian / JetBrains
UIs publish 500+ AX nodes (one reproduction in #22865 returned 597
elements against Obsidian), so a single capture could consume enough
context to trigger compression failures or render the session unusable.
The human-readable `_format_elements` summary is already capped at 40
lines, so the truncation gap was invisible to anyone reading the summary
output.

Add a `max_elements` argument to the tool schema, default 100, that
trims the AX `elements` array. When the cap fires, the response surfaces
`total_elements` and `truncated_elements` and appends a "raise
max_elements or pass app= to narrow" hint to the summary so the model
knows the JSON view is partial and can re-issue with a tighter scope.

Validation is centralized in `_coerce_max_elements`: missing /
non-integer / sub-1 inputs fall back to the default cap, so the
protection can never be silently disabled by a malformed tool-call
argument. The cap only affects AX-mode JSON; `mode='som'` and
`mode='vision'` keep returning a screenshot + image-aware summary
unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 19:07:32 -07:00
ddupont
e31f3b3c56 feat(computer-use): background focus-safe backend — set_value, structured windows, MIME detection
Extends the cua-driver computer-use backend to drive backgrounded macOS
windows without stealing keyboard or mouse focus from the foreground app.
All changes target the cua-driver MCP backend and the shared dispatcher.

## cua_backend.py

**Window-aware capture**: capture() now calls list_windows + get_window_state
instead of the removed capture tool. Prefers structuredContent.windows
(MCP 2024-11-05+ cua-driver) for zero-parse window enumeration; falls back
to regex-parsed text for older builds. Stores the selected (pid, window_id)
as sticky context so subsequent action calls do not need a redundant round-trip.

**Action routing**: click/scroll/type_text/key all carry the sticky pid
(and window_id for element-indexed clicks). type_text routes through
type_text_chars (individual key events) rather than AX attribute write --
WebKit AXTextFields reject attribute writes from backgrounded processes.

**Key parsing**: _parse_key_combo splits cmd+s-style strings into
(key, [modifiers]) and routes to hotkey (modifier present) or
press_key (bare key) -- cua-driver actual tool names.

**set_value method**: new set_value(value, element) calls the cua-driver
set_value MCP tool. For AXPopUpButton / HTML select in a backgrounded Safari,
AXPress opens the native macOS popup which closes immediately when the app is
non-frontmost; set_value AX-presses the matching child option directly
(no menu required, no focus steal).

**focus_app**: reimplemented as a pure window-selector (enumerates
list_windows, sets sticky pid/window_id) without ever raising the window
or stealing focus.

**list_apps**: fixed tool name from listApps to list_apps; handles plain-text
response via regex when structured data is absent.

**Structured-content extraction**: _extract_tool_result now surfaces
structuredContent from MCP results, enabling the list_windows window array
without text parsing.

**Helpers**: _parse_windows_from_text, _parse_elements_from_tree,
_split_tree_text, _parse_key_combo extracted as module-level functions.

## schema.py

Added set_value to the action enum with a description explaining when to
prefer it over click (select/popup elements, sliders, no focus steal).
Added value field for set_value payloads.

## tool.py

Routed set_value action through _dispatch to backend.set_value.
Added set_value to _DESTRUCTIVE_ACTIONS (approval-gated).
Fixed MIME-type detection in _capture_response: cua-driver may return
JPEG; detect from base64 magic bytes (/9j/ -> image/jpeg, else image/png)
rather than hardcoding image/png.

## agent/display.py + run_agent.py

Guard _detect_tool_failure and result-preview logic against non-string
function_result values: multimodal tool results (dicts with _multimodal=True)
are not string-sliceable; treat them as successes and fall back to str()
for length/preview.
2026-05-08 11:07:38 -07:00
Teknium
850413f120 feat(computer-use): cua-driver backend, universal any-model schema
Background macOS desktop control via cua-driver MCP — does NOT steal the
user's cursor or keyboard focus, works with any tool-capable model.

Replaces the Anthropic-native `computer_20251124` approach from the
abandoned #4562 with a generic OpenAI function-calling schema plus SOM
(set-of-mark) captures so Claude, GPT, Gemini, and open models can all
drive the desktop via numbered element indices.

- `tools/computer_use/` package — swappable ComputerUseBackend ABC +
  CuaDriverBackend (stdio MCP client to trycua/cua's cua-driver binary).
- Universal `computer_use` tool with one schema for all providers.
  Actions: capture (som/vision/ax), click, double_click, right_click,
  middle_click, drag, scroll, type, key, wait, list_apps, focus_app.
- Multimodal tool-result envelope (`_multimodal=True`, OpenAI-style
  `content: [text, image_url]` parts) that flows through
  handle_function_call into the tool message. Anthropic adapter converts
  into native `tool_result` image blocks; OpenAI-compatible providers
  get the parts list directly.
- Image eviction in convert_messages_to_anthropic: only the 3 most
  recent screenshots carry real image data; older ones become text
  placeholders to cap per-turn token cost.
- Context compressor image pruning: old multimodal tool results have
  their image parts stripped instead of being skipped.
- Image-aware token estimation: each image counts as a flat 1500 tokens
  instead of its base64 char length (~1MB would have registered as
  ~250K tokens before).
- COMPUTER_USE_GUIDANCE system-prompt block — injected when the toolset
  is active.
- Session DB persistence strips base64 from multimodal tool messages.
- Trajectory saver normalises multimodal messages to text-only.
- `hermes tools` post-setup installs cua-driver via the upstream script
  and prints permission-grant instructions.
- CLI approval callback wired so destructive computer_use actions go
  through the same prompt_toolkit approval dialog as terminal commands.
- Hard safety guards at the tool level: blocked type patterns
  (curl|bash, sudo rm -rf, fork bomb), blocked key combos (empty trash,
  force delete, lock screen, log out).
- Skill `apple/macos-computer-use/SKILL.md` — universal (model-agnostic)
  workflow guide.
- Docs: `user-guide/features/computer-use.md` plus reference catalog
  entries.

44 new tests in tests/tools/test_computer_use.py covering schema
shape (universal, not Anthropic-native), dispatch routing, safety
guards, multimodal envelope, Anthropic adapter conversion, screenshot
eviction, context compressor pruning, image-aware token estimation,
run_agent helpers, and universality guarantees.

469/469 pass across tests/tools/test_computer_use.py + the affected
agent/ test suites.

- `model_tools.py` provider-gating: the tool is available to every
  provider. Providers without multi-part tool message support will see
  text-only tool results (graceful degradation via `text_summary`).
- Anthropic server-side `clear_tool_uses_20250919` — deferred;
  client-side eviction + compressor pruning cover the same cost ceiling
  without a beta header.

- macOS only. cua-driver uses private SkyLight SPIs
  (SLEventPostToPid, SLPSPostEventRecordTo,
  _AXObserverAddNotificationAndCheckRemote) that can break on any macOS
  update. Pin with HERMES_CUA_DRIVER_VERSION.
- Requires Accessibility + Screen Recording permissions — the post-setup
  prints the Settings path.

Supersedes PR #4562 (pyautogui/Quartz foreground backend, Anthropic-
native schema). Credit @0xbyt4 for the original #3816 groundwork whose
context/eviction/token design is preserved here in generic form.
2026-05-08 11:07:38 -07:00
Teknium
e63364b8df revert: computer-use cua-driver (PR #16919) (#16927)
Reverts PR #16919 (commits dad10a78d, 413ee1a28, b4a8031b2, afb958829)
which was merged prematurely. Restoring the pre-merge state so #14817
and #15328 can be revisited as standing PRs.

Reverted commits:
- afb958829 fix(computer-use): harden image-rejection fallback + AUTHOR_MAP
- b4a8031b2 fix(computer-use): unwrap _multimodal tool results
- 413ee1a28 feat(computer-use): background focus-safe backend
- dad10a78d feat(computer-use): cua-driver backend, universal any-model schema

Co-authored-by: teknium1 <teknium@users.noreply.github.com>
2026-04-28 01:57:21 -07:00
ddupont
413ee1a286 feat(computer-use): background focus-safe backend — set_value, structured windows, MIME detection
Extends the cua-driver computer-use backend to drive backgrounded macOS
windows without stealing keyboard or mouse focus from the foreground app.
All changes target the cua-driver MCP backend and the shared dispatcher.

## cua_backend.py

**Window-aware capture**: capture() now calls list_windows + get_window_state
instead of the removed capture tool. Prefers structuredContent.windows
(MCP 2024-11-05+ cua-driver) for zero-parse window enumeration; falls back
to regex-parsed text for older builds. Stores the selected (pid, window_id)
as sticky context so subsequent action calls do not need a redundant round-trip.

**Action routing**: click/scroll/type_text/key all carry the sticky pid
(and window_id for element-indexed clicks). type_text routes through
type_text_chars (individual key events) rather than AX attribute write --
WebKit AXTextFields reject attribute writes from backgrounded processes.

**Key parsing**: _parse_key_combo splits cmd+s-style strings into
(key, [modifiers]) and routes to hotkey (modifier present) or
press_key (bare key) -- cua-driver actual tool names.

**set_value method**: new set_value(value, element) calls the cua-driver
set_value MCP tool. For AXPopUpButton / HTML select in a backgrounded Safari,
AXPress opens the native macOS popup which closes immediately when the app is
non-frontmost; set_value AX-presses the matching child option directly
(no menu required, no focus steal).

**focus_app**: reimplemented as a pure window-selector (enumerates
list_windows, sets sticky pid/window_id) without ever raising the window
or stealing focus.

**list_apps**: fixed tool name from listApps to list_apps; handles plain-text
response via regex when structured data is absent.

**Structured-content extraction**: _extract_tool_result now surfaces
structuredContent from MCP results, enabling the list_windows window array
without text parsing.

**Helpers**: _parse_windows_from_text, _parse_elements_from_tree,
_split_tree_text, _parse_key_combo extracted as module-level functions.

## schema.py

Added set_value to the action enum with a description explaining when to
prefer it over click (select/popup elements, sliders, no focus steal).
Added value field for set_value payloads.

## tool.py

Routed set_value action through _dispatch to backend.set_value.
Added set_value to _DESTRUCTIVE_ACTIONS (approval-gated).
Fixed MIME-type detection in _capture_response: cua-driver may return
JPEG; detect from base64 magic bytes (/9j/ -> image/jpeg, else image/png)
rather than hardcoding image/png.

## agent/display.py + run_agent.py

Guard _detect_tool_failure and result-preview logic against non-string
function_result values: multimodal tool results (dicts with _multimodal=True)
are not string-sliceable; treat them as successes and fall back to str()
for length/preview.
2026-04-28 01:46:36 -07:00
Teknium
dad10a78d0 feat(computer-use): cua-driver backend, universal any-model schema
Background macOS desktop control via cua-driver MCP — does NOT steal the
user's cursor or keyboard focus, works with any tool-capable model.

Replaces the Anthropic-native `computer_20251124` approach from the
abandoned #4562 with a generic OpenAI function-calling schema plus SOM
(set-of-mark) captures so Claude, GPT, Gemini, and open models can all
drive the desktop via numbered element indices.

- `tools/computer_use/` package — swappable ComputerUseBackend ABC +
  CuaDriverBackend (stdio MCP client to trycua/cua's cua-driver binary).
- Universal `computer_use` tool with one schema for all providers.
  Actions: capture (som/vision/ax), click, double_click, right_click,
  middle_click, drag, scroll, type, key, wait, list_apps, focus_app.
- Multimodal tool-result envelope (`_multimodal=True`, OpenAI-style
  `content: [text, image_url]` parts) that flows through
  handle_function_call into the tool message. Anthropic adapter converts
  into native `tool_result` image blocks; OpenAI-compatible providers
  get the parts list directly.
- Image eviction in convert_messages_to_anthropic: only the 3 most
  recent screenshots carry real image data; older ones become text
  placeholders to cap per-turn token cost.
- Context compressor image pruning: old multimodal tool results have
  their image parts stripped instead of being skipped.
- Image-aware token estimation: each image counts as a flat 1500 tokens
  instead of its base64 char length (~1MB would have registered as
  ~250K tokens before).
- COMPUTER_USE_GUIDANCE system-prompt block — injected when the toolset
  is active.
- Session DB persistence strips base64 from multimodal tool messages.
- Trajectory saver normalises multimodal messages to text-only.
- `hermes tools` post-setup installs cua-driver via the upstream script
  and prints permission-grant instructions.
- CLI approval callback wired so destructive computer_use actions go
  through the same prompt_toolkit approval dialog as terminal commands.
- Hard safety guards at the tool level: blocked type patterns
  (curl|bash, sudo rm -rf, fork bomb), blocked key combos (empty trash,
  force delete, lock screen, log out).
- Skill `apple/macos-computer-use/SKILL.md` — universal (model-agnostic)
  workflow guide.
- Docs: `user-guide/features/computer-use.md` plus reference catalog
  entries.

44 new tests in tests/tools/test_computer_use.py covering schema
shape (universal, not Anthropic-native), dispatch routing, safety
guards, multimodal envelope, Anthropic adapter conversion, screenshot
eviction, context compressor pruning, image-aware token estimation,
run_agent helpers, and universality guarantees.

469/469 pass across tests/tools/test_computer_use.py + the affected
agent/ test suites.

- `model_tools.py` provider-gating: the tool is available to every
  provider. Providers without multi-part tool message support will see
  text-only tool results (graceful degradation via `text_summary`).
- Anthropic server-side `clear_tool_uses_20250919` — deferred;
  client-side eviction + compressor pruning cover the same cost ceiling
  without a beta header.

- macOS only. cua-driver uses private SkyLight SPIs
  (SLEventPostToPid, SLPSPostEventRecordTo,
  _AXObserverAddNotificationAndCheckRemote) that can break on any macOS
  update. Pin with HERMES_CUA_DRIVER_VERSION.
- Requires Accessibility + Screen Recording permissions — the post-setup
  prints the Settings path.

Supersedes PR #4562 (pyautogui/Quartz foreground backend, Anthropic-
native schema). Credit @0xbyt4 for the original #3816 groundwork whose
context/eviction/token design is preserved here in generic form.
2026-04-28 01:46:36 -07:00