Files
hermes-agent/tools/bot_desktop/runtime.py
Teknium 399d956903 feat(bot_desktop): Bot Screen, computer_use and the browser run inside the terminal backend (#121169)
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends

The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.

docker/sandbox-desktop.Dockerfile is that base plus:
  - the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
    zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
  - the exact package set the Hermes -desktop image installs (TigerVNC,
    Xfce components, dbus, xauth, fonts)
  - Playwright's headed Chromium (same build as the -desktop image)
  - cua-driver 0.28.2 from its pinned release tarball

No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.

docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.

* feat(docker): sandbox desktop base on python3.13-nodejs26

Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.

* feat(docker): bake agent-browser into hermes-sandbox:desktop

The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.

* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend

A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.

Now the desktop lives where the terminal lives:

- tools/environments/streams.py: one primitive per spawn-per-call backend, the
  local argv prefix that runs its remainder inside the sandbox with stdio open
  (`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
  vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
  the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
  prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
  gateway). `auto` follows the backend; a sandbox that cannot host a screen
  REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
  lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
  there; the host driver's runtime contract is irrelevant then; check_fn is
  true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
  with the daemon, socket dir and profile inside the sandbox; screenshots are
  fetched back so MEDIA: paths keep working; recycle closes the sandbox
  daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
  in a sandbox, a unix socket otherwise.

Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.

* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only

DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).

* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read

Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.

* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order

Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host

sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.

* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox

Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).

* chore: retrigger CI (zero-job dispatch failure)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore(config): template stamps v47 and shows the new default sandbox image

The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.

* feat(sandbox): the default image change is a decision, not a surprise

A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.

Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.

Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.

Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.

* fix(config): both plain defaults that preceded the desktop sandbox image are template copies

main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.

* ci(docker): build the sandbox image on release/dispatch, not every main push

Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.

* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta

The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.

* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)

* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)

* fix(bot_desktop): placement is the authority; sandbox screen survives restarts

Review findings on the sandbox-hosted Bot Screen, each reproduced live first.

Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.

Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.

SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.

CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.

pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.

Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.

Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.

Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.

* docs(bot-screen): no literal tmp path in the profile-location note

* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)

* fix(bot_desktop): adopting a screen the sandbox kept records the marker

Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.

Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).

* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)

* docs(bot-screen): what the Apptainer path inherits from the image and what it does not

* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 03:34:07 -07:00

838 lines
39 KiB
Python

"""Bot Desktop runtime: one headless Xfce desktop per Hermes profile, served over RFB on a private
Unix socket, viewed and driven from Hermes Desktop.
Layout under ``<HERMES_HOME>/bot-desktop/``: ``display`` (allocated X display number), ``rfb.sock``
(Xvnc RFB Unix socket, 0600), ``Xauthority``, ``env`` (DISPLAY/XAUTHORITY/DBUS_SESSION_BUS_ADDRESS
published by the launcher once Xfce's bus exists), ``launcher.pid``, ``launcher.log``, ``xdg/``
(per-profile XDG_CONFIG_HOME so two profiles never share xfconf). Everything is profile-scoped via
``get_hermes_home()`` so N profiles in one gateway get N desktops: one screen per bot on the shared
machine.
The launcher is ``launcher.sh`` next to this module; :func:`desktop_env` is what cua-driver and headed
Chromium spawns merge in so the agent acts on this profile's screen and nowhere else.
"""
from __future__ import annotations
import contextlib
import logging
import os
import shutil
import signal
import subprocess
import sys
import time
from dataclasses import dataclass
from pathlib import Path
from typing import Dict, Optional
from hermes_constants import get_hermes_home
from tools.bot_desktop import placement
logger = logging.getLogger(__name__)
_LAUNCHER = Path(__file__).with_name("launcher.sh")
# Display numbers below 10 collide with real seats and default Xvfb recipes (:99 is popular too); scan a
# private band and record the choice so restarts reuse it.
_DISPLAY_MIN, _DISPLAY_MAX = 20, 89
# Binaries the launcher execs; the package hint is per distro family.
REQUIRED_BINARIES = ("Xvnc", "xfwm4", "xfce4-panel", "xfdesktop", "xfsettingsd", "dbus-run-session",
"xauth", "xdpyinfo", "setxkbmap", "xprop")
# Which package in each distro list ships each required binary. Fedora retired the xorg-x11-utils /
# xorg-x11-server-utils umbrellas (per-binary packages since F35) and dnf5 refuses the whole transaction on
# one unknown name, so every binary must map to a package that still resolves; the test suite checks that
# each mapped package is in PACKAGES for its manager.
BINARY_PACKAGES = {
"apt": {"Xvnc": "tigervnc-standalone-server", "xfwm4": "xfwm4", "xfce4-panel": "xfce4-panel",
"xfdesktop": "xfdesktop4", "xfsettingsd": "xfce4-settings", "dbus-run-session": "dbus-x11",
"xauth": "xauth", "xdpyinfo": "x11-utils", "setxkbmap": "x11-xkb-utils", "xprop": "x11-utils"},
# tigervnc-x11-server is the real package (tigervnc-server-minimal is only a Provides on it); dbus-run-session
# is in dbus-daemon (dbus-x11 ships dbus-launch only).
"dnf": {"Xvnc": "tigervnc-x11-server", "xfwm4": "xfwm4", "xfce4-panel": "xfce4-panel",
"xfdesktop": "xfdesktop", "xfsettingsd": "xfce4-settings", "dbus-run-session": "dbus-daemon",
"xauth": "xorg-x11-xauth", "xdpyinfo": "xdpyinfo", "setxkbmap": "setxkbmap", "xprop": "xprop"},
"pacman": {"Xvnc": "tigervnc", "xfwm4": "xfwm4", "xfce4-panel": "xfce4-panel", "xfdesktop": "xfdesktop",
"xfsettingsd": "xfce4-settings", "dbus-run-session": "dbus",
"xauth": "xorg-xauth", "xdpyinfo": "xorg-xdpyinfo", "setxkbmap": "xorg-setxkbmap", "xprop": "xorg-xprop"},
}
PACKAGES = {
"apt": ["tigervnc-standalone-server", "xfce4-panel", "xfwm4", "xfdesktop4", "xfce4-settings",
"xfce4-terminal", "dbus-x11", "x11-xserver-utils", "x11-utils", "x11-xkb-utils", "xauth",
"fonts-dejavu-core"],
"dnf": ["tigervnc-x11-server", "xfce4-panel", "xfwm4", "xfdesktop", "xfce4-settings",
"xfce4-terminal", "dbus-daemon", "xsetroot", "xset", "xdpyinfo", "xprop", "xorg-x11-xauth", "setxkbmap",
"dejavu-sans-fonts"],
"pacman": ["tigervnc", "xfce4-panel", "xfwm4", "xfdesktop", "xfce4-settings", "xfce4-terminal", "dbus",
"xorg-xsetroot", "xorg-xset", "xorg-xdpyinfo", "xorg-xprop", "xorg-xauth", "xorg-setxkbmap",
"ttf-dejavu"],
}
def state_dir() -> Path:
return get_hermes_home() / "bot-desktop"
def is_supported_host() -> bool:
return sys.platform.startswith("linux")
def missing_binaries() -> list[str]:
return [b for b in REQUIRED_BINARIES if shutil.which(b) is None]
def package_manager() -> Optional[str]:
for pm in ("apt-get", "dnf", "pacman"):
if shutil.which(pm):
return "apt" if pm == "apt-get" else pm
return None
def install_command() -> Optional[str]:
"""The distro command that installs the Bot Desktop packages, as the human would type it on THIS host:
prefixed with ``sudo`` unless Hermes already runs as root, so it is both what the pane shows and what
:mod:`tools.bot_desktop.install` runs. ``None`` when no package manager is present.
Not a promise that it can run here: see :func:`installable`. The published Docker image supervises
every service under ``s6-setuidgid hermes`` (UID 10000 by default) and ships no ``sudo`` binary, so an
install on a hosted instance is impossible no matter what this returns."""
pm = package_manager()
if pm is None:
return None
pkgs = " ".join(PACKAGES[pm])
body = {
"apt": f"apt-get install -y --no-install-recommends {pkgs}",
"dnf": f"dnf install -y {pkgs}",
"pacman": f"pacman -S --needed --noconfirm {pkgs}",
}[pm]
return body if is_root() else f"sudo {body}"
def is_root() -> bool:
return hasattr(os, "geteuid") and os.geteuid() == 0
def installable() -> bool:
"""Whether :func:`install_command` could actually succeed on this host.
False on an unprivileged process with no ``sudo`` to reach for, which is exactly the published Docker
image: services drop to the ``hermes`` user and no ``sudo`` binary is installed. The packages can only
arrive in the image there, so :func:`start` says that instead of printing a sudo line the user has no
way to run. ``status()`` still reports ``install_command`` for the pane; surfacing this there needs a
wire-contract change and is deliberately out of scope.
"""
if package_manager() is None:
return False
return is_root() or shutil.which("sudo") is not None
@dataclass
class DesktopStatus:
profile: str
supported: bool
installed: bool
missing: list[str]
running: bool
pid: Optional[int]
display: Optional[str]
socket: Optional[str]
geometry: str
install_command: Optional[str]
browser: Optional[str] # headed Chromium the dock's Browser icon and agent-browser share; None = no headed browser
blocker: Optional[str] = None # why start() would refuse right now (memory); None = may start
memory_available_mb: Optional[int] = None
memory_limit_mb: Optional[int] = None
placement: str = "gateway" # "gateway" | "terminal:<backend>" — where Xvnc runs
image_switch: Optional[Dict[str, object]] = None # pending default-image switch the pane can approve
def as_dict(self) -> Dict[str, object]:
return dict(self.__dict__)
def _read(path: Path) -> Optional[str]:
try:
return path.read_text(encoding="utf-8-sig").strip() or None
except OSError:
return None
def _pid_alive(pid: int) -> bool:
"""A zombie is dead for our purposes: a SIGKILLed launcher stays a zombie in the gateway until the next
Popen reaps it, and reporting it as running would hide its orphaned X server behind a live status."""
import psutil
try:
return psutil.Process(pid).status() != psutil.STATUS_ZOMBIE
except psutil.Error:
return False
def _create_time(pid: int) -> Optional[float]:
import psutil
try:
return psutil.Process(pid).create_time()
except (psutil.Error, OverflowError, ValueError):
return None
def _launcher_pid() -> Optional[int]:
"""The live launcher's pid, or None. ``launcher.pid`` holds ``"<pid> <create_time>"``: a recycled pid
with a different start time is somebody else's process and must never be reported as ours nor
killed by :func:`stop`. The pre-identity single-number format is treated as not running."""
raw = _read(state_dir() / "launcher.pid")
pid_s, _, born_s = (raw or "").partition(" ")
if not pid_s.isdigit() or not born_s:
return None
try:
pid, born = int(pid_s), float(born_s)
except ValueError:
return None
actual = _create_time(pid)
return pid if actual is not None and abs(actual - born) < 0.01 and _pid_alive(pid) else None
def _recorded_launcher_pid() -> Optional[int]:
"""The pid ``launcher.pid`` names, alive or not (the orphan sweep matches process groups against it)."""
pid_s, _, born_s = (_read(state_dir() / "launcher.pid") or "").partition(" ")
return int(pid_s) if pid_s.isdigit() and born_s else None
_X_LOCK_DIR = Path("/tmp") # no-tmp: ok — X servers write .X<n>-lock here by protocol (tests point it at a scratch dir)
_X_UNIX_TABLE = Path("/proc/net/unix") # the kernel's list of bound Unix sockets (tests point it at a fixture)
def _x_lock_pid(num: int) -> Optional[int]:
try:
return int((_X_LOCK_DIR / f".X{num}-lock").read_text(encoding="utf-8-sig").strip())
except (OSError, ValueError):
return None
def _x_socket_bound(num: int) -> bool:
"""A running X server keeps ``@/tmp/.X11-unix/X<n>`` (abstract namespace) bound for its whole life; the
kernel drops it only when the process exits. A /tmp reaper can remove the lock file under a live Xvnc,
and only this binding then still says the number is taken (a new server on it dies 'already running')."""
try:
lines = _X_UNIX_TABLE.read_text(encoding="utf-8-sig").splitlines()
except OSError:
return False
# no-tmp: ok — detects the X server's display socket at the path the X11 protocol fixes
return any(line.split()[-1].lstrip("@") == f"/tmp/.X11-unix/X{num}" for line in lines if line.strip())
def _display_in_use(num: int) -> bool:
"""A live X server owns ``:num``: its lock file names a running pid, or its X11 socket is bound. A lock
left by a crashed server (dead pid, no socket) does not count, so the number can be reclaimed."""
pid = _x_lock_pid(num)
return (pid is not None and _pid_alive(pid)) or _x_socket_bound(num)
def _reap_orphaned_server(sd: Path) -> bool:
"""Caller holds ``start.lock`` and has established that no live launcher exists. The launcher runs Xvnc in
its own session, so a SIGKILLed launcher leaves the X server alive, holding the display and ``rfb.sock``;
``status()`` keys on the launcher and says stopped, and a naive restart allocates a second server next
to it and overwrites the socket path both now claim. The X lock of the recorded display names that
server: it is ours when it sits in the dead launcher's process group or its command line binds OUR
socket. When a /tmp reaper took the lock too (or the failed launch dropped ``display``), the socket path
on the Xvnc command line is the remaining handle — without it one X server leaks per occurrence. Kill it
(group first), drop the state it left, and report whether anything was signalled."""
import psutil
def _binds_our_socket(cmdline: list) -> bool:
return "Xvnc" in Path(cmdline[0] if cmdline else "").name and str(sd / "rfb.sock") in cmdline
recorded = _read(sd / "display")
pid = _x_lock_pid(int(recorded)) if recorded and recorded.isdigit() else None
if pid is None or not _pid_alive(pid):
pid = next((p.pid for p in psutil.process_iter(["cmdline"]) if _binds_our_socket(p.info["cmdline"] or [])), None)
if pid is None:
return False
launcher = _recorded_launcher_pid()
try:
pgid = os.getpgid(pid) # windows-footgun: ok — Linux-only runtime (is_supported_host gates start/stop)
cmdline = psutil.Process(pid).cmdline()
except (ProcessLookupError, psutil.Error):
return False
if pgid != launcher and not _binds_our_socket(cmdline):
return False # somebody else's server took the number after we died; never touch it
logger.warning("Bot Desktop launcher %s is gone but its X server (pid %s) survived on :%s; reaping",
launcher, pid, recorded or "?")
_kill_group_then_wait(pgid if pgid == launcher else None, pid)
if recorded:
(_X_LOCK_DIR / f".X{recorded}-lock").unlink(missing_ok=True)
for name in ("launcher.pid", "env", "rfb.sock"):
(sd / name).unlink(missing_ok=True)
return True
def _kill_group_then_wait(pgid: Optional[int], pid: int, grace: float = 2.0) -> None:
"""SIGTERM the group (or the lone pid), SIGKILL whatever is still there after ``grace``."""
def _signal(sig: int) -> None:
with contextlib.suppress(ProcessLookupError, PermissionError):
if pgid is not None:
os.killpg(pgid, sig) # windows-footgun: ok — Linux-only runtime (is_supported_host gates start/stop)
else:
os.kill(pid, sig)
def _anything_left() -> bool:
# The leader dying first is the common case (bash exits on TERM, Xvnc traps it); the group is
# done only when killpg(0) finds nobody, else a TERM-ignoring descendant keeps the display.
if pgid is None:
return _pid_alive(pid)
try:
os.killpg(pgid, 0) # windows-footgun: ok — Linux-only runtime (is_supported_host gates start/stop)
except ProcessLookupError:
return False
except PermissionError:
return True
return True
def _reap_if_ours() -> None:
# The launcher was Popen'd by whichever gateway started it; a later gateway that stops it holds
# no Popen, so the dead leader would sit as a zombie in our table (and count as "left").
with contextlib.suppress(ChildProcessError, OSError):
os.waitpid(pid, os.WNOHANG) # windows-footgun: ok — Linux-only runtime (is_supported_host gates start/stop)
_signal(signal.SIGTERM)
deadline = time.monotonic() + grace
while time.monotonic() < deadline:
_reap_if_ours()
if not _anything_left():
return
time.sleep(0.05)
_signal(signal.SIGKILL) # windows-footgun: ok — Linux-only runtime (is_supported_host gates start/stop)
time.sleep(0.1)
_reap_if_ours()
# Host-wide (every profile allocates from one band), so it lives outside any profile home. A predictable
# name must not be squattable: XDG_RUNTIME_DIR is the boundary — 0700 from logind, or from
# docker/stage2-hook.sh in containers, which have none.
_ALLOC_LOCK = Path(os.environ.get("XDG_RUNTIME_DIR") or Path.home() / ".cache") / "hermes-bot-desktop-alloc.lock"
@contextlib.contextmanager
def _flocked(path: Path):
import fcntl # windows-footgun: ok — Linux-only runtime (is_supported_host gates start)
with open(path, "a+", encoding="utf-8") as fh: # windows-footgun: ok — Linux-only runtime
fcntl.flock(fh.fileno(), fcntl.LOCK_EX)
try:
yield fh
finally:
fcntl.flock(fh.fileno(), fcntl.LOCK_UN)
def _pick_display() -> int:
"""Caller holds ``_ALLOC_LOCK``. The recorded number is only reused when no OTHER server holds it now:
after profile A stops, B may have taken A's number, and A's launcher must never unlink B's socket."""
recorded = _read(state_dir() / "display")
if recorded and recorded.isdigit() and not _display_in_use(int(recorded)):
return int(recorded)
for num in range(_DISPLAY_MIN, _DISPLAY_MAX + 1):
if not _display_in_use(num):
return num
raise RuntimeError("no free X display number in the Bot Desktop band")
def _allocate_display() -> int:
_ALLOC_LOCK.parent.mkdir(parents=True, exist_ok=True, mode=0o700)
with _flocked(_ALLOC_LOCK):
return _pick_display()
def desktop_env(base_env: Optional[Dict[str, str]] = None) -> Dict[str, str]:
"""``base_env`` (default ``os.environ``) with this profile's DISPLAY/XAUTHORITY/DBUS_SESSION_BUS_ADDRESS
merged in when its desktop is running. Unchanged otherwise, so hosts with a real seat keep it.
Pure: never starts anything (it is called from env builders, status probes and tests)."""
env = dict(os.environ if base_env is None else base_env)
published = published_env()
if published:
touch_activity() # a browser / cua-driver spawn is the agent using its screen
env.update(published)
env.pop("WAYLAND_DISPLAY", None) # X11 desktop; a leaked Wayland socket flips GTK/Chromium backends
from tools.bot_desktop.browser import env_for_agent
env_for_agent(env) # same binary + user-data-dir as the dock's Browser icon
return env
def ensure_started_for_tool() -> None:
"""Tool-boundary hook (``computer_use`` dispatch and the headed Chromium spawn sites of the browser tool).
Under ``terminal`` placement the sandbox screen is brought up on first use unconditionally: the sandbox is
the boundary the user chose, a screen inside it touches nothing outside it, and the tools below refuse
rather than run on the host when it is not up. On the gateway host, ``bot_desktop.auto_start`` (opt-in,
default off) starts a Linux host's screen when it has NO display and the packages installed, so a headless
gateway works the first time. Host failure is not an error here (the tool's own "no display" diagnosis is
the right message then); a sandbox start failure is logged and left for :func:`tool_placement` to raise."""
if published_env().get("DISPLAY"):
touch_activity()
return
where = placement.resolve()
if where.where == placement.REFUSED:
return # tool_placement() raises the reason at the spawn site
if where.where == placement.TERMINAL:
try:
start()
except Exception as exc:
logger.info("Bot Desktop sandbox start deferred to the tool: %s", exc)
return
if not _should_auto_start(os.environ):
return
try:
start()
except Exception as exc:
logger.info("Bot Desktop auto-start skipped: %s", exc)
def tool_placement() -> str:
"""Where the browser / cua-driver a tool is about to spawn MUST run: ``placement.GATEWAY`` or
``placement.TERMINAL``. Policy first, liveness second: a ``terminal`` placement whose screen is down gets
it started here (raising with the blocker when it cannot come up) and a ``refused`` placement raises its
reason. Neither ever yields the host — a screen that happens to be down is not permission to run the
agent's browser outside the sandbox the user chose."""
where = placement.resolve()
if where.where == placement.REFUSED:
raise RuntimeError(where.reason)
if where.where == placement.GATEWAY:
return placement.GATEWAY
if not sandbox_screen_running() or not published_env().get("DISPLAY"):
start()
return placement.TERMINAL
def _should_auto_start(env: Dict[str, str]) -> bool:
if not is_supported_host() or env.get("DISPLAY") or env.get("WAYLAND_DISPLAY"):
return False
if missing_binaries():
return False
from hermes_cli.config import load_config_readonly
cfg = load_config_readonly().get("bot_desktop") or {}
return bool(cfg.get("auto_start", False))
# ---- idle auto-stop -------------------------------------------------------------------------------
# A desktop nobody is using still holds ~220 MiB (plus whatever browser was left open). Every use —
# a computer_use action, a browser spawn onto the screen, a viewer attached, a human takeover — stamps
# ``activity``; the gateway's display watcher stops a screen idle past ``bot_desktop.idle_stop_minutes``
# unless a human holds it. The next use starts it again (auto_start or the pane's Start).
DEFAULT_IDLE_STOP_MINUTES = 30
def touch_activity() -> None:
path = state_dir() / "activity"
try:
path.touch()
os.utime(path, None)
except OSError:
pass
def idle_seconds() -> Optional[float]:
"""Seconds since the last stamped use; None when the screen never recorded one (falls back to the
env file's publish time so a screen started and then forgotten still ages)."""
for name in ("activity", "env"):
try:
return max(0.0, time.time() - (state_dir() / name).stat().st_mtime)
except OSError:
continue
return None
def idle_stop_seconds() -> float:
from hermes_cli.config import load_config_readonly
cfg = load_config_readonly().get("bot_desktop") or {}
try:
minutes = float(cfg.get("idle_stop_minutes", DEFAULT_IDLE_STOP_MINUTES))
except (TypeError, ValueError):
minutes = DEFAULT_IDLE_STOP_MINUTES
return max(0.0, minutes) * 60
def stop_if_idle() -> bool:
"""Stop this profile's screen when it has been idle past the limit and no human holds it. True when
it was stopped."""
limit = idle_stop_seconds()
if limit <= 0 or not is_running():
return False
idle = idle_seconds()
if idle is None or idle < limit:
return False
from tools.bot_desktop import lease as _bd_lease
if _bd_lease.get().holder == _bd_lease.HUMAN:
return False
logger.info("Bot Desktop for profile %s idle for %.0f min; stopping", _profile_name(), idle / 60)
return stop()
def published_env() -> Dict[str, str]:
"""Variables the launcher wrote once Xfce's private bus existed; empty when the desktop is down.
Pure file reads on the common path: this is called from every browser / cua-driver env builder, so it
must not load config (that initializes HERMES_HOME). A sandbox-hosted screen leaves a host-side marker at
start; only its presence routes to the sandbox probe."""
from tools.bot_desktop import sandbox_host
if sandbox_host._read_marker():
return _sandbox_published_env()
if _launcher_pid() is None:
return {}
raw = _read(state_dir() / "env")
if not raw:
return {}
out: Dict[str, str] = {}
for line in raw.splitlines():
key, sep, value = line.partition("=")
if sep:
out[key.strip()] = value.strip()
return out
def rfb_socket_path() -> Optional[Path]:
"""Host path of the RFB socket; None when down OR when the screen lives in a sandbox (use
:func:`open_rfb_stream` there: the socket is not on this filesystem)."""
from tools.bot_desktop import sandbox_host
if sandbox_host._read_marker():
return None
sock = state_dir() / "rfb.sock"
return sock if _launcher_pid() is not None and sock.exists() else None
def in_sandbox() -> bool:
"""True when this profile's screen is placed inside the terminal backend (policy; reads config)."""
return placement.resolve().where == placement.TERMINAL
def sandbox_screen_running() -> bool:
"""True when a sandbox-hosted screen is UP for this profile: the start marker exists AND the sandbox it
names is alive — registered in this process, or (after a gateway restart emptied the registry) the
marker's recorded container still running. Only a sandbox that is provably gone (its container removed
out from under us) drops the marker, so the next start rebuilds instead of the browser exec-ing into a
dead container; an unregistered-but-alive one is re-attached by ``_sandbox_env(create=True)`` at the
next spawn (the terminal planner reuses the persisted container by label)."""
from tools.bot_desktop import sandbox_host
marker = sandbox_host._read_marker()
if not marker:
return False
if _sandbox_env(create=False) is not None:
return True
if sandbox_host.marker_sandbox_alive(marker):
return True
sandbox_host._marker().unlink(missing_ok=True)
return False
def _owned_sandbox_env():
"""The environment hosting the screen the marker records, re-attaching after a restart when the recorded
sandbox is still alive; None when there is no marker or its sandbox is gone. Never builds a sandbox for
a screen that is not there."""
from tools.bot_desktop import sandbox_host
marker = sandbox_host._read_marker()
if not marker:
return None
env = _sandbox_env(create=False)
if env is not None:
return env
if sandbox_host.marker_sandbox_alive(marker):
return _sandbox_env(create=True)
return None
def is_running() -> bool:
return bool(published_env().get("DISPLAY"))
def open_rfb_stream() -> "subprocess.Popen":
"""Popen whose stdin/stdout carry RFB bytes for a sandbox-hosted screen (``in_sandbox()`` only)."""
from tools.bot_desktop import sandbox_host
env = _sandbox_env(create=False)
if env is None:
raise RuntimeError("the sandbox hosting this screen is not running")
return sandbox_host.open_rfb_stream(env, _profile_name())
def geometry() -> str:
from hermes_cli.config import load_config_readonly
cfg = load_config_readonly().get("bot_desktop") or {}
return str(cfg.get("geometry") or "1440x900")
def status(profile: Optional[str] = None) -> DesktopStatus:
from tools.bot_desktop import browser as _bd_browser
from tools.bot_desktop import resources
from tools.bot_desktop import sandbox_host
where = placement.resolve()
if where.where == placement.TERMINAL or sandbox_host._read_marker():
# A screen already running inside a sandbox is reported (and stoppable) even after the placement
# setting moved: the recorded owner wins over the current policy until it is stopped.
return _sandbox_status(profile, where)
missing: list[str] = missing_binaries() if is_supported_host() else list(REQUIRED_BINARIES)
pid = _launcher_pid()
env = published_env()
running = pid is not None and bool(env.get("DISPLAY"))
mem = resources.memory_info() if is_supported_host() else resources.MemoryInfo(None, None)
return DesktopStatus(
profile=profile or _profile_name(),
supported=is_supported_host(),
installed=not missing,
missing=missing,
running=running,
pid=pid,
display=env.get("DISPLAY"),
socket=str(rfb_socket_path()) if rfb_socket_path() else None,
geometry=geometry(),
install_command=install_command() if missing else None,
browser=_bd_browser.executable() if is_supported_host() else None,
# A running screen is never "blocked": the check guards the allocation, not the session.
blocker=None if running or missing or not is_supported_host() else resources.memory_blocker(mem),
memory_available_mb=mem.available_mb,
memory_limit_mb=mem.limit_mb,
)
def _sandbox_status(profile: Optional[str], where) -> DesktopStatus:
"""Status of a sandbox-placed screen. Package presence is only known once the sandbox exists; before
that the pane shows "installed" with the image hint carried in ``install_command`` so Start can explain."""
from tools.bot_desktop import sandbox_host
env = _owned_sandbox_env() or _sandbox_env(create=False)
missing = sandbox_host.missing_binaries(env) if env is not None else []
published = sandbox_host.published_env(env, profile or _profile_name()) if env is not None else {}
running = bool(published.get("DISPLAY"))
# No install_command: the pane's Install button runs apt on the HOST, which is the wrong machine here.
# A sandbox missing the stack is a blocker (shown in place of Start) naming the image that has it.
blocker = None
image_switch = None
if missing:
blocker = (f"The terminal backend's sandbox image lacks {', '.join(missing)}. Use "
f"{sandbox_host.SANDBOX_IMAGE_HINT} as terminal.{where.backend}_image (the default sandbox base "
f"plus the desktop stack), or set bot_desktop.placement: gateway.")
if where.backend == "docker":
# The usual reason on an upgraded install: the persisted container predates the default
# flip and was kept on purpose. The pane offers the switch instead of a config hint.
from hermes_cli.sandbox_image_switch import pending
sw = pending()
if sw is not None:
image_switch = {"current_image": sw.current_image, "target_image": sw.target_image,
"containers": len(sw.containers)}
blocker = (f"Your sandbox container still runs {sw.current_image}, which has no desktop. "
f"Switch it to {sw.target_image}: files in /root and /workspace stay, packages "
f"installed inside the container are reinstalled on demand.")
return DesktopStatus(
profile=profile or _profile_name(),
supported=True,
installed=True,
missing=missing,
running=running,
pid=None,
display=published.get("DISPLAY"),
socket=None,
geometry=geometry(),
install_command=None,
browser=None,
blocker=blocker,
memory_available_mb=None,
memory_limit_mb=None,
placement=f"{placement.TERMINAL}:{where.backend}",
image_switch=image_switch,
)
def _profile_name() -> str:
try:
from hermes_cli.profiles import get_active_profile_name
return get_active_profile_name() or "default"
except Exception:
return "default"
# The gate lives in ``resources`` so start() and status() cannot disagree about it. Measured in the
# official image: gateway idle 304 MiB, +216 for Xvnc/Xfce, 1073 MiB with one Chromium page. The OOM
# killer picks by score, so on a small instance the casualty is the dashboard or the gateway, not the
# desktop that caused the pressure.
def start(*, wait_seconds: float = 15.0) -> DesktopStatus:
"""Start this profile's desktop (idempotent). Blocks until the launcher publishes its env file or
``wait_seconds`` pass; raises ``RuntimeError`` naming the blocker.
``bot_desktop.placement`` decides WHERE: inside the configured terminal backend (docker/ssh/singularity;
``sandbox_host``), or on the gateway host (the rest of this function). A sandbox backend that cannot host
a screen refuses rather than silently falling back to the host beside it.
Two locks: the per-profile ``start.lock``, held from the running-check to the launcher's publish so two
start() calls for one profile spawn one launcher (the loser sees it running), and the host-wide
display-allocation lock, held only until this Xvnc has written ``/tmp/.X<n>-lock`` (a second profile
picking the same number before that would fail and its stale-lock cleanup could remove our socket).
Holding it for the whole Xfce bring-up serialized every profile's start behind one desktop launch."""
where = placement.resolve()
if where.where == placement.REFUSED:
raise RuntimeError(where.reason)
if where.where == placement.TERMINAL:
return _start_in_sandbox(wait_seconds)
if not is_supported_host():
raise RuntimeError("Bot Desktop runs on Linux gateway hosts only")
missing = missing_binaries()
if missing:
# Three dead ends: an operator told "unprivileged, no sudo" while running as root hunts the wrong bug.
need = f"Bot Desktop needs {', '.join(missing)} on the gateway host"
if package_manager() is None:
raise RuntimeError(
f"{need}, and no supported package manager (apt/dnf/pacman) is available to install them. "
"Install TigerVNC (Xvnc) and the Xfce core components with this distro's own tooling.")
if not installable():
raise RuntimeError(
f"{need}, and this host cannot install them: the process is unprivileged and there is no "
"sudo. On the published Docker image the packages have to be baked in, so this needs a "
"newer image rather than an install.")
raise RuntimeError(f"{need}. Install: {install_command()}")
sd = state_dir()
sd.mkdir(parents=True, exist_ok=True)
os.chmod(sd, 0o700)
with _flocked(sd / "start.lock"):
if _launcher_pid() is not None and published_env().get("DISPLAY"):
return status()
from tools.bot_desktop import resources
floor = resources.min_free_mb()
mem = resources.memory_info()
if (blocker := resources.memory_blocker(mem, need=floor)) is not None:
raise RuntimeError(blocker)
if mem.available_mb is not None and mem.available_mb < resources.tight_headroom_mb(floor):
logger.warning(
"Bot Desktop starting with %d MB available; a browser with a few pages open can use most "
"of that.", mem.available_mb)
if _launcher_pid() is None:
_reap_orphaned_server(sd)
_ALLOC_LOCK.parent.mkdir(parents=True, exist_ok=True, mode=0o700)
try:
return _spawn_and_wait(sd, wait_seconds)
except RuntimeError:
# _pick_display reuses the recorded number first: left in place after a failed launch (e.g. Xvnc
# 'server already running' on it), every retry would pick the same number and the profile wedges.
(sd / "display").unlink(missing_ok=True)
raise
def _spawn_and_wait(sd: Path, wait_seconds: float) -> DesktopStatus:
"""Caller holds ``start.lock``. Takes ``_ALLOC_LOCK`` itself, from picking the number to Xvnc's claim."""
with contextlib.ExitStack() as alloc:
alloc.enter_context(_flocked(_ALLOC_LOCK))
num = _pick_display()
(sd / "display").write_text(str(num), encoding="utf-8")
env_file = sd / "env"
env_file.unlink(missing_ok=True)
child_env = {k: v for k, v in os.environ.items() if k not in {
"DISPLAY", "XAUTHORITY", "WAYLAND_DISPLAY", "DBUS_SESSION_BUS_ADDRESS", "SESSION_MANAGER"}}
child_env.update({
"HERMES_BD_PROFILE": _profile_name(),
"HERMES_BD_DISPLAY_NUM": str(num),
"HERMES_BD_SOCKET": str(sd / "rfb.sock"),
"HERMES_BD_XAUTH": str(sd / "Xauthority"),
"HERMES_BD_ENV_FILE": str(env_file),
"HERMES_BD_CONFIG_HOME": str(sd / "xdg"),
"HERMES_BD_GEOMETRY": geometry(),
})
from tools.bot_desktop.browser import dock_exec_line, dock_launch
if (browser := dock_launch()) is not None:
# The bare executable (the launcher checks it exists) and the ready-made, spec-quoted Exec= line.
child_env["HERMES_BD_BROWSER_EXEC"] = browser[0]
child_env["HERMES_BD_BROWSER_EXEC_LINE"] = dock_exec_line(*browser)
# Truncated per start: the log is a diagnostic for THIS launch, and nothing rotates it otherwise.
log = open(sd / "launcher.log", "wb") # noqa: SIM115 — handed to the child, closed by it
proc = subprocess.Popen( # windows-footgun: ok — Linux-only runtime (is_supported_host)
["bash", str(_LAUNCHER)], env=child_env, stdin=subprocess.DEVNULL, stdout=log, stderr=log,
start_new_session=True, close_fds=True)
log.close()
born = _create_time(proc.pid)
(sd / "launcher.pid").write_text(f"{proc.pid} {born if born is not None else 0}", encoding="utf-8")
deadline = time.monotonic() + wait_seconds
while time.monotonic() < deadline:
if _x_lock_pid(num) is not None:
alloc.close() # the number is Xvnc's now; other profiles may allocate (idempotent)
if proc.poll() is not None:
tail = (sd / "launcher.log").read_bytes()[-2000:].decode("utf-8", "replace")
raise RuntimeError(f"Bot Desktop launcher exited with {proc.returncode}:\n{tail}")
if env_file.exists() and (sd / "rfb.sock").exists():
logger.info("Bot Desktop for profile %s up on :%s", _profile_name(), num)
touch_activity()
return status()
time.sleep(0.1)
# Giving up must take the launch down: left alone, the launcher publishes DISPLAY and rfb.sock a moment
# later and a screen whose start() reported failure stays up as "running". The launcher is its own
# session leader, so its group is exactly this launch (Xvnc, dbus, Xfce) and nothing else.
_kill_group_then_wait(proc.pid, proc.pid)
proc.wait()
for name in ("launcher.pid", "env", "rfb.sock"):
(sd / name).unlink(missing_ok=True)
raise RuntimeError(f"Bot Desktop did not publish its display within {wait_seconds:.0f}s (see {sd / 'launcher.log'})")
def _sandbox_env(*, create: bool):
"""The terminal environment hosting this profile's screen, or None. ``create=False`` for status probes
(a status call must never build a container)."""
return placement.terminal_environment(create=create)
def _start_in_sandbox(wait_seconds: float) -> DesktopStatus:
from tools.bot_desktop import sandbox_host
env = _sandbox_env(create=True)
if env is None:
raise RuntimeError("the terminal backend's sandbox could not be started, so there is nowhere to put the screen")
sd = state_dir()
sd.mkdir(parents=True, exist_ok=True)
os.chmod(sd, 0o700)
with _flocked(sd / "start.lock"):
# The dock's Browser icon runs the sandbox's own Playwright Chromium on the profile agent-browser
# uses there: the human's browser is the bot's browser, as on the host.
from tools.bot_desktop.browser import dock_exec_line
exe = sandbox_host.chromium_executable(env)
browser = (exe, dock_exec_line(exe, sandbox_host.browser_profile_dir(env), sandbox_bypass=True)) if exe else (None, None)
published = sandbox_host.start(env, _profile_name(), geometry=geometry(), wait_seconds=max(wait_seconds, 20.0),
browser_exec=browser[0], browser_exec_line=browser[1])
(sd / "env").write_text("".join(f"{k}={v}\n" for k, v in published.items()), encoding="utf-8")
touch_activity()
return status()
def _sandbox_published_env() -> Dict[str, str]:
"""Published env of a sandbox-hosted screen (the caller saw the host-side marker written at start)."""
from tools.bot_desktop import sandbox_host
env = _sandbox_env(create=False)
if env is None:
return {}
return sandbox_host.published_env(env, _profile_name())
def stop() -> bool:
"""Stop this profile's desktop; True when a running launcher (or the X server a dead one left behind)
was signalled."""
from tools.bot_desktop import sandbox_host
if sandbox_host._read_marker() or placement.resolve().where == placement.TERMINAL:
# The marker names the sandbox that owns the screen; stop THAT one (re-attaching after a restart),
# whatever the placement setting says now. A marker whose sandbox is gone is simply dropped.
env = _owned_sandbox_env()
stopped = sandbox_host.stop(env, _profile_name()) if env is not None else False
sandbox_host._marker().unlink(missing_ok=True)
for name in ("env", "activity"):
(state_dir() / name).unlink(missing_ok=True)
return stopped
if not is_supported_host():
return False
sd = state_dir()
sd.mkdir(parents=True, exist_ok=True)
with _flocked(sd / "start.lock"):
return _stop_locked(sd)
def _stop_locked(sd: Path) -> bool:
pid = _launcher_pid()
if pid is None:
reaped = _reap_orphaned_server(sd)
(sd / "env").unlink(missing_ok=True)
return reaped
# The launcher runs in its own session; killing the group takes Xvnc, dbus and Xfce with it.
_kill_group_then_wait(pid, pid, grace=5.0)
for name in ("launcher.pid", "env", "activity"):
(sd / name).unlink(missing_ok=True)
return True