A hand-written Linux daemon unit (systemd user unit / XDG autostart entry
running `cua-driver serve`) can be dead for days — crash loop, stopped, never
started — while `hermes computer-use doctor` and `status` report a healthy
binary: the runtime contract only checks the binary (`manifest`), and nothing
ever connected to the daemon socket (#114748).
- tools/computer_use/cua_backend.py::cua_daemon_listening — socket-level
liveness via `cua-driver status [--socket PATH]` (rc 0 = a daemon answered,
"not running" = dead, anything else = unknown). Never raises.
- tools/computer_use/doctor.py::cua_daemon_units — one scan of the units that
run cua-driver (kind, unit, exec target, runs `serve`, `--socket` path with
`%h` expanded); the pruned-Exec guard now filters that list instead of
re-scanning.
- doctor.py::_apply_daemon_liveness_guard — per `serve` unit: `pass` when its
socket answers, `fail` (degrading `ok`) when not, with the hint that a driver
reinstall does not start the daemon. Unconfigured daemons are never probed:
on Linux the MCP runtime needs none, so a silent default socket is normal.
- `hermes computer-use status` prints the dead-daemon line and exits 1.
- docs: the Linux daemon-unit paragraph under the doctor section.
Not changed: _maybe_repair_runtime_contract. The repair is gated on the
binary-level contract only; a dead daemon never enters it, so a reinstall was
never triggered by the daemon (the `.release_installed/<version>` marker the
reporter saw is written by cua-driver itself on any first run of the binary —
observed via strace of `cua-driver status` under a fresh HOME).
Supersedes #114928 (@Finn763): same finding, ~600-line implementation with a
new module and a repair-gate rewrite; this is the ~90-line version on the
existing doctor seams.