After _CRASH_LIMIT consecutive spawn/execution failures the breaker opened and
never closed again: check_command_security returns early while _circuit_open
is set, so the reset below it was unreachable for the rest of the process. A
transient failure (a slow disk during install, a binary being replaced) meant
every later command silently skipped scanning until restart — a security
control that fails open permanently, not temporarily.
The breaker now half-opens after _CIRCUIT_RETRY_S: exactly one caller claims
the probe slot and runs a real scan. Claiming re-arms _circuit_open_at under
_breaker_lock first, so concurrent callers see a fresh TTL and stay on the
fail-open path — single-flight, one probe per TTL window. The lock guards only
the claim (a TTL comparison and a timestamp write); it is never held across
the subprocess, so it cannot reintroduce the #41400 hang. Crash counting stays
lock-free, matching the mcp_tool.py error counters.
Any completed scan closes the breaker. exit 0/1/2 are all verdicts —
allow/block/warn — and all three prove the binary is healthy, so recovery no
longer depends on the next command happening to be clean. That also fixes the
streak never resetting on block/warn.
time.monotonic() throughout, matching the MCP breaker: a wall-clock jump
backwards must not strand the breaker open past its TTL.
Rebased onto current main; the file was reformatted upstream in the tools/
compaction wave, so the change is re-expressed in the current structure
(_verdict / _EXIT_ACTIONS). The latch itself is unchanged upstream.
(cherry picked from commit 2ff4ee6c745060e9aaf4f4ee2d04ac68ba55b2e8)