EmptyCompletionError is raised by complete_task; without a handler the
CLI and the tool would surface a traceback instead of the actionable
refusal (add --result/--summary). Same shape as the HallucinatedCards gate.
Part of #117483
When the owner dies before a background child records a result, the
recovered event now includes the last lines of each live transcript and a
one-line git snapshot of the owner cwd, and the single-delegation renderer
shows the recovery diagnostics the batch renderer already had.
Fixes#116000
_Capture.read polls read_available against the stop event before every
InputStream.read, close() aborts before closing, and _halt_thread joins
outside the detector lock and keeps the handle while the thread is alive.
Fixes#117096
The binding half of _verify_reapable_browser_daemon accepted the socket-dir
basename anywhere in argv. That basename is predictable from the session
name, so a recycled PID whose argv merely mentioned it (a grep, a shell)
passed both gates and was tree-killed. Require the full normalized path as
an argv token (bare or --flag=path) or the environ match.
Fixes#116884
Review follow-up: _request_google_offline_access re-wrapped
context.redirect_handler on every _perform_authorization, so N
re-authorizations ran N nested wrappers. Mark the wrapper and return early
when it is already installed; it reads the issuer at call time so the
single wrapper is right for any issuer. The six test functions fold into
two — browser flow (offline access + consent once, wrap-once, other-issuer
and lookalike controls) and device flow (issuer slash-normalization +
access_type on the device request).
Google issues refresh tokens only for authorization requests carrying
access_type=offline (its idiom where OIDC servers use the offline_access
scope MCP discovery would advertise), so a Google-hosted MCP server
(Gmail/Calendar) authorized in a browser dies with the short-lived access
token: later reconnects (gateway, cron) find no refresh token and fail
back to an interactive login they cannot perform (#117510).
The device flow could not even start: the advertised authorization server
reaches issuer validation root-slash-stripped on one side only, so
Google's slash-less document issuer is rejected against a slash-terminated
expected issuer. Compare both sides normalized, the convention
_metadata_issuer and the refresh-token issuer binding already use; any
other mismatch is still rejected.
(cherry picked from commit bc08ef9572616d09ecfb9a0085b32b3582af5f5e)
Review follow-up for #116905: bot_mode_probe was the third private copy of
the regex (profiles, update_restart_recovery) with a sync test policing the
drift the copy created. One canonical object next to named_profile_is_live;
the three modules import it and the sync test is gone.
A marker-carrying directory whose name is not a profile id (a parked backup
like _backup_removed_20260920, a .staging-area dotdir) passed the roster's
identity predicate and was injected as a phantom teammate into every Bot's
system prompt, while profile list hides it via _PROFILE_ID_RE. Add the same
name gate to _roster so the roster agrees with every other profiles/
enumerator; friendly-name aliases and the message_agent target set no longer
admit dirs the CLI refuses to treat as profiles.
Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
A spawner that skipped setsid (the Darwin gateway's posix_spawn shim) leaves
tool children in the gateway's process group, and _kill_process_group_posix
then killpg()s the gateway itself — launchd logs 'Killed: 9' and KeepAlive
respawns it. Branch on pgid == os.getpgrp() as a plain if/else (not a raised
PermissionError routed into the EPERM handler) and signal the wrapper plus
its snapshotted descendants by PID; the EPERM fallback shares that helper.
Fixes#107029
On macOS a search that reaches its `limit` takes the early-stop branch that
TERMs rg's process group; rg has often already exited, and macOS answers
killpg on a zombie-only group with EPERM instead of ESRCH. The
PermissionError escaped _kill_process_group_posix, unwound to search_tool
and the drained matches were replaced by
`{"error": "[Errno 1] Operation not permitted"}` (#116855, same symptom
as #104696). The helper now treats EPERM like "nothing left to signal"
and falls back to killing the known PIDs so a live child cannot escape.
It also never killpg's the caller's own group (#107029): a child that
shares our pgid is torn down by PID.
Diagnosis credit: #116949 (@liuhao1024), redone slim.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Review follow-up: hoist the credential-path alternation the JavaScript and
Python read-secrets patterns inlined three times into _CRED_FILE /
_CRED_FILE_LITERAL / _PY_CRED_FILE_ARG (compiled patterns unchanged), and
fold the four scan_file tests into two — one drives plugin_guard.scan_plugin
on a plugin dir containing the open() read and asserts the verdict flips to
dangerous, the other pins write modes, .pub and unrelated files as clean.
RED-verified by review: open(os.path.expanduser(...))/Path(os.path.expanduser(...))
required a literal quote right after the opening paren, and Path(...) only matched
.read_text(), so read_bytes()/readlines()/readline() on a credential file stayed clean.
(cherry picked from commit 3e3b1fa6e374ba21ed7de1a2ee03b6f45513779a)
py_read_secrets_file's open(...) branch fired on any open() call naming a
credential-shaped path regardless of mode, so open(".env", "w")/open(".npmrc",
"a") -- a plugin/skill writing its own setup files, not exfiltration -- landed
critical and unconditionally blocked install. Exclude a literal write/append/
exclusive mode (2nd-arg string or mode= kwarg containing w/a/x), mirroring the
shell sibling's `cat > file` exclusion. Path(...).read_text() has no mode
argument so needs no change.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 545787712a8d721b9940a2104caee9dc1aac04ac)
`tools/skills_guard.py`'s THREAT_PATTERNS covers shell `cat <secrets-file>`
(read_secrets_file) and JS readFileSync/readTextFile (js_read_secrets_file)
at critical severity, but the same Python read (`open("~/.hermes/.env")`,
`Path(...).read_text()`) only tripped the `hermes_env_access` mention
pattern, which plugin_guard's SEVERITY_REMAP demotes to medium — so a
Python plugin reading a credential file passed admission as "safe".
Adds py_read_secrets_file, the Python twin of js_read_secrets_file, over
the same credential-file-name class (id_rsa/id_ed25519/etc, .env,
credentials, .netrc, .pgpass, .npmrc, .pypirc).
Fixes#116950
(cherry picked from commit 02758d94bf75dea76fa43d10a693febe72f5fb6f)
Path.write_text(newline=None) translates LF to os.linesep on Windows, so the
quarantined SKILL.md landed as CRLF while bundle_content_hash re-encodes the
str as LF bytes — content_hash of the installed copy could never match and
every hub skill reported update_available forever (#117181). Pass newline=""
to keep the bundle's bytes verbatim, mirroring the bytes branch.
A reminder created from a top-level Slack message delivered to the channel
root while its confirmation landed in the thread the assistant opened under
that message: with platforms.slack.extra.reply_in_thread at its default true
the whole exchange lives in the thread keyed on the asking message's own id,
so the synthetic per-message drop rule threw away the conversation's actual
location. Scope the drop by the job's first-fire horizon: a one-shot firing
within 60 minutes keeps its thread; recurring jobs and one-shots further out
keep today's flat delivery, which is what the rule protects.
Fixes#117306
_windows_bash_candidates in tools/environments/local.py appended
shutil.which("bash") to the candidate list without filtering — when Git Bash
is installed but PATH resolves `bash` to WSL's stub (C:\Windows\System32\bash.exe)
or WindowsApps bash.exe, that stub enters the candidate list.
Fix: skip paths whose normalized form contains `system32` or `windowsapps`,
logging at debug level. Git Bash (already searched explicitly in the roots list)
remains the first usable candidate.
Fixes#116818
A child that stalls under a configured delegation.child_timeout_seconds used to
learn about the budget only by dying, losing its whole context. The liveness
wait now queues a one-line "[delegation budget warning]" through the child's
steer channel once the idle window is 80% spent (delivered at the child's next
iteration boundary), so a slow-but-recoverable child can wrap up and return its
summary. The warning fires once per idle window and re-arms when progress
resets the window; a progressing child never sees it.
Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged.
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a
plain `settled.wait(timeout=child_timeout)`, so any child that outlived the
budget was abandoned even while the provider was actively serving it.
The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could
not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint
— but the mechanism it names is exactly this gate: a child waiting on an
in-flight LLM completion, killed with a nearly-finished context. A slow child is
already bounded elsewhere (the per-call stale watchdog, the heartbeat's
staleness verdict), so this cap could only ever kill children the runtime had
judged healthy.
`child_timeout_seconds` now measures time with NO progress: the wait runs in
slices and restarts the window on the same signals the heartbeat's stale verdict
reads (completed call, tool change, activity-clock tick). A frozen child is
still abandoned when the window elapses; a progressing one is never killed for
taking long.
Timeout entries also carry `last_event_age` (how long the child had been silent),
so operators can tell a slow provider from a runaway without transcript
forensics.
Fixes the mechanism reported in #116001. The budget warning and
continuation-respawn items in that issue are separate features and are not part
of this change.
The idle sweep lives inside _acquire_kernel, so it only fires when the next
kernel request arrives. A host that stays alive but stops executing anything
(a pids-exhausted container fail-closing every tool call) never acquires
again: idle kernels, their runners and thread pools survive indefinitely,
which is how four orphaned runners helped exhaust pids_limit=256 for hours
(#117169). A low-frequency daemon reaper — started on the first kernel
spawn, using the acquire-path criteria verbatim — retires idle kernels on
its own schedule, independent of tool traffic.
The same pass sweeps hermes_kernel_* staging dirs untouched for over a
week: a live host rmtrees each dir within one idle timeout of the kernel's
last use, so a week-old dir belongs to a host that died without cleanup
(SIGKILL / container restart). Younger dirs are left alone because a
concurrently running host's live kernel may own one; rmtree rejects
symlinks rather than following them.
Review follow-up: the split delegate guard dropped singular `child` and the
`to` determiner from the possessed-worker group, so `Send the child the
context it needs` and `Send to workers the context they need` were flagged
context_exfil. Restore both; bare `child`/`workers` after the verb is still
exfil.
The delegation skip treated any child/workers/delegates word after send/share as an in-process handoff. That hid Send child context to the operator. Keep the skip for subagents and for possessed workers (each worker), and match the bare-recipient forms.
(cherry picked from commit 8d6d500ae0bdf8214e0f4ecf5561eea8d8a79f6c)
The PATHEXT retry swapped os.environ["PATHEXT"] around shutil.which, publishing one
server's per-profile value to every other thread in a multiplexed gateway for the
duration of the lookup; the extension walk now happens inside the resolver instead.
When a server env omits PATH, shutil.which(cmd, path=None) silently consulted the
parent's PATH, letting a command resolve against an env the child never sees; the
lookup is now skipped (absent PATH is a miss; empty PATH keeps its cwd-only meaning).
A 200 response with empty content (finish_reason="length" after a
reasoning model spends the whole max_tokens budget on hidden reasoning)
mapped to ESCALATE via a plain _VERDICTS dict miss with no log above
DEBUG — indistinguishable from a genuine strict verdict for days
(#117428, failure mode of #108163/#68263). The exception path already
warns (#93809); this closes the truncation path's asymmetry.
Fixes#117428
The dangerous-pattern matcher lowercases its input and compiles every pattern
with re.IGNORECASE, which erases the only difference between git's safe
merged-only branch delete (-d, refused by git when unmerged) and the
destructive force delete (-D). The result was backwards: -d was gated as
"git branch force delete" while --delete, the long spelling of the same safe
flag, passed untouched.
Scope the flag group as (?-i:-D) so this one pattern opts out of the
module-wide case folding, and prepare the detection input with
_lower_preserving_flags(), which lowercases everything except dash-prefixed
tokens. All other patterns match case-insensitively, so preserved flag case is
invisible to them. Approvals stored under the pattern's pre-change legacy key
keep matching via an alias.
Rebuilt from #32876, which proposed this direction before the approval.py ->
approval_detection.py split left it unmergeable.
Artifact staging is atomic with the completion write, so a worker's
pre-completion kanban_attachments readback is always empty and workers
narrated "registered at completion: none" even when the rows landed
(#117360). Return the card's durable attachment set in the completion
result, in the same shape the readback tool returns.
On the reporting macOS host with fake-IP TUN DNS, the resolver answers an IPv4
name with the plain address and its IPv4-translated form ::ffff:0:a.b.c.d
(RFC 2765). ipaddress reads it as unrelated IPv6 space (ipv4_mapped is None),
so a declared security.fake_ip_ranges block could not excuse it — every fetch
on such a host stayed blocked (pre-flight and connect-time) — and the
cloud-metadata floor missed its translated form, leaving it dialable when
allow_private_urls is on.
_embedded_ipv4() extracts the wrapped IPv4 for both wrapper forms; the three
classification sites classify by it.
Fixes#117100.
A hand-written Linux daemon unit (systemd user unit / XDG autostart entry
running `cua-driver serve`) can be dead for days — crash loop, stopped, never
started — while `hermes computer-use doctor` and `status` report a healthy
binary: the runtime contract only checks the binary (`manifest`), and nothing
ever connected to the daemon socket (#114748).
- tools/computer_use/cua_backend.py::cua_daemon_listening — socket-level
liveness via `cua-driver status [--socket PATH]` (rc 0 = a daemon answered,
"not running" = dead, anything else = unknown). Never raises.
- tools/computer_use/doctor.py::cua_daemon_units — one scan of the units that
run cua-driver (kind, unit, exec target, runs `serve`, `--socket` path with
`%h` expanded); the pruned-Exec guard now filters that list instead of
re-scanning.
- doctor.py::_apply_daemon_liveness_guard — per `serve` unit: `pass` when its
socket answers, `fail` (degrading `ok`) when not, with the hint that a driver
reinstall does not start the daemon. Unconfigured daemons are never probed:
on Linux the MCP runtime needs none, so a silent default socket is normal.
- `hermes computer-use status` prints the dead-daemon line and exits 1.
- docs: the Linux daemon-unit paragraph under the doctor section.
Not changed: _maybe_repair_runtime_contract. The repair is gated on the
binary-level contract only; a dead daemon never enters it, so a reinstall was
never triggered by the daemon (the `.release_installed/<version>` marker the
reporter saw is written by cua-driver itself on any first run of the binary —
observed via strace of `cua-driver status` under a fresh HOME).
Supersedes #114928 (@Finn763): same finding, ~600-line implementation with a
new module and a repair-gate rewrite; this is the ~90-line version on the
existing doctor seams.
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The live-work docks retire a finished background process ~60 s after it
ends; the registry only knew started_at, so the age of an exit was not
observable. _move_to_finished is the single choke point every exit path
(reader loop, reconcile, kill) passes through, so the stamp lives there.
completion_reason rides along so a killed process can read "killed"
instead of "exit -15".
Local message_agent, the Desktop relay, `hermes peer dm` and `hermes peer run`
all hand a turn to the live Bot Chat owner's mailbox and then need the same
thing: poll the receipt until the owner settles it, the budget lapses, or the
caller wants out. Each lane carried (or lacked) its own copy of that loop,
which is how the class drifted — the relay answered with a receipt sentence
(#115316), the two peer lanes never asked the owner at all (#114959, #115174).
`tools/bot_live_delivery.py::await_delivery` (+ `await_delivery_async` for the
aiohttp handlers, which must not park a worker thread for 300 s) is now the one
loop; `_wait_live_dm`, `bot_relay.deliver`, `_answer_through_live_bot_chat`
and `_execute_run_via_live_owner` call it. `should_stop` carries the peer-run
`/stop` check the runs lane needs.
process(action='wait'/'poll'), the already_exited and kill snapshots still used a
hard-coded 2000-char tail with no output_cut, so the polling fallback bot-mode.md
documents for api_server/one-shot senders delivered a silently truncated reply.
One _completion_output() helper now sizes every delivery route.
skills.sh repos that keep SKILL.md at the repo root (no skill directory, e.g.
orzcls/win-disk-cleaner) are listed by `hermes skills search` but `inspect` and
`install` failed with "Could not find '<identifier>' in any source": discovery only
matched `<name>/SKILL.md` paths and the repo-root scan skipped non-directory entries by
construction, so no identifier form could express "the skill directory IS the repo root".
Install goes through the same resolver, so those hub-listed skills were uninstallable.
- GitHubSource._find_repo_root_skill() returns the `owner/repo/` identifier (empty path
segment = repo root) when the tree holds exactly one SKILL.md and it sits at the root,
so a categorized multi-skill repo never resolves this way.
- SKILL.md, support-file, bundle-name and source_url paths go through _skill_file_path(),
which collapses an empty skill path to the repo root; the bundle is named after the repo
(a root skill has no directory to be named after).
- SkillsShSource._discover_identifier() consults the root fallback after the tree walk.
Switching local.py to the folded _is_provider_env_blocklisted helper dropped the
blocklist name from its import list, which removed the re-export that
tests/providers/test_auth_registry_import_order.py (and any out-of-tree caller)
reaches through tools.environments.local. Re-import it alongside the helper so
the module keeps exposing the set; behaviour is unchanged.
The provider-credential blocklist and every adjacent env-name check ran
case-sensitive membership tests while the Windows environment block
resolves names case-insensitively. A skill or terminal.env_passthrough
entry registering openai_api_key was accepted, then resolved to the real
OPENAI_API_KEY by os.getenv() and forwarded into SSH/Docker exec envs,
the same GHSA-rhgp-j443-p4rf tunnel the blocklist closes. The same gap
let variant-cased credentials survive the source-side strip
(_filter_secret_env, _scrub_credentials), evade the docker_forward_env
and docker_extra_args egress collision guards, and ride
strip_launch_profile_env residue into a routed sibling profile's child.
Add _is_provider_env_blocklisted (exact + folded membership) and apply
it at every layer: registration refusal, both scrub paths (Tier-1 and
plugin strip sets fold too; inherit_credentials still inherits), the
remote exec-env builder, and the docker egress collision checks. Fold
_matches_terminal_first_party_prefix symmetrically so a lowercase-stored
buzz_private_key keeps its terminal carve-out, and fold the launch-
residue pop (selection folds too so a lowercase path in .env stays
global). _HERMES_FORCE_* opt-in transport and docker_env literal
container names stay case-sensitive on purpose.
An MCP server task inherits the dashboard OAuth flow for the life of the task
(the ContextVar sits set inside the transport's auth flow), so once that handle
had ended, every retry/revival re-entered `publish_authorization_url` and got
`RuntimeError: OAuth flow already ended`. That error is transient to the connect
ladder, so the server burned all initial-connect attempts and parked
("failed initial connection after 3 attempts") -- and parked forever: no number
of reconnects could get past the same stale handle.
Re-mint the flow in place for the next authorization attempt instead of raising:
clear the ended attempt's URL, state and (spent) authorization code so the new
URL's state is the one the callback is validated against, and the SDK is never
handed a code the provider already burned. Everything stays under the flow lock,
so concurrent/re-entrant publishes leave exactly one live (url, state) pair.
A user cancellation stays terminal: `mark_error(..., cancelled=True)` (both the
dashboard DELETE route and the gateway cancel) is not re-minted by a retrying
worker -- a cancel must keep freeing the per-server slot promptly.
#116727 recognised Deno's bare `eval` subcommand only as args[0]. Deno's CLI
is `deno [OPTIONS] [COMMAND]`, so any global option before the subcommand —
`deno -q eval "..."`, `deno --quiet eval`, `deno -L debug eval`,
`deno --log-level=debug eval` — made the scan stop at the positional `eval`
and the inline script ran without approval, while `node --no-warnings -e`
was still caught. Verified against deno 2.9.6: each of those forms executes
the script; `deno -- eval` opens the REPL and stays unflagged.
Treat the FIRST positional token after the option scan as the exec marker
(the existing loop already skips options and their values), and register
`-L/--log-level` as Deno's one value-taking global so `-L debug eval` is not
mistaken for `deno -L debug` + a data token. `deno run eval.ts` stays data.
Follow-up to #116727 (independent review finding).
A finished (non-timeout) agent-browser failure at the backend level — the CLI
exiting 101 against a stale session daemon, or empty/non-JSON output from a dead
one — returned normally as {"success": False}, so the cached local session record
was never marked suspect and every later browser command re-ran against the same
dead daemon until the process was recycled by hand (#115184).
- _interpret_browser_command_output carries `returncode` on its three
protocol-level failure dicts (parsed page-level errors never carry one)
- _is_recoverable_local_backend_failure classifies those on plain local Chromium
sessions only (cloud/CDP/real-profile/Lightpanda own their recovery)
- _recycle_local_session is the local half of _handle_browser_command_timeout,
extracted so the timeout path and the new path share the alive→suspect /
dead→tree-kill+evict split
- _run_browser_command retries once on the replacement session; `close` is
exempt (a dead daemon is already closed and cleanup must not spawn a session
just to close it)
- argv construction + spawn moved into _dispatch_browser_command so the retry
loop stays a loop over one call
Based on #115206 by @liuhao1024 (returncode on the failure dict, the
recoverability predicate, the one-retry loop); reshaped so the recycle helper is
shared with the timeout path instead of calling the timeout handler.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
`hermes mcp login`, the dashboard re-auth (web_server_mcp.py) and the Desktop
re-auth (tui_gateway/mcp_oauth_sessions.py) each bounded the login probe at
`max(connect_timeout, 315)`. The 315 was the default 300 s callback window
plus headroom, frozen: a user who set `oauth: {timeout: 3600}` still had the
probe cancelled at 315 s. That expiry was a bare `asyncio.TimeoutError`, whose
`str()` is '', so the CLI printed a blank `✗ Authentication failed:` line
(#116278, reporter's steps 5).
`tools/mcp_oauth.py::login_connect_timeout(config)` computes the bound once —
`max(connect_timeout, oauth.timeout + 15)` — and all three call sites use it.
`_probe_single_server` re-raises its wait_for expiry as a TimeoutError naming
the server, the elapsed bound and both governing knobs, so every caller's
`humanized or exc` renders a reason.
Slimmer redo of the blank-line half of #114527 (@liuhao1024): that PR races
the connect against `asyncio.wait` so an inner exception can win; with the
window now sized from oauth.timeout the callback waiter's own
`OAuth callback timed out` message wins on its own, and the described
TimeoutError covers the remaining case in one place.
Completes the timeout atom of #116278 (closed by #116658); refs #103633
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
`MCP: registered N tool(s) from M server(s) (2 failed)` left the failing
identity diagnosable only by elimination from the per-server `registered`
lines. The per-server WARNING fires only on the immediate-failure path; a
candidate skipped for its retry cooldown (a failure from an earlier pass) is
counted as failed with no line of its own at all.
`_connected_summary` now returns `(name, reason)` pairs and `_log_summary`
prints them inline: `(2 failed: github (Connection closed); notion (HTTP 401
...))`. The reason is the recorded `_server_connect_errors` entry (already
credential-scrubbed by `_format_connect_error`); a candidate never attempted
this pass reads `not attempted (in retry cooldown)`. One line, no second
WARNING per server on the path that already warns.
Slimmer redo of #114794 (@liuhao1024, earliest) and #114872 (@Finn763): both
name the failures via an extra WARNING per server; inline on the summary keeps
the immediate-failure path at its current two WARNINGs and still covers the
cooldown case. #114872's extra `_sanitize_error` pass is redundant with the
recorder.
Fixes#114746
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
SIGKILL/taskkill are asynchronous; poll()/isalive() right after kill() still
say alive, so the exact #115490 scenario (a SIGTERM-ignoring child that needed
escalation) was reported as 'Kill incomplete' on every run (10/10 in a live
probe) and the session stayed in _running until the reader thread finalised it
as a plain exit. Re-probe survivors for up to 1s before deciding.
Post-kill, check the session tree (Popen poll, PTY aliveness,
PID-scope identity plus descendant scan). Survivors keep the session
running instead of persisting a false killed receipt and pruning it.
Fixes#115490.