_mcp_registry_scope() became profile-keyed for served profiles with the
multiplex flag off, but check_fn_cache_scope() still returned None in that
mode, so the process-wide availability cache stayed keyed (fn, None) across
profiles. A served profile whose mcp__x__* check_fn now correctly resolves to
its own (absent) connection cached False for the TTL window and the launch
profile that owns the live connection lost its tools for that window.
Both sites now call one helper, agent.secret_scope.serves_routed_profile()
(multiplex on, or a HERMES_HOME override naming a home other than the
process home), so the registry scope and the cache key can no longer drift.
Review finding: served profile B's check_fn verdict shadowed the launch profile's live mcp tools via the unscoped check_fn cache.
A dashboard/desktop backend (and the per-profile cron ticker) serves sessions of
several profiles through the HERMES_HOME contextvar override while
gateway.multiplex_profiles stays off. _mcp_registry_scope() keyed every MCP
connection by the bare server name in that mode, so the first profile to
discover `zernio` owned the only connection and every later served profile —
including one whose config carries a different Authorization header — called
the server through it and got the other account's data back (#111151).
The registry scope now follows the served home: a routed profile (an override
naming a home other than the process home) gets the same per-profile overlay
the multiplexer uses, so a same-named server with other credentials is a
separate connection, discovery for profile B is a connect candidate instead of
"already connected", and status/tool views stay per profile. Single-profile
processes (no override) keep bare keys, byte-identical to before.
Fixes#111151
Credit: #111158 by @KoNit-K located the inert flag on hermes_cli surfaces; its
fix (activating fail-closed multiplex secret scoping from config.yaml on the
dashboard) is not taken — the connection-key seam, not the secret-scope mode,
is what leaks the connection, and flipping the process-wide mode from the
dashboard would change credential resolution for every code path in it.
_parallel_safe_servers was keyed by the raw server name while _servers
moved to (scope, name) connection keys (ceaf622c6d). With profile A's `x`
serial and profile B's same-named `x` opted into parallel calls, B's
discovery pass flipped A's tool to parallel-safe and the batch planner put
two A calls in one parallel segment against a server that never opted in
(review of #108352, finding B).
The opt-in is now recorded under the discovering profile's own key and
is_mcp_tool_parallel_safe() looks it up under the calling profile's key,
so one profile's policy never reaches another's same-named server.
Single-profile processes keep the bare-name key, byte for byte.
Under gateway.multiplex_profiles a profile whose mcp_servers entry has the
same route and credentials as another profile's live connection adopts that
connection instead of opening its own. _same_server_route() compares only
the connection identity, so a `trust: untrusted` profile adopted a
`trust: full` profile's connection; _trust_gate_check() then resolved the
OWNER's connection key, read `full`, and let the untrusted profile run
write-capable tools without the approval prompt its config demands
(review of #108352, finding A; regression from ceaf622c6d, where the
name-keyed ledger let the last registrant's trust win instead).
`trust` is the consuming profile's policy, not a property of the
connection: _server_trust_levels is now keyed by the calling profile's own
key (recorded at its own registration and at adoption, dropped when its
overlay is removed), while readOnlyHint stays under the connection key
because it describes the server's tools. Sharing the connection is still
allowed — only the gate is per profile.
Relocated onto the decomposed module layout and hardened:
- Trigger covers the rejection CLASS, not just literal 400: SSE-only
servers' load balancers answer the chunked Streamable HTTP initialize
POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
rejections as an opaque -32603 'Server returned an error response'
(error class per #104363 by @RohithPariki). Timeouts and 5xx never
trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
clears _ready before re-entering the transport, so the original guard
also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
connects + latches; no fallback on reconnect/timeout/5xx; both-fail
error is actionable.
The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.
Fixes#53676
Application errors (isError payloads) keep counting as breaker strikes:
that is #10447's point (a server answering errors made the model hammer
it 8x in 10s) and #109180 just reasserted it. What #11113 actually hit is
the open-breaker MESSAGE: after three rejected fetches the model was told
the server was "unreachable" and went to the user instead of fixing its
URL. Track whether the streak was all application errors and word the
pause accordingly; one transport strike restores the unreachable text.
Under gateway.multiplex_profiles every connection ledger in tools/mcp_tool.py
(_servers, _server_scope_keys/_server_tool_scopes, connecting/error/cooldown
maps, the circuit breaker, lazy schema-cache configs, trust metadata) was keyed
by the bare server NAME. The common per-tenant layout — each profile names its
server `github`/`notion` with its own token — gave only the first profile a
connection: the second profile's register_mcp_servers saw the name as "already
connected", refused to adopt it (different credentials, 4ddbcbd35e), and left
the profile silently tool-less with a healthy-looking `configured` status
(#106005 Bug 1/2, #91654). Siblings of the same bug: profile A's failing `x`
put profile B's healthy `x` into A's 10-minute connect cooldown and A's open
circuit breaker short-circuited B's calls; toolsets._resolve_toolset_memo was
not scope-keyed, so B resolved A's `mcp-<server>` tool names.
Keys are now the connection key from the new tools/mcp_tool_scope.py: the bare
name outside a multiplexer (single-profile processes are unchanged) and
(owner_scope, name) under one. Call-time lookups (_resolve_server_key) prefer
the calling scope's own connection, then a shared connection it adopted, so
identical-route profiles still share one subprocess. _select_new_servers,
the cooldown/breaker/trust maps, lazy registration and get_mcp_status all read
and write through the composite key; teardown resolves a task's key by
identity (the MCP loop has no profile context). The toolset memo key includes
registry.current_scope_key().
An owner's scoped /reload-mcp tore down its connection and, with it, every
adopting profile's tool overlay; nothing re-ran the adopters' discovery until
they reloaded. shutdown_mcp_servers(scope=) now records the orphaned adopters
and register_mcp_servers re-registers them under their own home + secret scope
once the owner's rediscovery pass completes.
Docs: multi-profile-gateways.md states the per-profile connection rule.
Fixes#106005Fixes#91654
Co-authored-by: Bergmann89 <info@bergmann89.de>
Co-authored-by: Izzy-Gottz <srulynj@gmail.com>
_paginate_full_list wrapped the paginated list call in try/except TypeError
to detect the mcp 1.x calling convention. The same except also caught
TypeErrors raised INSIDE the modern list call — e.g. a server response
decode failure — and retried with the legacy cursor= keyword, replacing the
real error with a misleading 'unexpected keyword argument cursor' and
making genuine MCP pagination failures undiagnosable.
Probe list_method's signature instead (_list_method_accepts_params): the
legacy cursor= fallback fires only when the method genuinely doesn't accept
the mcp 2.0 params= keyword (or takes **kwargs), so a TypeError from inside
the list call propagates to the caller. Regression tests: the decode
TypeError surfaces and the legacy retry doesn't run; a genuinely 1.x-shaped
method keeps using the cursor fallback.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Review follow-up, and a real platform bug. On Windows npm's `.bin` holds three
siblings per binary — the extensionless sh script, `<name>.cmd` and
`<name>.ps1`. Spawning the sh one from a Windows process fails, and
`os.access(X_OK)` there is effectively an existence check, so the previous
candidate test could not tell them apart: it would have picked a file that
cannot run and broken server startup that works today via npx.
Select by extension instead (`.cmd`, then `.exe`), the same precedence
hermes_constants._candidate_node_command_names already uses for npm/npx/node,
and fall back to npx when no launcher is present.
The platform branch moved into `_npx_bin_candidates(..., windows=...)` so it is
testable by injection. Patching `os.name` instead took pytest's own traceback
formatting down with an INTERNALERROR — a test that cannot report its own
failure is worse than no test.
Also from review: an `npx pkg -y` shape (flag AFTER the spec) now falls back to
npx, since those args are forwarded verbatim and would hand the server a flag
npx would have consumed. And the structural OSV-ordering guard reports a rename
explicitly instead of raising a bare ValueError that reads like a broken test.
`npx` resolves a package and then FORKS: it stays alive as the real MCP
server's parent for the whole process lifetime while doing no work. Measured
on a 4-agent host, that is ~48 MB of private memory per stdio MCP server —
and it buys nothing here, because Hermes already wraps the child in its own
parent-death watchdog, so npx's supervision is a second parent nobody reads.
The process tree for one server is watchdog -> `npm exec <pkg>` -> node.
When the package is already in npx's cache, spawn its binary directly and
drop the middle process. A cache miss changes nothing: the command stays
`npx`, so a cold machine still installs on first run.
The swap deliberately happens AFTER the OSV malware preflight.
`_infer_ecosystem` keys off the command basename being `npx`/`uvx`/`pipx`, so
rewriting first makes `check_package_for_malware` return None and silently
turns the gate into a no-op. Two tests pin that ordering — one behavioural,
one structural — because a future edit that moved the swap earlier would
disable the malware check without failing any test of either piece alone.
This is also why doing it in code beats the config workaround of pointing
`command` straight at the cached path: that loses the preflight.
Conservative by construction — falls back to npx for a version-pinned spec
(npx owns that resolution), an ambiguous or absent `bin` map, a missing or
non-executable binary, an unreadable cache entry, or no cache at all.
Measured on a Mac mini running four agent gateways plus a dashboard, ten
stdio MCP servers total: resident `npm exec` parents 5 -> 0, tracked
footprint 1531 MB -> 1324 MB across 30 -> 25 processes, free memory
464 MB -> 697 MB on a host that was actively swapping. Every MCP server kept
working; tool calls verified against a live Linear server afterwards.
content and structuredContent are now alternatives — never both forwarded
to the model. Spec-following servers render their data into content (the
verbatim dual-emit SHOULD or a faithful reorganisation), so forwarding
both sent the same information twice per tool call. structuredContent
fills in only when content blocks render effectively empty (whitespace-only
text counts as empty), preserving structuredContent-only servers. _meta
passthrough is unchanged.
Also surfaces unsupported content-block drops to the MODEL as
'[MCP content dropped: unsupported block (...)]' notices carrying
type/mime/uri/size handles (kimi-code#3227) instead of a log-only warning;
drop notices do not count as usable content for the arbitration.
Live E2E: real stdio MCP server (mcp 2.x structured_output tool) through
register_mcp_servers + _make_tool_handler — origin/main forwarded the
payload twice, branch forwards content only; text-only path unchanged.
With servers configured but none selected by -t/--toolsets, discovery still
paid the ~260ms mcp SDK import before discovering it had nothing to spawn.
Filter first; the test now asserts the SDK probe is not reached.
Phase-2c finding on #101598: the empty-set release closed the pipe and
dropped the Popen without reaping it, so an idle gateway that once
connected a stdio server held one zombie until the next Popen in the
process. wait(timeout=5) after close -- the supervisor exits on EOF
immediately with nothing registered (E2E: ps shows no entry at all).
Also: the replay test asserted against the last line only, which could
not catch the regression it names; assert against the full stream and a
post-unregister replay. Module docstring now states the stdlib-only /
no tools/ import constraint and why _reap's sweep is a deliberate copy.
Addresses @andrexibiza's Blocker 1 on #93517 (reproduced): after a
broken-pipe write dropped the supervisor while groups were still
registered, the no-spawn fast path was keyed on the incoming verb
(`unregister` + no proc => return), so a clean teardown of one server
left the survivors recorded in _supervised_pgids but unsupervised until
the next register. Key the fast path on the supervised set being empty
instead; any later call respawns and replays the survivors.
Regression test: two live groups, write fails, unregister one ->
replacement receives `register` for the survivors. Mutation-checked
against the old verb-keyed guard.
Follow-up to the salvaged #93517. Two gaps found in review:
- After the last unregister the supervisor stayed resident for the life of
the process (~15 MB + a pipe) in any gateway/cron that ever connected a
stdio server; main's per-server watchdog exited with its server. Close
our write end when the supervised set empties: EOF with nothing
registered makes the supervisor exit without reaping, and the next
register already respawns and replays.
- A child that raced and exited before os.getpgid dropped its group from
coverage entirely. The SDK spawns stdio servers as session leaders
(pgid == pid), so fall back to the pid instead; the prune forgets the
group once nothing in it is alive.
Tests: fake-supervisor release/re-spawn sequence, and a real-process EOF
release (supervisor exits 0, unregistered child untouched). Mutation
checked: disabling the release branch fails both.
Review raised a real gap: we reap by pgid, so a registration is only as
meaningful as the group's identity. A group we deliberately keep registered --
an orphan teardown failed to kill, such as the `node` mcp-remote leaves behind
-- can later exit on its own, after which the kernel may hand that pgid to an
unrelated process owned by the same user. An ungraceful Hermes death while the
registration is stale would then signal a stranger. `_is_safe_target` cannot
catch it, because the value is stale rather than invalid.
Prune registrations whose group has no members left on every registration
change, and tell the supervisor to forget them. Signal 0 is a pure existence
question -- it cannot terminate anything -- so this is cheap and safe to run on
the hot path. An ambiguous answer (EPERM: exists but not ours) keeps the
registration, since dropping real coverage is the more expensive mistake.
This narrows the window rather than closing it: a group can still die and its
pgid be recycled between two probes. Closing it completely means proving group
identity at reap time, e.g. stamping children with a boot-unique env marker and
checking a member still carries it. That was judged not worth putting a `ps`
parse into the one process whose job is to stay simple enough to always work,
so the residual is now documented in the module docstring instead, along with
the note that Hermes's existing killpg-based orphan cleanup already carries the
same exposure (upstream #88350).
Also records why supervisor recovery is deliberately two-step: a failed write
drops the handle, and the next call rebuilds coverage from `_supervised_pgids`,
which -- not the pipe -- is the record of what needs reaping.
The protocol tests register synthetic pgids that were never real process
groups, so they now state that precondition through an `all_groups_alive`
fixture instead of depending on pid-space luck.
29 tests pass. Verified the change adds no failures: same selection, my changes
stashed vs applied, 42 failed either way (pre-existing in a locally rebuilt
venv) with passed going 617 -> 619 for the two new tests. Footgun linter clean.
(cherry picked from commit 3a7d219620a8cb57a7fe2c2346ce3f055320e743)
Every stdio MCP server was wrapped in its own CPython watchdog that polled
getppid() every 2s to notice an ungraceful Hermes exit (kill -9, OOM, crash,
force-quit), since macOS has no PR_SET_PDEATHSIG. That is a whole interpreter
per server for a job that does nothing until the moment Hermes dies: 10.1 MB
physical footprint each, measured on macOS/arm64.
Replace the fleet of pollers with a single supervisor per Hermes process
holding the read end of a pipe only Hermes writes to. Death detection becomes
EOF on that pipe -- exact and instant, rather than up to a poll interval late.
Hermes sends `register <pgid>` / `unregister <pgid>` as servers come and go;
on EOF the supervisor killpg's whatever is still registered, which is exactly
the set whose teardown never ran. A clean shutdown unregisters as it goes, so
EOF then finds nothing to kill.
Servers are now spawned unwrapped. The MCP SDK already starts each stdio child
in its own session, so the pgid recorded for killpg is the server's own group
and the existing cleanup paths reach it unchanged. That also deletes the
signal-forwarding layer the wrapper needed: wrapping had put the real server in
a different session from the pgid being tracked, so a graceful killpg would
have hit only the wrapper.
Measured on a 5-gateway host: 10 watchdogs (~98 MB) -> 5 supervisors (~49 MB).
One supervisor costs about what one watchdog did (9.9 vs 10.1 MB), so the win
is (servers_per_process - 1) x ~10 MB, and a process with no stdio servers now
spawns nothing at all.
The supervisor reads length-capped lines rather than iterating the stream: a
writer that never sends a newline would otherwise grow it without bound, which
it must not be vulnerable to when it is the last defense against leaked
servers. Found by feeding it /dev/zero, where it reached 15 GB.
Verified beyond unit coverage: a real stdio MCP server connects and its tools
are discovered on the unwrapped path; with a live Hermes holding a real
connection, kill -9 reaped the server, its grandchild (in the server's group,
the mcp-remote `node` case), and the supervisor exited on its own. The reap
tests were sabotage-checked in both directions -- a no-op reaper fails all
three, while the test pinning that a cleanly unregistered server survives keeps
passing -- and each wiring half fails independently when removed.
(cherry picked from commit a252d4ce7ff1722f687635fdbf0cff79f538c3f1)
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:
* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
agent.tools from live probes with no predecessor to preserve. Persist
the session's resolved tool-name order in a new `sessions.tool_names`
JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
alongside the system prompt and re-pinned on every published refresh
(so /reload-mcp and compaction naturally reset it; /new mints a new
row). On restore-for-existing-session the fresh definitions are folded
onto the saved order via the SAME `_merge_preserving_prefix` helper —
a probe-flipped tool is carried forward from the registry schema, a
deregistered one dropped, new tools appended at the tail.
* /reload-mcp (CLI, gateway, TUI RPC) now also calls
`reprobe_tool_availability()` — drops the check_fn verdict cache and the
get_tool_definitions memo — so a user can consciously pick up a
credential/daemon that appeared mid-session. Docs updated.
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:
* a tool whose `check_fn` merely flapped (headless browser probe, expired
credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.
Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.
`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.
Refs #100336
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.
- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
served profile inside `_profile_runtime_scope`, carried into the
executor via `copy_context()` (same shape as
`_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
caller (e.g. button-confirm callback) did not; shut down / rediscover /
report only that profile's servers; refresh only that profile's cached
agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
`_server_scope_keys` ownership map; leaves the shared MCP loop running
while other profiles' servers are live. Unscoped call keeps the full
historical behavior.
- MCP tools register into the owning profile's registry overlay
(`registry.register(scope=...)`), and `registry.deregister()` gains a
matching `scope=` kwarg. Plugin callers still cannot name another
profile's scope; the plugin-vs-global guard is unchanged for them.
Fixes#95518
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
A stdio server that never answers the optional ping (no -32601, no
response at all) produced a bare TimeoutError that _keepalive_probe
classified as a dead transport, tearing down and respawning a healthy
subprocess on every keepalive tick. On a first ping timeout, confirm with
list_tools before declaring death; if it answers, latch _ping_unsupported
and use list_tools from then on. If both fail, propagate as before.
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.
The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.
Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.
The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.
Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
exhausted, parked, tools deregistered -- no respawn loop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/proc/<pid>/task/<tid>/children is per-thread. stdio_client() spawns the
MCP subprocess from the background loop thread, so the main-thread-only
read returned an empty set on Linux and _stdio_child_pids/_stdio_pids
never tracked the child: the #81995 dead-child fast-fail, the #96452
respawn signal and the killpg shutdown sweep were all no-ops. Union the
children of every task instead.
Follow-up to the salvaged #96044 hunk: drop the 'or callable(...)' arm —
callable(MagicMock) is True, which would have flipped stubbed sessions
into the fast-fail race the surrounding comment explicitly routes to the
plain-await path. inspect.iscoroutinefunction alone reproduces the old
isawaitable(call) split exactly (real async def / AsyncMock -> race,
MagicMock -> plain await) without creating the leaked coroutine.
The fast-fail gate probed the stdio child watcher by CALLING it —
inspect.isawaitable(_watch_children()) — creating a fresh coroutine on
every stdio MCP tool call that was never awaited (RuntimeWarning spam +
gc churn). Inspect the function instead of invoking it.
Salvaged (unique hunk only) from PR #96044; the bundled
_stdio_children_dead polarity fix was already on main via #94339.
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.