- test_code_kernel::TestModelFacingReset: reset stays in the schema under a stale
kernel_mode and a model-shaped reset call really discards kernel state (#96787).
- test_computer_use::TestStartupTimeoutPhaseDetail: cua-driver ready timeout names
the wedged phase + doctor hint, and the session stays un-started (#57025/#69372).
- test_file_operations::TestEscapeNativeToolArg (windows_only, was FAKE_OS): rg
search and the node --check linter get native C:/ paths, not /c/ (#84303).
- test_file_tools::TestSSHConfigWriteGate (2): a ~/.ssh/config write goes through
the real approval gate, BLOCKED with nobody present / in -q mode, nothing written (#93201).
- test_lazy_deps: a newer compatible plugin SDK is satisfied, never re-pinned down
(#98407, #86992); windows_only real Matrix probe reports unsupported.
- test_mcp_empty_error_message: empty/whitespace exception messages still yield a
diagnostic (#19417).
- test_mcp_stdio_children_dead: the fast-fail probe never calls the watcher on the
plain path, no un-awaited coroutine (#96044).
- test_subprocess_utf8_encoding: real _op_whoami survives invalid UTF-8 child output (#53428/#55339).
- test_tts_output_timestamp: same-second TTS calls get distinct output files (#43911).
- test_windows_agent_loop_papercuts::TestAutocompleteDevicePaths: a relpath
ValueError on a device path is skipped, not raised (#42016).
- test_windows_native_support::TestCronSchedulerBashResolution (2): cron .sh uses
bash from PATH; missing bash is an actionable error, not WinError 2.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
The idle sweep lives inside _acquire_kernel, so it only fires when the next
kernel request arrives. A host that stays alive but stops executing anything
(a pids-exhausted container fail-closing every tool call) never acquires
again: idle kernels, their runners and thread pools survive indefinitely,
which is how four orphaned runners helped exhaust pids_limit=256 for hours
(#117169). A low-frequency daemon reaper — started on the first kernel
spawn, using the acquire-path criteria verbatim — retires idle kernels on
its own schedule, independent of tool traffic.
The same pass sweeps hermes_kernel_* staging dirs untouched for over a
week: a live host rmtrees each dir within one idle timeout of the kernel's
last use, so a week-old dir belongs to a host that died without cleanup
(SIGKILL / container restart). Younger dirs are left alone because a
concurrently running host's live kernel may own one; rmtree rejects
symlinks rather than following them.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
A delegated child's execute_code kernel was keyed correctly
(<owner>:🧒:<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.
A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
_callback_api() now yields (getter, setter) pairs for every per-thread prompt
(approval, sudo, vault unlock); the kernel cell captured and restored the old
fixed 4-tuple. Iterate the table so a cell carries every callback and a future
addition needs no change here. Test recorder unpacks the new shape.
Also: perfectionist import order in ui-tui interfaces.ts (CI lint).
Prepare dependency generations before selecting them. Keep shipped tool
bytes separate from writable additions, and store facts beside their entries.
Validate proposed plugin sets before config publication. Restore the previous
config if the facts write fails.
Consolidate duplicate updater, backup, setup, and voice helpers. Repair
launcher selection, dependency consumers, download ownership, update feeds,
and native Windows process and file handling.
Verification: 206 changed/prior-failing Python files reported 4630 passed,
one failed, and 330 skipped. Fix the remaining Hindsight fixture boundary.
The final targeted rerun reported 234 passed and two skipped. The store
review regression batch reported 83 passed and one skipped. Desktop
TypeScript checks, 56 selected Electron tests, 24 release tests, and the
removed-import/compatibility guards passed.
This is an integration checkpoint, not full audit acceptance. The complete
Python suite has not run on this fixed tree. Crash-atomic plugin publication,
generation cleanup, receipt correlation, and packaged lifecycle acceptance
remain open in docs/pm-audit-status.md.
pytest-xdist is gone: run_tests.sh no longer dispatches on host, the
Windows xdist (--dist loadfile) arm is deleted, and pytest-xdist is
dropped from the dev extra (uv lock removes it + execnet). The
per-file subprocess runner (run_tests_parallel.py) is THE runner on
every host — the shape the CI linux lane already used.
- tests-os.yml: the macOS/Windows lanes now run scripts/run_tests.sh
like every other lane, passing the marker-narrowed file list via
--files and -m after --. The runner natively tolerates per-file
empty collections (exit 5, platform-gated file) and fails the run
when NOTHING collected — replacing the hand-rolled exit-5 branch.
- tests/gateway/conftest.py: the hasattr(config, 'workerinput')
controller-guard was xdist-only dead code; the file-locked cache
already handled the per-file model. Removed.
- tests/tools/test_browser_supervisor.py: the port came from xdist's
worker_id (fixed 9225 under per-file isolation — concurrent files
collide). Binds an ephemeral port instead (bind 0, read back).
- ~35 comment sites named xdist as the isolation mechanism; they now
describe the shared-process hazard they actually guard against
(bare pytest runs, same-file ordering) without naming a runner that
no longer exists.
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.
Widens the Windows parent-death fix to the class: the kernel inherits
the read end of a pipe whose only write end lives in the host, so host
death by any means (SIGKILL, OOM, crash) is EOF and the kernel exits.
Stdin EOF alone only reaches the runner between cells, so a kernel
SIGKILLed mid-cell used to outlive its host indefinitely. The
integration test now runs on every platform and is trimmed to the
contract (kill host mid-cell, kernel gone), not the handle plumbing.
Session kernels: a kernel mid-spawn (proc=None) read as dead, so every
concurrent cell for one owner replaced the registry entry and the
winner's process leaked outside the registry (110 live kernels, 1.1 GB,
330 threads under one 4-capped process). Reap/evict also tore down
kernels with cells attached, rmtree-ing the staging dir under the
spawner. Kernels now track attached cells: only settled kernels are
reaped/evicted, a kernel dropped while busy is torn down by its last
cell, and in-cell registry pops never remove a replacement.
LSP: a directory holding __init__.py is a package, not a project root.
hermes_cli/setup.py matched the python marker list and gave every
worktree a second pyright rooted at hermes_cli/ (70 of 105 reaped
clients in one session, ~40 servers / 8.7 GB live).
* refactor(execute_code): integrate kernel persistence into the core description (712 -> 654 tok/call, -8%)
* fix(execute_code): honest interpreter note — Hermes's own python is the common case; project venv only when VIRTUAL_ENV/CONDA_PREFIX is active
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)
* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel
* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen
* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session)
execute_code spawns a fresh Python process per call, so every multi-step
data task re-loads its inputs: a CSV parsed in call one is gone by call
two, and scripts route state through temp files to survive. Hermes
already rewards programmatic tool calling (execute_code-only turns
refund the iteration budget), which makes the missing half — state that
survives between calls — the bottleneck.
Add opt-in `code_execution.kernel_mode: session`: one persistent kernel
per (task, mode, interpreter, cwd, tool-set). Variables, imports, and
loaded data persist across calls; `reset=true` discards state on demand.
The default `per-call` keeps today's behavior byte-for-byte.
Safety posture is unchanged by design: the child env comes from the same
builder as the per-call path (extracted, not duplicated, so the secret
scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same
`_rpc_server_loop` with the same token and a per-cell tool budget, and
output passes the same ANSI strip + secret redaction. A timed-out or
interrupted cell kills the whole kernel tree and the next call respawns
— a wedged kernel can never hang the agent. The kernel env is frozen at
spawn; the schema and config comment say so.
Wire protocol: NDJSON requests on the kernel's stdin; responses framed
on stdout behind a per-kernel random sentinel, with unframed bytes
(fd-level output from user-spawned subprocesses) attributed to the
serialized current cell. The generated RPC client reconnects once when
HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC
server's 300s idle window between cells.
Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in
tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel,
timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough,
schema surface, mode fallback) plus the existing
test_code_execution.py / test_code_execution_modes.py suites (81 passed).
* fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority
Addresses the blocking review on the session-kernel design: two
authority/lifecycle boundaries were wrong.
1. Ownership and bounded lifetime. The kernel key's first component is
now the conversation's approval session key (_resolve_owner), not the
per-turn task id run_agent mints per top-level invocation — so state
genuinely survives across user turns of one conversation, and delegated
subagent sessions isolate naturally under their own keys (the task id
remains only the last-resort owner for embeds/tests with no session
context). Lifetime is bounded on four edges: kernels are disposed at the
same session boundary that clears the owner's approval/yolo state
(tools.approval.clear_session -> shutdown_kernels_for_owner), reaped
after code_execution.kernel_idle_timeout seconds idle (default 1800,
swept on every entry), capped process-wide at
code_execution.max_session_kernels live children (default 4, LRU
evicted), and still torn down by reset/death/atexit as before. The
ownership + disposal + idle-reap + cap shape deliberately carries
forward the lifecycle invariants of the earlier session-persistent
implementation in #88637 by @z80dev.
2. Per-cell RPC authority. The serving thread no longer freezes the
spawning cell's context/callbacks for the kernel's life. Each cell
installs a CellAuthority — captured on the calling thread exactly as
propagate_context_to_thread would for a per-call RPC thread — before its
request is written, and retires it on every settle path; _rpc_server_loop
gains a dispatch hook the kernel uses to route each tool call through
the CURRENT cell's context, callbacks, and task id. A call arriving with
no active cell is refused. Interpreter state persists; RPC authority
does not.
Composition with the per-script static guard (see the config note): a
persistent namespace lets cell N+1 invoke objects cell N created, which
a single-cell static scan cannot see — the runtime RPC boundary
(allow-list by name, per-cell budget, per-cell authority) is the
operative cross-cell enforcement in this mode, and the adversarial
alias test pins exactly that.
Tests (9 new): state survives across turns of one conversation;
sessions isolate; clear_session disposes the owner's kernels (and the
next turn starts fresh); the live-kernel cap LRU-evicts with evicted
children proven dead; idle kernels are reaped; a later cell's RPC runs
under that cell's approval callback; a cross-cell alias dispatches under
the CURRENT cell's authority; a settled cell's authority refuses
dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81
code-execution tests, ruff clean. The 7 test-order failures in the
tools/-k-approval selection reproduce identically on the clean branch
base (pre-existing pollution, not this change).
* fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions)
---------
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>