3c4e84c1663107a0305247c59b2cfbea4c2d18b2
1 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b39d76d902 |
feat(tools): session-persistent kernels for execute_code (kernel_mode: session) (#94647)
* feat(tools): session-persistent kernels for execute_code (kernel_mode: session) execute_code spawns a fresh Python process per call, so every multi-step data task re-loads its inputs: a CSV parsed in call one is gone by call two, and scripts route state through temp files to survive. Hermes already rewards programmatic tool calling (execute_code-only turns refund the iteration budget), which makes the missing half — state that survives between calls — the bottleneck. Add opt-in `code_execution.kernel_mode: session`: one persistent kernel per (task, mode, interpreter, cwd, tool-set). Variables, imports, and loaded data persist across calls; `reset=true` discards state on demand. The default `per-call` keeps today's behavior byte-for-byte. Safety posture is unchanged by design: the child env comes from the same builder as the per-call path (extracted, not duplicated, so the secret scrubbing / PYTHONPATH hygiene cannot drift), the RPC server is the same `_rpc_server_loop` with the same token and a per-cell tool budget, and output passes the same ANSI strip + secret redaction. A timed-out or interrupted cell kills the whole kernel tree and the next call respawns — a wedged kernel can never hang the agent. The kernel env is frozen at spawn; the schema and config comment say so. Wire protocol: NDJSON requests on the kernel's stdin; responses framed on stdout behind a per-kernel random sentinel, with unframed bytes (fd-level output from user-spawned subprocesses) attributed to the serialized current cell. The generated RPC client reconnects once when HERMES_RPC_PERSISTENT=1, because a kernel legitimately outlives the RPC server's 300s idle window between cells. Tested on macOS 15 (Apple Silicon), Python 3.11: 13 new tests in tests/tools/test_code_kernel.py (persistence, reset, error-keeps-kernel, timeout-kills-kernel, sys.exit ends kernel, subprocess fd passthrough, schema surface, mode fallback) plus the existing test_code_execution.py / test_code_execution_modes.py suites (81 passed). * fix(tools): session kernels get a stable owner, bounded lifetime, and per-cell RPC authority Addresses the blocking review on the session-kernel design: two authority/lifecycle boundaries were wrong. 1. Ownership and bounded lifetime. The kernel key's first component is now the conversation's approval session key (_resolve_owner), not the per-turn task id run_agent mints per top-level invocation — so state genuinely survives across user turns of one conversation, and delegated subagent sessions isolate naturally under their own keys (the task id remains only the last-resort owner for embeds/tests with no session context). Lifetime is bounded on four edges: kernels are disposed at the same session boundary that clears the owner's approval/yolo state (tools.approval.clear_session -> shutdown_kernels_for_owner), reaped after code_execution.kernel_idle_timeout seconds idle (default 1800, swept on every entry), capped process-wide at code_execution.max_session_kernels live children (default 4, LRU evicted), and still torn down by reset/death/atexit as before. The ownership + disposal + idle-reap + cap shape deliberately carries forward the lifecycle invariants of the earlier session-persistent implementation in #88637 by @z80dev. 2. Per-cell RPC authority. The serving thread no longer freezes the spawning cell's context/callbacks for the kernel's life. Each cell installs a CellAuthority — captured on the calling thread exactly as propagate_context_to_thread would for a per-call RPC thread — before its request is written, and retires it on every settle path; _rpc_server_loop gains a dispatch hook the kernel uses to route each tool call through the CURRENT cell's context, callbacks, and task id. A call arriving with no active cell is refused. Interpreter state persists; RPC authority does not. Composition with the per-script static guard (see the config note): a persistent namespace lets cell N+1 invoke objects cell N created, which a single-cell static scan cannot see — the runtime RPC boundary (allow-list by name, per-cell budget, per-cell authority) is the operative cross-cell enforcement in this mode, and the adversarial alias test pins exactly that. Tests (9 new): state survives across turns of one conversation; sessions isolate; clear_session disposes the owner's kernels (and the next turn starts fresh); the live-kernel cap LRU-evicts with evicted children proven dead; idle kernels are reaped; a later cell's RPC runs under that cell's approval callback; a cross-cell alias dispatches under the CURRENT cell's authority; a settled cell's authority refuses dispatch; each cell installs a fresh authority. 22/22 kernel tests, 81 code-execution tests, ruff clean. The 7 test-order failures in the tools/-k-approval selection reproduce identically on the clean branch base (pre-existing pollution, not this change). * fix(code-kernel): delegated children get their own kernels — child contexts inherit the parent approval key, so qualify the owner with the delegation session id (live-verified leak, both directions) --------- Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com> |