The rm of the cell result file runs after the runner has executed the cell.
A transport failure there propagated out of _run_remote_cell, so
_run_attached_cell evicted the kernel and re-raised, and the caller's
per-call fallback then ran the same code a second time.
Catch and debug-log that failure (a leftover cell_res_* is harmless because
seq is monotonic), so the atomic ship is the only remote call that can raise
inside the discard scope and the handler's "runs exactly once" comment holds.
The cell ship now fails closed (RuntimeError). _execute_remote catches it
and falls back to per-call execution, but the kernel stayed registered
with cell_seq bumped. The next call would then reuse a kernel whose
state silently missed this cell, with no state_lost flag. Discard (kill)
it the way the timeout path does, so the next call starts a fresh
kernel. The request never reached the runner (atomic publish), so the
fallback still runs the code exactly once.
Gate follow-ups on the remote lockdown:
- _execute_checked(env, cmd, what, **kw) in code_execution_rpc replaces
the three copy-pasted "execute, raise if returncode != 0" blocks
(per-call setup, kernel dir setup, file ship via
_remote_write(check=True)). env.execute always returns a dict, so the
isinstance/(r or {}) guards go; the error carries command output only,
never the payload.
- _ship_env_file_and_launch_prefix returned a half-built "( ... && "
that both callers had to close; a caller that dropped the ")" or
composed it differently would lose the load-bearing subshell. It now
takes the launch command, builds the shared env map (RPC dir, token,
PYTHONDONTWRITEBYTECODE, routed TZ) itself, and returns the complete
command; the kernel passes only HERMES_KERNEL_DIR/PYTHONPATH, which
drops its duplicated TZ block and lazy hermes_time import.
- _private_dirs_cmd(root, *subdirs): every caller spelled each path
twice for mkdir and chmod.
- _run_remote_cell publishes the cell request with one atomic
_remote_write instead of ship-to-.tmp then a separate unchecked mv,
saving a backend round-trip per cell.
On shared remote backends the execute_code channel created kernel and
sandbox dirs under shared temp at the process umask (775 group-writable
under umask 002), wrote request/result files group-readable, and carried
HERMES_RPC_TOKEN on remote command lines where co-tenant users read argv
via ps for the whole run. A co-tenant could read tool arguments and
results, and on group-writable dirs forge RPC requests dispatched under
the user's approval context.
- All remote dirs are created owner-only (umask 077 + chmod 700, checked
fail-closed) and every Hermes file write is mode 600.
- The token travels in a sourced env file inside a subshell so the vars
never enter the backend's session-snapshot dump, and ships via stdin on
pipe-capable backends so it never enters argv at all.
- The RPC poll loop rejects non-int seq requests before dispatch instead
of replaying them every cycle.
- tool_result_storage gets the same owner-only treatment for archived
tool output.
(cherry picked from commit aef21731d7fb8a4e0a6ada4ff9889264df4a8893)
A delegated child's execute_code kernel was keyed correctly
(<owner>:🧒:<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.
A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
The hermes_tools stub module a remote kernel imports is generated from sandbox_tools once,
at spawn, but the registry key was (owner, env_type, task_env_id) only: a later
execute_code call with a different tool set (skill loaded, toolset toggled) reused the
kernel and got stale stubs. The tool set is now part of the key; a different set gets
its own kernel and the over-cap eviction keeps the newest.
Salvaged from #97265 by @Liuzikaii.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
Same attached-cell guard as the local kernel host (#101861): a remote
kernel mid-cell is never reaped or cap-evicted, so a fan-out never has
its runner killed under a live poll loop.
Local session kernels sweep idle-expired entries and enforce a
process-wide cap (DEFAULT_MAX_SESSION_KERNELS) on every call
(tools/code_kernel.py's _reap_unlocked / _evict_over_cap_unlocked). The
new remote kernel host (#96991) never got the same treatment:
_REMOTE_KERNELS only shrinks lazily when a specific key is revisited and
found dead, so an owner that opens kernels for several distinct
(env_type, task_env_id) combinations (or delegated children) and never
revisits some of them accumulates host-side bookkeeping entries for the
life of the gateway process.
Note this is narrower than the local case: the remote runner already
self-reaps on its own idle timeout, and SSH/Docker connections are
independently bounded by their own transport-level lifecycles (SSH
ControlPersist, Docker's session-scoped idle-timeout in terminal_tool.py)
— so nothing here leaks a live remote connection. What's missing is
purely the host-side dict/cap bookkeeping symmetry with local kernels.
Adds _reap_unlocked/_evict_over_cap_unlocked mirroring the local
implementation, reusing the same max_session_kernels config as an
independent cap on _REMOTE_KERNELS.