The kernel eviction on a failed cell ship and the raise in _execute_checked
had no test teeth: removing either left the suite green. Extend the existing
shared-host lockdown test (no new test functions) so a failed cell ship must
raise and empty the registry, and a failed dir setup must spawn nothing and
ship nothing.
Also reuse the loop's parsed sandbox.env ship/token in the per-call lockdown
test instead of re-parsing it, and fix a comment that still described the
removed pipe-vs-heredoc stdin branch.
The lockdown tests pinned literal substrings ("( set -a", "exec python3",
"umask 077", "chmod 700"), which is change-detection. Replay the recorded
commands through a real bash under a temp root instead: dirs come out
0700 and the archive 0600, the sourced env reaches the child, the
child's exit code is the command's, and the token does not leak into the
outer shell (the session-snapshot class). The token-never-in-argv
assertions stay. No new test functions; the replay is skipped on
Windows, where the stdin/argv invariants are still checked.
_remote_write branched on getattr(env, "_stdin_mode", "pipe") and only
passed stdin_data on pipe backends, echoing base64 into argv elsewhere.
BaseEnvironment.execute already embeds stdin_data as a heredoc for
heredoc-mode backends (modal/daytona/vercel), and managed_modal forwards
it as stdinData; _write_to_sandbox already relies on that for every
backend. The branch duplicated base-class logic, and its defensive
getattr default meant a fake env with neither _stdin_mode nor a
stdin_data parameter raised TypeError on every RPC response write. The
poll loop swallowed the error, so no res_* file appeared and
test_code_execution_file_rpc hung forever (it passes on base).
Collapse to one path that always passes stdin_data, and teach the
file-RPC Shell fake to accept it and feed it as input. ScriptedEnv no
longer needs its _stdin_mode stub.
Keep one invariant per spawn path (persistent remote kernel and per-call
sandbox): token never on argv, dirs owner-only. Drop the command-string
duplicates, the malformed-seq replay test (the guard itself stays in
tools/code_execution_rpc.py) and the real-fs E2E class. The dropped
st_mode == 0o600 assertions were the Windows-failing ones flagged in
review, so no POSIX guard is needed on what remains.
On shared remote backends the execute_code channel created kernel and
sandbox dirs under shared temp at the process umask (775 group-writable
under umask 002), wrote request/result files group-readable, and carried
HERMES_RPC_TOKEN on remote command lines where co-tenant users read argv
via ps for the whole run. A co-tenant could read tool arguments and
results, and on group-writable dirs forge RPC requests dispatched under
the user's approval context.
- All remote dirs are created owner-only (umask 077 + chmod 700, checked
fail-closed) and every Hermes file write is mode 600.
- The token travels in a sourced env file inside a subshell so the vars
never enter the backend's session-snapshot dump, and ships via stdin on
pipe-capable backends so it never enters argv at all.
- The RPC poll loop rejects non-int seq requests before dispatch instead
of replaying them every cycle.
- tool_result_storage gets the same owner-only treatment for archived
tool output.
(cherry picked from commit aef21731d7fb8a4e0a6ada4ff9889264df4a8893)
test_eviction_skips_kernels_with_a_running_cell polled `_REMOTE_KERNELS.values()`
while the worker thread was inserting into the same dict, so the poll itself
raised "dictionary changed size during iteration" (seen on main and on PR runs
that never touch the file). Iterate a snapshot and yield between polls.
Same attached-cell guard as the local kernel host (#101861): a remote
kernel mid-cell is never reaped or cap-evicted, so a fan-out never has
its runner killed under a live poll loop.
Local session kernels sweep idle-expired entries and enforce a
process-wide cap (DEFAULT_MAX_SESSION_KERNELS) on every call
(tools/code_kernel.py's _reap_unlocked / _evict_over_cap_unlocked). The
new remote kernel host (#96991) never got the same treatment:
_REMOTE_KERNELS only shrinks lazily when a specific key is revisited and
found dead, so an owner that opens kernels for several distinct
(env_type, task_env_id) combinations (or delegated children) and never
revisits some of them accumulates host-side bookkeeping entries for the
life of the gateway process.
Note this is narrower than the local case: the remote runner already
self-reaps on its own idle timeout, and SSH/Docker connections are
independently bounded by their own transport-level lifecycles (SSH
ControlPersist, Docker's session-scoped idle-timeout in terminal_tool.py)
— so nothing here leaks a live remote connection. What's missing is
purely the host-side dict/cap bookkeeping symmetry with local kernels.
Adds _reap_unlocked/_evict_over_cap_unlocked mirroring the local
implementation, reusing the same max_session_kernels config as an
independent cap on _REMOTE_KERNELS.