16 Commits

Author SHA1 Message Date
kshitijk4poor
30c44c7555 fix(code-kernel): make post-result cell cleanup best-effort
The rm of the cell result file runs after the runner has executed the cell.
A transport failure there propagated out of _run_remote_cell, so
_run_attached_cell evicted the kernel and re-raised, and the caller's
per-call fallback then ran the same code a second time.

Catch and debug-log that failure (a leftover cell_res_* is harmless because
seq is monotonic), so the atomic ship is the only remote call that can raise
inside the discard scope and the handler's "runs exactly once" comment holds.
2026-09-27 00:56:31 +05:30
kshitijk4poor
9bf2c3c9fa fix(code-kernel): evict the remote kernel when a cell request fails to ship
The cell ship now fails closed (RuntimeError). _execute_remote catches it
and falls back to per-call execution, but the kernel stayed registered
with cell_seq bumped. The next call would then reuse a kernel whose
state silently missed this cell, with no state_lost flag. Discard (kill)
it the way the timeout path does, so the next call starts a fresh
kernel. The request never reached the runner (atomic publish), so the
fallback still runs the code exactly once.
2026-09-27 00:56:31 +05:30
kshitijk4poor
1a748cc85a refactor(code-execution): one checked-execute helper, complete launch command
Gate follow-ups on the remote lockdown:
- _execute_checked(env, cmd, what, **kw) in code_execution_rpc replaces
  the three copy-pasted "execute, raise if returncode != 0" blocks
  (per-call setup, kernel dir setup, file ship via
  _remote_write(check=True)). env.execute always returns a dict, so the
  isinstance/(r or {}) guards go; the error carries command output only,
  never the payload.
- _ship_env_file_and_launch_prefix returned a half-built "( ... && "
  that both callers had to close; a caller that dropped the ")" or
  composed it differently would lose the load-bearing subshell. It now
  takes the launch command, builds the shared env map (RPC dir, token,
  PYTHONDONTWRITEBYTECODE, routed TZ) itself, and returns the complete
  command; the kernel passes only HERMES_KERNEL_DIR/PYTHONPATH, which
  drops its duplicated TZ block and lazy hermes_time import.
- _private_dirs_cmd(root, *subdirs): every caller spelled each path
  twice for mkdir and chmod.
- _run_remote_cell publishes the cell request with one atomic
  _remote_write instead of ship-to-.tmp then a separate unchecked mv,
  saving a backend round-trip per cell.
2026-09-27 00:56:31 +05:30
beardthelion
5b8fd7fc32 fix(code-execution): lock down remote kernel/RPC dirs, keep RPC token out of argv
On shared remote backends the execute_code channel created kernel and
sandbox dirs under shared temp at the process umask (775 group-writable
under umask 002), wrote request/result files group-readable, and carried
HERMES_RPC_TOKEN on remote command lines where co-tenant users read argv
via ps for the whole run. A co-tenant could read tool arguments and
results, and on group-writable dirs forge RPC requests dispatched under
the user's approval context.

- All remote dirs are created owner-only (umask 077 + chmod 700, checked
  fail-closed) and every Hermes file write is mode 600.
- The token travels in a sourced env file inside a subshell so the vars
  never enter the backend's session-snapshot dump, and ships via stdin on
  pipe-capable backends so it never enters argv at all.
- The RPC poll loop rejects non-int seq requests before dispatch instead
  of replaying them every cycle.
- tool_result_storage gets the same owner-only treatment for archived
  tool output.

(cherry picked from commit aef21731d7fb8a4e0a6ada4ff9889264df4a8893)
2026-09-27 00:56:31 +05:30
teknium1
9a49b3c984 fix(execute_code): subagent kernels survive the LRU cap for the child's lifetime
A delegated child's execute_code kernel was keyed correctly
(<owner>:🧒:<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.

A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
2026-09-15 03:45:41 -07:00
liuzikaii
c99dbba33e fix(code-execution): remote kernels are keyed by their sandbox tool set (#97263)
The hermes_tools stub module a remote kernel imports is generated from sandbox_tools once,
at spawn, but the registry key was (owner, env_type, task_env_id) only: a later
execute_code call with a different tool set (skill loaded, toolset toggled) reused the
kernel and got stale stubs. The tool set is now part of the key; a different set gets
its own kernel and the over-cap eviction keeps the newest.

Salvaged from #97265 by @Liuzikaii.
2026-09-05 18:46:53 +05:30
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
0177c16903 fix(code-execution): remote kernel reap/evict skip kernels with a running cell
Same attached-cell guard as the local kernel host (#101861): a remote
kernel mid-cell is never reaped or cap-evicted, so a fan-out never has
its runner killed under a live poll loop.
2026-09-02 23:55:14 -07:00
nftpoetrist
d4126c6f49 fix(code-execution): reap idle and cap remote session kernels
Local session kernels sweep idle-expired entries and enforce a
process-wide cap (DEFAULT_MAX_SESSION_KERNELS) on every call
(tools/code_kernel.py's _reap_unlocked / _evict_over_cap_unlocked). The
new remote kernel host (#96991) never got the same treatment:
_REMOTE_KERNELS only shrinks lazily when a specific key is revisited and
found dead, so an owner that opens kernels for several distinct
(env_type, task_env_id) combinations (or delegated children) and never
revisits some of them accumulates host-side bookkeeping entries for the
life of the gateway process.

Note this is narrower than the local case: the remote runner already
self-reaps on its own idle timeout, and SSH/Docker connections are
independently bounded by their own transport-level lifecycles (SSH
ControlPersist, Docker's session-scoped idle-timeout in terminal_tool.py)
— so nothing here leaks a live remote connection. What's missing is
purely the host-side dict/cap bookkeeping symmetry with local kernels.

Adds _reap_unlocked/_evict_over_cap_unlocked mirroring the local
implementation, reusing the same max_session_kernels config as an
independent cap on _REMOTE_KERNELS.
2026-09-02 23:55:14 -07:00
Teknium
303af7a014 refactor(tools): tighten remote-kernel result assembly and per-call staging 2026-09-02 23:00:08 -07:00
Teknium
071f966bf2 refactor(tools): pack stub/doc tables and kernel attribute inits 2026-09-02 22:52:47 -07:00
Teknium
02f92ec707 refactor(tools): unify remote result assembly, compact execute_code docs and comments 2026-09-02 22:33:03 -07:00
Teknium
89881ad022 refactor(tools): simplify execute_code stack — drop kernel_mode shim, extract per-call remote path, reuse terminal config helpers 2026-09-02 22:10:19 -07:00
Teknium
606cb2de92 refactor(tools/code_exec): unify code_kernel local/remote helpers, split checkpoint_manager god methods, compact spill helpers 2026-09-02 14:44:15 -07:00
Teknium
5f75ec197b feat(code-execution): remote kernel host — session persistence for docker/ssh/modal backends (closes #96873) (#96991) 2026-08-28 01:39:33 -07:00