fix(code-kernel): make post-result cell cleanup best-effort

The rm of the cell result file runs after the runner has executed the cell.
A transport failure there propagated out of _run_remote_cell, so
_run_attached_cell evicted the kernel and re-raised, and the caller's
per-call fallback then ran the same code a second time.

Catch and debug-log that failure (a leftover cell_res_* is harmless because
seq is monotonic), so the atomic ship is the only remote call that can raise
inside the discard scope and the handler's "runs exactly once" comment holds.
This commit is contained in:
kshitijk4poor
2026-09-26 23:19:42 +05:30
committed by kshitij
parent c21acb3c1b
commit 30c44c7555

View File

@@ -321,7 +321,13 @@ def _run_remote_cell(kernel: RemoteKernel, code: str, timeout: int) -> Tuple[str
status = payload.get("status", "error")
except ValueError:
payload, status = {}, "protocol-error"
kernel.sh(f"rm -f {q_cells}/{q_res}", timeout=10)
try:
kernel.sh(f"rm -f {q_cells}/{q_res}", timeout=10)
except Exception:
# Best-effort: the cell already ran, so raising here would send
# the caller to its per-call fallback and run the code twice.
# A leftover result file is harmless (seq is monotonic).
logger.debug("remote kernel: cell result cleanup failed", exc_info=True)
return status, payload
time.sleep(_CELL_POLL_INTERVAL)
return "timeout", {}
@@ -381,7 +387,8 @@ def _run_attached_cell(kernel: RemoteKernel, key: Tuple, code: str, *, env, task
try:
cell_status, cell_payload = _run_remote_cell(kernel, code, timeout)
except Exception:
# The request never reached the runner (the atomic ship failed), so the
# The atomic ship is the only remote call here that can raise (the poll
# and result cleanup are best-effort), so the request never reached the runner and the
# caller's per-call fallback runs the code exactly once. Kill the
# kernel as the timeout path does: leaving it registered would let the
# next call reuse a kernel whose state silently missed this cell.