fix(delegate): nested orchestrators get their workers' results back — delegate_task exempt from the 420 s tool deadline; summary budget uses current prompt, not the session sum
Independent review found two holes in the first fix. A parent with no usage
row yet was treated as 0 tokens used, so a 190K/200K prompt received a
384K-char dynamic summary budget instead of ~4K; the budget now returns None
(static ceiling only) when nothing is known. And under MoA the folded usage
includes advisor prompts that are not in the parent's context, over-stating
the prompt size and wrongly truncating summaries; turn_usage now records the
aggregator's pre-fold prompt_tokens as _last_prompt_size_tokens and the
budget reads that first.
Tests (2 new): unknown usage -> None; MoA-folded and unfolded parents with
the same real prompt get the same budget.
web_extract / browser_snapshot / delegate_task spill their full text under HERMES_HOME/cache,
which is mounted read-only into docker/modal (at /root/.hermes) and synced under ~/.hermes for
ssh/daytona/vercel — but the footer told the agent the HOST path, so read_file inside the
sandbox got 'File not found'. Translate through the existing
credential_files.to_agent_visible_cache_path (what tool_result_storage already does); local and
singularity are unchanged.
Salvaged from #72429 by @JonthanaHanh (the web_extract sites), widened to every footer.
Two defects in the same path, both measured on the 1,393-agent refactor run.
1. A nested orchestrator (depth > 0) runs delegate_task synchronously by
design: it needs its workers' results inside its own turn. But the
sequential tool runner put every tool call under the generic 420 s
deadline, and delegate_task was not exempt, so every batch longer than
seven minutes returned "timed out after 420.0s" while the children kept
running as orphans. 332 such timeouts in 234 orchestrator sessions; only
89 nested delegate_task calls in the whole run ever returned a real result.
The orchestrators then spent 388 h of wall time polling: 1,526 reads of
the live transcript files, 551 list actions, 242 h of explicit sleep,
about $4k of API turns. delegate_task is now exempt from the sequential
deadline (the batch owns its liveness: per-child heartbeats, the stale
monitor, delegation.child_timeout_seconds).
Live A/B, depth-1 orchestrator dispatching a 75 s leaf with the deadline
set to 40 s (glm-5.3-flash via Nous): main -> "Error executing tool
'delegate_task': timed out after 40.0s", leaf result lost; branch ->
orchestrator blocked 161 s and returned the leaf's LEAF_DONE_MARKER.
2. _parent_summary_char_budget computed the parent's context headroom from
session_prompt_tokens, which is the running SUM of prompt tokens over
every API call in the session. After a few hundred calls it exceeds any
window, headroom goes negative, and every child summary collapses to the
2,000-char floor with the full text spilled to disk. All 1,393 child
summaries in the run were truncated this way; the orchestrator planned
from stubs. The budget now reads the last call's prompt_tokens from
_last_turn_usage.
Tests: delegate_task is in the exempt set and the set is narrow; budget for
a long-lived parent equals the budget for a fresh parent with the same
current prompt, and exceeds the floor.