45 Commits

Author SHA1 Message Date
Halldrix
0d4dbde837 fix(delegate): mark elided timeout-diagnostic goal and eval tool output
The subagent timeout diagnostic and the session_search eval harness still
appended a bare "...[truncated]" marker, the imitable wording #121548
replaced everywhere else. Route both through agent.compression_marker.elide
so every elision in the tree mints the same counted, guard-matched marker.

Salvaged from #122392 (only the two call-site hunks; base already ships
the elide helpers the PR re-defined). Refs #121572.
2026-09-27 18:17:20 +05:30
teknium1
97beaeeeb3 feat(delegate): warn the child at 80% of its inactivity window before abandoning it
A child that stalls under a configured delegation.child_timeout_seconds used to
learn about the budget only by dying, losing its whole context. The liveness
wait now queues a one-line "[delegation budget warning]" through the child's
steer channel once the idle window is 80% spent (delivered at the child's next
iteration boundary), so a slow-but-recoverable child can wrap up and return its
summary. The warning fires once per idle window and re-arms when progress
resets the window; a progressing child never sees it.

Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged.
2026-09-20 15:50:20 -07:00
finn763
03973bd02b fix(delegate): child_timeout_seconds bounds inactivity, not total runtime
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a
plain `settled.wait(timeout=child_timeout)`, so any child that outlived the
budget was abandoned even while the provider was actively serving it.

The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could
not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint
— but the mechanism it names is exactly this gate: a child waiting on an
in-flight LLM completion, killed with a nearly-finished context. A slow child is
already bounded elsewhere (the per-call stale watchdog, the heartbeat's
staleness verdict), so this cap could only ever kill children the runtime had
judged healthy.

`child_timeout_seconds` now measures time with NO progress: the wait runs in
slices and restarts the window on the same signals the heartbeat's stale verdict
reads (completed call, tool change, activity-clock tick). A frozen child is
still abandoned when the window elapses; a progressing one is never killed for
taking long.

Timeout entries also carry `last_event_age` (how long the child had been silent),
so operators can tell a slow provider from a runaway without transcript
forensics.

Fixes the mechanism reported in #116001. The budget warning and
continuation-respawn items in that issue are separate features and are not part
of this change.
2026-09-20 15:50:20 -07:00
teknium1
c38f7f6d4c fix: bind the delegated child to the entry actually leased, not the pool's shared cursor
_lease_child_credential resolved the leased entry via child_pool.current()
after acquire_lease(); the pool is shared with the parent and siblings, so
_current_id is a mutable cursor that may already point at another caller's
pick. Resolve by the returned lease id from child_pool.entries() instead.
2026-09-19 09:34:09 -07:00
Ofer LaOr
d75a51872c fix: delegated child never leases a same-provider pool entry for another endpoint
A delegated child with an explicit base_url (e.g. an Azure OpenAI resource
under provider "openai") shared the parent's "openai" credential pool on bare
provider equality, and _lease_child_credential bound whatever entry
acquire_lease() picked. _swap_credential adopts the entry's base_url as well,
so the child was silently rebound to https://api.openai.com/v1, sent the
Azure key there, got HTTP 401 and only then fell back (#68237).

_resolve_child_credential_pool now requires endpoint coherence on top of
provider identity (credential_pool_matches_provider + at least one entry whose
base_url matches the child's) for both the shared parent pool and the
provider's loaded pool; a mismatched pool is not attached and the child keeps
its fixed credential. _lease_child_credential validates the leased entry
against the child's base_url and, on a mixed pool, releases the wrong-host
lease and leases an endpoint-matching entry by id instead. Entries and
adapters without endpoint metadata cannot rebind and are accepted unchanged.

Salvaged from #68240 (@oferlaor); the acquire_lease/leased_entry filter
parameters and the extra pool-level helper were reduced to the two small
delegate-side predicates.
2026-09-19 09:34:09 -07:00
teknium1
b71359066c fix: /stop halts background subagents and returns their partial results as interrupted completions
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:

- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
  already did this in `_handle_reset_command`, the second request is
  idempotent) and the idle `_handle_stop_command` tail, which replied
  "No active task to stop." while a background child was running; it now
  stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
  spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.

The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.

An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.

Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.

Part of #114456
2026-09-18 20:34:42 -07:00
teknium1
6b1cf3a68e fix(delegation): a stale verdict under a configured cap names the stale threshold, not the cap
When delegation.child_timeout_seconds is set, the heartbeat stale verdict
pre-empts the cap but the entry still said "timed out after 3600s" with
timeout_seconds=3600 although the wait ended at the ~450s stale threshold.
The heartbeat now records the threshold it fired at and await_child
branches on that instead of on "no cap configured", so both the message and
timeout_seconds report what actually ended the wait.
2026-09-16 16:56:46 -07:00
teknium1
90663bff30 fix(delegation): a stale subagent ends the sync delegation wait instead of holding the turn forever
On a finite runtime (`hermes chat -Q`, the Bot Chat one-shot delivery child) delegate_task falls
back to a synchronous batch and had no bound at any layer: delegate_task is exempt from the
sequential tool deadline, delegation.child_timeout_seconds defaults to off, and the heartbeat
only stopped itself on its stale verdict ("let gateway timeout fire") — a watchdog a -Q turn does
not have. A worker wedged after its final answer therefore held the parent turn, the Bot Chat
session lease, and the process forever.

_Heartbeat.settled is set on the stale verdict and _ChildRun.await_child waits on it (the
worker future done-callback sets it too), so the wedge ends through the existing timeout /
deferred-close path with a structured `timeout` entry naming the stalled progress.

Slim redo of the mechanism proposed in #109762 (@fangliquanflq, earliest) and #113275
(@JoaoMarcos44); their quarantine subsystem, 50 ms polling loop and collect-grace were dropped.

Fixes #109749

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-16 16:56:46 -07:00
teknium1
6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1
8e16bde491 test: fold the schema-retry context test into the output-schema module
Reuse tests/tools/test_delegate_output_schema.py's _StubChild instead of a
new one-test file with its own double; the invariant (the retry turn sees
is_delegated_child_context() True and the flag is restored afterwards) is
unchanged. Trim the source comment to the WHY.
2026-09-15 18:41:19 -07:00
moep90
e002cdb92d fix(delegation): run the schema-retry turn in the delegated-child context
_validate_child_output_schema issues a second run_conversation on the child when
the first answer fails the declared output_schema. The main child turn is wrapped
in delegated_child_context; this one was not. It runs on the parent worker's
thread, where HERMES_KANBAN_TASK is set and nothing marks the execution as a
child, so every identity gate keyed on is_delegated_child_context() fails open.

The visible effect is the kanban stop guard: it nudges the child to call
kanban_complete or kanban_block. A child owns no board task and carries no kanban
toolset, so it cannot, and the nudge text ("do not narrate intent", "finish any
remaining deliverable") displaces the structured answer the retry exists to
produce. The retry then fails the same schema and delegate_task reports an error
for a child whose work was already complete.

Observed with four children, each nudged during its retry:

  [subagent-0] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-2] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-3] Kanban worker tried to exit without kanban_complete/kanban_block
  [subagent-1] Kanban worker tried to exit without kanban_complete/kanban_block
  4/4 - Final answer does not satisfy the declared output_schema (after 1 retry)

Wrap the retry the same way the main turn is wrapped. The context is entered and
exited around the single call, so nothing outside the retry sees it.

Signed-off-by: moep90 <volleyballlive@googlemail.com>
2026-09-15 18:41:19 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
45ab3ad57f fix(delegate_task): return a schema-invalid child's raw text instead of failing the task
When a child's final answer still missed its output_schema after the one
bounded retry, the result entry flipped to status=failed with the error
"Final answer does not satisfy the declared output_schema" — the completion
line printed ✗ and orchestrators read a finished audit as a failure. Five
audits of 413-4103 s were lost this way in the Sep 10-14 retrospective and
the parent had to mine the live transcripts; in four of them the "violation"
was a ```json fence around a valid array, which the candidate extractor
sliced to its first..last object.

Now: status stays completed, `summary` is the child's raw final text,
`schema_valid: false` + `schema_errors` carry the verdict and a `schema_note`
says the text is unvalidated; the sync completion line shows ⚠ with the
reason. The extractor tries the earliest-opening bracket span and keeps the
first that parses (fenced arrays validate). The OUTPUT CONTRACT the child
sees now says "ONLY the JSON value — no prose, no code fence" and what a miss
costs. One bounded retry is unchanged.
2026-09-15 03:45:41 -07:00
teknium1
9a49b3c984 fix(execute_code): subagent kernels survive the LRU cap for the child's lifetime
A delegated child's execute_code kernel was keyed correctly
(<owner>:🧒:<session>) but counted against the process-wide
max_session_kernels LRU cap (default 4) like any other kernel. In a fan-out
wider than the cap every child's first cell spawned a kernel and evicted the
oldest sibling's, so the sibling's next cell started a fresh interpreter and
NameError'd on state its own previous cell had set — while the tool schema
promised "variables, imports, and loaded data survive across execute_code
calls". Finished children's kernels also squatted the cap for
kernel_idle_timeout (1800 s) after the child was gone. 48 NameErrors across 28
subagent lanes in the Sep 10-14 retrospective.

A live child's kernel (local and remote) is now pinned: exempt from LRU
eviction while the child runs, disposed by the delegation cleanup path
(shutdown_kernels_for_delegated_child) as soon as the child finishes. Top-level
sessions keep the existing cap and idle reaping unchanged.
2026-09-15 03:45:41 -07:00
teknium1
60559d4e0e fix(delegation): late-attached children take the parent's soft/hard stop kind; dedupe the fallback replay
_attach_child now mirrors a pending parent stop with the same split
interrupt() uses for its own fan-out (hard -> hard_interrupt, soft ->
interrupt), so a redirect is not turned into a cancel on a child that was
attached late. _restore_parent_cancellation collapses to re-attaching the
rejected unit's children: the replay is the attach step's job now.

Test fixture: _Batch gained origin_session_history_delivery on main after
the salvaged PR was written.

Co-authored-by: illidan <noequal666@gmail.com>
2026-09-13 21:31:57 -07:00
teknium1
230ca004a7 fix(delegate): forward inline data-URL images to vision children; trim tests to invariants
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.

Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
2026-09-13 21:05:42 -07:00
Teknium
f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
teknium1
dd497c3d59 fix(delegate): grandchildren spawned after their orchestrator was stopped now die with it
AIAgent.interrupt() fans the stop out to a snapshot of _active_children. A child
that is attached after that snapshot — an orchestrator subagent still building
its fan-out siblings, or one that has not yet hit its next iteration check —
started with no signal and ran to completion as an orphan while its parent had
already reported `interrupted`. Live repro (mid orchestrator, stop delivered
between grandchild A and B builds): grandchild B kept its `sleep 20` alive and
stayed in the subagent registry after the mid returned.

_attach_child now mirrors a pending parent stop onto the newcomer, so the whole
spawn tree dies with the node that was stopped. Tests: unit (late attach gets
the stop, normal attach does not) + the real _build_child_agent path with a
stopped orchestrator.
2026-09-13 20:13:12 -07:00
finn763
1562b87d5d fix(agent): isolate periodic scheduler callbacks from blocking siblings (#102574) 2026-09-09 12:17:57 +05:30
Teknium
ef9239571d feat(delegation): report a child's exited-but-unread notify processes to the parent
A process that finishes while the child is alive needs no handoff, but if the child never
polls/waits/logs it, the result vanished: the completion notice is suppressed in the parent
and the child's summary never mentions it. Finalization now attaches exit code + output
tail as unread_completions, rendered in the parent's delegation notice.
2026-09-07 12:50:29 -07:00
Teknium
3c0d90e8ef feat(delegation): subagents hand background processes to the parent; leftovers are named, not trusted
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).

- process_manage(action="handoff", session_id, data="<purpose>"), children
  only: process_registry.transfer_ownership flips owner_task_id/task_id/
  session_key to the parent under the registry lock, so the completion is
  stamped with the parent's owner at exit, passes the parent's sa- filter,
  and is reaped by the parent, not the child. Cap 3 per child; an exited,
  foreign, or non-child request is a tool error. The purpose rides the
  event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
  notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
  orphaned_processes on the result before cleanup kills the leftovers; the
  parent's delegation block renders both.
2026-09-07 12:50:29 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
9cf7c74030 refactor(delegate): module headers condensed (one-line docstrings, logger comment inline) 2026-09-02 19:43:56 -07:00
Teknium
0a51863181 refactor(delegate): child_run failure entry/emit_complete tightened 2026-09-02 19:39:28 -07:00
Teknium
defca3639a refactor(delegate): reflow comments/docstrings to 118 cols (word-identical, AST-identical) 2026-09-02 19:25:04 -07:00
Teknium
657628f64d refactor(delegate): child_run — worktree setup via _quiet, packed registry record, diag header table 2026-09-02 19:21:01 -07:00
Teknium
1283ccd855 refactor(delegate): docstring compaction by hand (every rule/why kept) 2026-09-02 19:16:07 -07:00
Teknium
eabec2ff50 refactor(delegate): AST-identical re-layout pass 2 2026-09-02 19:04:50 -07:00
Teknium
660235a228 refactor(delegate): child_run diagnostic sections via one helper; tighter heartbeat/result-entry bodies 2026-09-02 19:02:55 -07:00
Teknium
fe9c26ee5d refactor(delegate): _ChildRun dataclass replaces _WorktreeReporter/_ChildWorkspace/_ChildFailure plumbing 2026-09-02 18:27:59 -07:00
Teknium
ce19685282 refactor(delegate): merge single-caller _handle_child_wait_failure into _await_child; _diag_section helper 2026-09-02 18:22:16 -07:00
Teknium
378692be4f refactor(delegate): fold _construct_child_agent/_announce_child_spawn into _build_child_agent; _quiet best-effort sites 2026-09-02 18:04:43 -07:00
Teknium
b18228f377 refactor(delegate): AST-identical re-layout (pack signatures/call args, join split literals) 2026-09-02 17:59:28 -07:00
Teknium
901ba7db6f refactor(delegate): child_run best-effort blocks via _quiet; shared failure-entry tail 2026-09-02 17:54:08 -07:00
Teknium
3da2ff88e3 refactor(delegate): pack re-import block; single blank line between top-level defs (AST-identical) 2026-09-02 17:03:55 -07:00
Teknium
987bc82574 refactor(delegate): compact docstrings/comments by hand (every WHY kept); collapse orchestrator toolset branches 2026-09-02 17:03:05 -07:00
Teknium
016f31375d refactor(delegate): join wrapped logical lines that fit in 120 cols (AST-identical) 2026-09-02 16:59:26 -07:00
Teknium
e204008f10 refactor(delegate): _knob() unifies config>env>default int/float knobs; timeout diagnostic sections -> _diag_sizes/_diag_threads + attr table; drop dead _handle_child_wait_failure re-import 2026-09-02 16:57:22 -07:00
Teknium
884c697291 refactor(delegate): split _finalize_child_results (memory/hooks/cost rollup), lift _create_isolated_worktree and _report_child_done, collapse dead worktree-attach import fallback 2026-09-02 16:54:15 -07:00
Teknium
338da30a45 refactor(delegate): _resolve_child_runtime returns AIAgent kwargs (drop _ChildRuntime; routing-filter table); dispatch reuses _detach_child/_signal_child_stop; control actions -> _CONTROL_OUTCOMES table; attribution lookup dedupe 2026-09-02 16:50:57 -07:00
Teknium
52419bda5d refactor(delegate): split delegate_task into _normalize_task_list/_coerce_task_schemas/_build_children/_execute_and_aggregate(_Batch); _build_child_agent -> _open_child_session_db/_construct_child_agent/_announce_child_spawn; _run_single_child -> _lease_child_credential/_await_child/_merge_late_steer 2026-09-02 16:46:39 -07:00
Teknium
4711e8105d refactor(delegate): child_run dedupe — shared close/attach/detach/stop helpers, _defer_close_after_timeout, _build_tool_trace, _num; single _fabricated_entry 2026-09-02 16:37:47 -07:00
Teknium
c1f8fd4e40 wip(delegate): child_run dedupe snapshot (unverified; worker died at 401) 2026-09-02 16:33:00 -07:00
Teknium
05526b028a refactor(tools/delegate): split delegate_tool into child_run/config/dispatch/progress/registry/results; compact delegation helpers 2026-09-02 14:45:15 -07:00