* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
790 lines
43 KiB
Python
790 lines
43 KiB
Python
#!/usr/bin/env python3
|
|
"""
|
|
Delegate Tool -- Subagent Architecture
|
|
|
|
Spawns child AIAgent instances with a fresh conversation, their own task_id
|
|
(terminal session, file-ops cache), the parent's toolsets minus child-blocked
|
|
tools, and a focused system prompt built from goal + context. Single-task and
|
|
batch (parallel) modes; top-level model calls run in the background while
|
|
orchestrator children wait for their own workers. The parent only ever sees
|
|
the delegation call and the summary result, never the child's intermediate
|
|
tool calls or reasoning.
|
|
"""
|
|
|
|
import logging
|
|
import time
|
|
import weakref
|
|
from typing import Any, Dict, List, Optional
|
|
|
|
from tools.terminal_tool import set_approval_callback as _set_subagent_approval_cb # noqa: F401 (used via _ChildRun.await_child)
|
|
from utils import is_truthy_value
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
# The delegate_tool_* siblings hold the pieces split out of this module; every name callers or patching tests reach as
|
|
# ``tools.delegate_tool.<name>`` is re-imported here. Mutable flag globals live only in their owning module.
|
|
from tools.delegate_tool_child_run import ( # noqa: F401
|
|
_ChildRun, _attach_child, _build_child_goal_message, _build_result_entry, _dump_subagent_timeout_diagnostic, _fabricated_entry,
|
|
_lease_child_credential, _merge_late_steer, _register_child, _start_heartbeat, _validate_child_output_schema,
|
|
)
|
|
from tools.delegate_tool_config import ( # noqa: F401
|
|
_DEFAULT_MAX_CONCURRENT_CHILDREN, _get_child_timeout, _get_max_async_children, _get_max_concurrent_children,
|
|
_get_max_spawn_depth, _get_oneshot_max_children, _get_orchestrator_enabled, _get_subagent_approval_callback, _get_worktree_isolation,
|
|
_inherit_parent_capabilities, _load_config, _merge_request_overrides, _resolve_child_credential_pool,
|
|
_resolve_child_runtime, _resolve_delegation_credentials,
|
|
_subagent_auto_approve, _subagent_auto_deny,
|
|
)
|
|
from tools.delegate_tool_dispatch import _Batch, _announce_batch, _capture_origin, _run_batch
|
|
from tools.delegate_tool_progress import ( # noqa: F401
|
|
DelegateEvent, SUBAGENT_FAILURE_STATUSES, _batch_prefix, _build_child_progress_callback,
|
|
_build_child_system_prompt, _clean_error_text, _emit_parent_console, _quiet, _resolve_workspace_hint,
|
|
_safe_progress, format_batch_tag, format_subagent_failure_line,
|
|
)
|
|
from tools.delegate_tool_registry import ( # noqa: F401
|
|
_CONTROL_ACTIONS, _active_subagents, _active_subagents_lock, _capture_gateway_steer_authority,
|
|
_handle_control_action, _is_descendant_of, _owns_subagent_record, _register_subagent, _unregister_subagent,
|
|
get_subagent_attribution, interrupt_subagent, is_spawn_paused, list_active_subagents, set_spawn_paused,
|
|
steer_subagent,
|
|
)
|
|
from tools.delegate_tool_tasks import ( # noqa: F401
|
|
_MAX_TASK_IMAGES, _coerce_task_images, _coerce_task_schemas, _normalize_task_images, _normalize_task_list,
|
|
)
|
|
from tools.delegate_tool_toolsets import ( # noqa: F401
|
|
DELEGATE_BLOCKED_TOOLS, _expand_parent_toolsets, _resolve_child_toolsets, _strip_blocked_tools,
|
|
)
|
|
from tools.delegate_tool_results import ( # noqa: F401
|
|
_apply_summary_budget, _build_child_preserving_parent_tools, _run_child_lifecycle, _summarize_tool_arguments,
|
|
)
|
|
|
|
_ROLES = frozenset({"leaf", "orchestrator"})
|
|
|
|
# Nested delegation is granted by depth/role in _build_child_agent, never by the
|
|
# model naming toolsets (there is no model-facing toolsets argument).
|
|
def _normalize_role(r: Optional[str]) -> str:
|
|
"""'leaf' | 'orchestrator'; None/empty/unknown -> 'leaf' (unknown warns)."""
|
|
r_norm = str(r).strip().lower() if r else "leaf"
|
|
if r_norm not in _ROLES:
|
|
logger.warning("Unknown delegate_task role=%r, coercing to 'leaf'", r)
|
|
return "leaf"
|
|
return r_norm
|
|
|
|
DEFAULT_MAX_ITERATIONS = 250
|
|
_HEARTBEAT_INTERVAL = 30 # seconds between parent activity heartbeats during delegation
|
|
# Stale-heartbeat thresholds (cycles of _HEARTBEAT_INTERVAL with no progress). Progress = iteration, current_tool OR
|
|
# last_activity_ts advancing; an in-flight model wait refreshes last_activity_ts, so slow models are not "idle". Idle
|
|
# stays tight so a truly wedged child doesn't mask the gateway timeout; in-tool is much higher so legitimately long
|
|
# tools can finish.
|
|
_HEARTBEAT_STALE_CYCLES_IDLE = 15 # 450s idle between turns → stale
|
|
_HEARTBEAT_STALE_CYCLES_IN_TOOL = 40 # 1200s stuck on same tool → stale
|
|
|
|
def check_delegate_requirements() -> bool:
|
|
"""Delegation has no external requirements -- always available."""
|
|
return True
|
|
|
|
|
|
def _open_child_session_db(parent_agent) -> Any:
|
|
"""DEDICATED SessionDB handle for the child, or None: the parent's handle can be closed by its own lifecycle while
|
|
a background child still flushes (transcript silently dropped). It MUST open the same db FILE as the parent's
|
|
handle (non-launch profiles), else lineage / session_search break; released by the child's close() via
|
|
_owns_session_db."""
|
|
# Each child gets a DEDICATED SessionDB connection instead of the parent's live object. The parent's
|
|
# handle is owned by the parent's lifecycle (cron run_job's finally block, gateway session end, /new)
|
|
# and can be closed while a fire-and-forget background child is still flushing on a daemon thread —
|
|
# every subsequent flush then hits the closed handle and the child's transcript is silently dropped
|
|
# (#81267). It MUST point at the same database FILE as the parent's handle: parents can hold non-default
|
|
# per-profile handles (tui_gateway opens SessionDB(db_path=<profile>/ state.db) for non-launch
|
|
# profiles), and a bare SessionDB() would write the child's transcript into the launch profile's db,
|
|
# breaking parent_session_id lineage and session_search. AsyncSessionDB wrappers (gateway) forward
|
|
# .db_path via __getattr__, so this works through them.
|
|
parent_session_db = getattr(parent_agent, "_session_db", None)
|
|
if parent_session_db is None:
|
|
return None
|
|
with _quiet("subagent: failed to open dedicated SessionDB; child persistence disabled", exc_info=True):
|
|
from hermes_state_registry import acquire
|
|
_parent_db_path = getattr(parent_session_db, "db_path", None)
|
|
return acquire(_parent_db_path) if _parent_db_path is not None else acquire()
|
|
return None
|
|
|
|
def _apply_child_cache_ttl(child) -> None:
|
|
"""A delegated child never uses the 1h cache tier. The tier is priced for a person who steps
|
|
away between turns (2x write vs 1.25x for 5m, #14971); a subagent calls every few seconds for
|
|
minutes and is gone, so it pays the 2x on every tool result and never collects the retention.
|
|
Caching itself stays exactly as configured (disabled stays disabled)."""
|
|
if getattr(child, "_cache_ttl", None) == "1h":
|
|
child._cache_ttl = "5m"
|
|
|
|
_CHILD_CAP_MIN = 16_000 # below this a child compresses on every call; treat as a config error
|
|
|
|
|
|
def _child_compression_cap_tokens(raw) -> "int | None":
|
|
"""Validated ``delegation.compression_threshold_tokens``: an int >= 16000, or None for "no cap".
|
|
|
|
Unset / ``0`` / ``false`` / ``null`` mean no subagent-specific cap: the child compacts at the
|
|
same ratio trigger as everyone else (0.50 x window). A bool ``true`` (YAML) would coerce to 1
|
|
and make every call compress; a string like ``"200k"`` would silently read as no cap. Both are
|
|
config errors: warn and treat as unset so a typo never changes compaction behaviour."""
|
|
if raw is None or raw is False or raw == 0:
|
|
return None
|
|
if isinstance(raw, bool) or not isinstance(raw, (int, float)) or int(raw) < _CHILD_CAP_MIN:
|
|
logger.warning(
|
|
"delegation.compression_threshold_tokens=%r is not a token count >= %d; ignoring it "
|
|
"(children keep the ratio trigger).", raw, _CHILD_CAP_MIN,
|
|
)
|
|
return None
|
|
return int(raw)
|
|
|
|
|
|
def _apply_child_compression_cap(child, delegation_cfg: dict) -> None:
|
|
"""Optional absolute cap on the child's compaction trigger, ``delegation.compression_threshold_tokens``
|
|
(lower of it and any global ``compression.threshold_tokens``). Off by default: a 1M-window child
|
|
compacts where its parent does. The compressor applies the cap on first window resolution, which
|
|
happens after construction, so setting it here is exactly equivalent to config."""
|
|
from agent.context_compressor import ContextCompressor
|
|
|
|
cc = getattr(child, "context_compressor", None)
|
|
if not isinstance(cc, ContextCompressor):
|
|
return
|
|
cap = _child_compression_cap_tokens((delegation_cfg or {}).get("compression_threshold_tokens"))
|
|
if cap is None:
|
|
return
|
|
existing = cc.threshold_tokens_cap
|
|
cc.threshold_tokens_cap = min(cap, existing) if isinstance(existing, int) and existing > 0 else cap
|
|
if cc._threshold_tokens is not None: # already resolved: re-clamp now
|
|
cc._apply_threshold_tokens_cap()
|
|
|
|
|
|
def _build_child_agent(
|
|
task_index: int,
|
|
goal: str,
|
|
context: Optional[str],
|
|
toolsets: Optional[List[str]],
|
|
model: Optional[str],
|
|
max_iterations: int,
|
|
task_count: int,
|
|
parent_agent,
|
|
# Credential overrides from delegation config
|
|
override_provider: Optional[str] = None,
|
|
override_base_url: Optional[str] = None,
|
|
override_api_key: Optional[str] = None,
|
|
override_api_mode: Optional[str] = None,
|
|
override_request_overrides: Optional[Dict[str, Any]] = None,
|
|
|
|
# ACP transport overrides from trusted delegation config.
|
|
override_acp_command: Optional[str] = None,
|
|
override_acp_args: Optional[List[str]] = None,
|
|
# Configuration block that owns the selected provider/model route. Internal
|
|
# callers such as /review pass auxiliary.review here so fallback policy is
|
|
# not accidentally read from the general delegation block.
|
|
routing_cfg: Optional[Dict[str, Any]] = None,
|
|
# Legacy; accepted for wire compat but ignored (capability is depth-derived).
|
|
role: str = "leaf",
|
|
):
|
|
"""Build (don't run) a child AIAgent on the main thread. override_* (from delegation config) replace parent
|
|
inheritance so children can run on a different provider:model pair."""
|
|
import uuid as _uuid
|
|
from run_agent import AIAgent
|
|
from agent.delegation_context import delegated_child_context
|
|
# Role is depth-derived: a child may delegate iff the kill switch is on and
|
|
# depth budget remains below max_spawn_depth. The `role` arg is ignored.
|
|
child_depth = getattr(parent_agent, "_delegate_depth", 0) + 1
|
|
max_spawn = _get_max_spawn_depth()
|
|
effective_role = "orchestrator" if _get_orchestrator_enabled() and child_depth < max_spawn else "leaf"
|
|
|
|
# One subagent_id shared by the progress callback, spawn_requested event and
|
|
# the live registry; parent_id is set when THIS parent is itself a subagent.
|
|
subagent_id = f"sa-{task_index}-{_uuid.uuid4().hex[:8]}"
|
|
parent_subagent_id = getattr(parent_agent, "_subagent_id", None)
|
|
|
|
# General delegation behavior (reasoning, compression, capabilities) stays
|
|
# global. Only fallback policy follows the owner of a per-call route such
|
|
# as auxiliary.review.
|
|
delegation_cfg = _load_config()
|
|
child_toolsets, child_disabled_toolsets = _resolve_child_toolsets(parent_agent, toolsets, effective_role)
|
|
child_prompt = _build_child_system_prompt(
|
|
goal, context, workspace_path=_resolve_workspace_hint(parent_agent), role=effective_role,
|
|
max_spawn_depth=max_spawn, child_depth=child_depth,
|
|
)
|
|
parent_api_key = getattr(parent_agent, "api_key", None)
|
|
if (not parent_api_key) and hasattr(parent_agent, "_client_kwargs"):
|
|
parent_api_key = parent_agent._client_kwargs.get("api_key")
|
|
|
|
# Shared ref: session_id once the child exists, delegation_id once
|
|
# delegate_task stamps it — both ride on every relayed event.
|
|
child_session_ref: Dict[str, Any] = {}
|
|
child_progress_cb = _build_child_progress_callback(
|
|
task_index, goal, parent_agent, task_count, subagent_id=subagent_id, parent_id=parent_subagent_id,
|
|
depth=max(0, child_depth - 1), # 0 = first-level child for the UI
|
|
model=model or getattr(parent_agent, "model", None), toolsets=child_toolsets, session_ref=child_session_ref,
|
|
)
|
|
rt = _resolve_child_runtime(
|
|
parent_agent, delegation_cfg, parent_api_key, model=model, override_provider=override_provider,
|
|
override_base_url=override_base_url, override_api_key=override_api_key, override_api_mode=override_api_mode,
|
|
override_acp_command=override_acp_command,
|
|
override_acp_args=override_acp_args,
|
|
routing_cfg=routing_cfg,
|
|
)
|
|
if override_request_overrides is not None:
|
|
# honored whenever set, incl. the inherit branch where
|
|
# _resolve_delegation_credentials already merged OVER the parent's
|
|
request_overrides = dict(override_request_overrides)
|
|
else:
|
|
request_overrides = {} if override_provider else dict(getattr(parent_agent, "request_overrides", {}) or {})
|
|
parent_sid = getattr(parent_agent, "session_id", None)
|
|
child_session_db = _open_child_session_db(parent_agent)
|
|
with delegated_child_context():
|
|
try:
|
|
child = AIAgent(
|
|
**rt, max_iterations=max_iterations, prefill_messages=getattr(parent_agent, "prefill_messages", None),
|
|
enabled_toolsets=child_toolsets, disabled_toolsets=child_disabled_toolsets, quiet_mode=True,
|
|
ephemeral_system_prompt=child_prompt, log_prefix=f"[subagent-{task_index}]", platform="subagent",
|
|
side_agent=True,
|
|
skip_context_files=True, skip_memory=True, clarify_callback=None,
|
|
thinking_callback=(
|
|
(lambda text: _safe_progress(child_progress_cb, "_thinking", text) if text else None)
|
|
if child_progress_cb else None
|
|
),
|
|
session_db=child_session_db, parent_session_id=parent_sid, request_overrides=request_overrides,
|
|
tool_progress_callback=child_progress_cb,
|
|
iteration_budget=None, # fresh budget per subagent
|
|
)
|
|
except BaseException:
|
|
# No child close() will ever run: release the dedicated handle here.
|
|
if child_session_db is not None:
|
|
with _quiet(None):
|
|
from hermes_state_registry import release_or_close
|
|
release_or_close(child_session_db)
|
|
raise
|
|
child._print_fn = getattr(parent_agent, "_print_fn", None)
|
|
_apply_child_cache_ttl(child)
|
|
if child_session_db is not None:
|
|
child._owns_session_db = True # released by the child's close(), never by the parent
|
|
# Ownership transfer for the dedicated handle: the child's close() must release it (nothing else holds a
|
|
# reference), and no parent teardown can close it out from under a background child (#81267).
|
|
child_session_ref["session_id"] = getattr(child, "session_id", "") or ""
|
|
child._progress_identity_ref = child_session_ref
|
|
child._delegate_depth, child._delegate_role = child_depth, effective_role # post-degrade role
|
|
child._subagent_id, child._parent_subagent_id = subagent_id, parent_subagent_id
|
|
_apply_child_compression_cap(child, delegation_cfg)
|
|
# Ownership chain for action=list/steer/stop; weakref so a finished parent
|
|
# can be collected while a detached child record lingers in the registry.
|
|
try:
|
|
child._delegate_parent_ref = weakref.ref(parent_agent)
|
|
except TypeError:
|
|
child._delegate_parent_ref = None # non-weakref-able test doubles
|
|
# Sidebar marker: subagent sessions stay out of session pickers even when a
|
|
# parent delete orphans them (mirrors /branch's ``_branched_from``).
|
|
if parent_sid and getattr(child, "_session_init_model_config", None) is not None:
|
|
child._session_init_model_config["_delegate_from"] = parent_sid
|
|
# Shared pool lets children rotate credentials on rate limits.
|
|
child_pool = _resolve_child_credential_pool(
|
|
rt["provider"], parent_agent, rt["base_url"], effective_requested_provider=rt.get("requested_provider"),
|
|
)
|
|
if child_pool is not None:
|
|
child._credential_pool = child_pool
|
|
|
|
_attach_child(parent_agent, child) # interrupt propagation
|
|
# spawn_requested now — the child may queue for seconds when the pool is
|
|
# saturated — then the subagent_start lifecycle hook.
|
|
_safe_progress(child_progress_cb, "subagent.spawn_requested", preview=goal)
|
|
with _quiet("subagent_start hook invocation failed", exc_info=True):
|
|
from hermes_cli.lifecycle import invoke_hook as _invoke_hook
|
|
_invoke_hook(
|
|
"subagent_start", parent_session_id=parent_sid,
|
|
parent_turn_id=getattr(parent_agent, "_current_turn_id", "") or "", parent_subagent_id=parent_subagent_id,
|
|
child_session_id=getattr(child, "session_id", None), child_subagent_id=subagent_id,
|
|
child_role=effective_role, child_goal=goal,
|
|
)
|
|
return child
|
|
|
|
def _run_single_child(
|
|
task_index: int, goal: str, child=None, parent_agent=None, *, owner_session_id: Optional[str] = None,
|
|
owner_transport: Any = None, owner_session_record: Any = None, **_kwargs,
|
|
) -> Dict[str, Any]:
|
|
"""Run a pre-built child agent (called from a worker thread) and return its result entry.
|
|
|
|
Contract, derived from the child's structured completion fields:
|
|
status ∈ {completed, interrupted, failed} — a structured failure
|
|
(failed=True / non-empty error) or an invalid terminal state
|
|
is "failed" even when a summary exists.
|
|
exit_reason ∈ {completed, max_iterations, interrupted, error} —
|
|
"max_iterations" only for genuine budget exhaustion
|
|
(completed=False with no failure fields), never for errors.
|
|
truncated == (exit_reason == "max_iterations").
|
|
|
|
* ``"completed"`` — normal finish. See #97655.
|
|
"""
|
|
child_progress_cb = getattr(child, "tool_progress_callback", None)
|
|
child_pool, leased_cred_id = _lease_child_credential(child)
|
|
# Heartbeat keeps the parent's _last_activity_ts moving so the gateway inactivity timeout doesn't fire while the
|
|
# child works; once the child looks stale (see _HEARTBEAT_STALE_CYCLES_*) it also ends await_child's wait.
|
|
heartbeat = _start_heartbeat(child, parent_agent, task_index)
|
|
# TUI/RPC registry entry (kill/pause/status by subagent_id); None for test
|
|
# doubles without a stable id. Unregistered in the finally block.
|
|
_subagent_id = _register_child(
|
|
child, parent_agent, goal, owner_session_id=owner_session_id, owner_transport=owner_transport,
|
|
owner_session_record=owner_session_record,
|
|
)
|
|
run = _ChildRun(child, parent_agent, task_index, goal, _subagent_id, child_progress_cb, heartbeat=heartbeat)
|
|
# Set when a timed-out Future still owns the child: closing it from this
|
|
# thread before the worker settles races the conversation's finally path.
|
|
_child_close_deferred = False
|
|
try:
|
|
heartbeat.start()
|
|
_safe_progress(child_progress_cb, "subagent.start", preview=goal)
|
|
run.seed_workspace()
|
|
result, failure_entry, _child_close_deferred = run.await_child()
|
|
if failure_entry is not None:
|
|
return failure_entry
|
|
|
|
schema = _validate_child_output_schema(child, result, task_index, run.child_task_id, run.relay_text)
|
|
_merge_late_steer(result, _subagent_id, child)
|
|
# Flush any remaining batched progress to gateway
|
|
if child_progress_cb and hasattr(child_progress_cb, "_flush"):
|
|
with _quiet("Progress callback flush failed: %s"):
|
|
child_progress_cb._flush()
|
|
|
|
duration = run.elapsed()
|
|
entry = _build_result_entry(child, result, task_index, duration, schema)
|
|
run.append_sibling_write_reminder(entry)
|
|
run.account_background_processes(entry)
|
|
run.emit_complete(result, entry, duration)
|
|
return run.attach_worktree(entry)
|
|
except Exception as exc:
|
|
# Close steer acceptance before any completion callback (see _merge_late_steer).
|
|
_late_pending_steer = run.close_steering()
|
|
logging.exception(f"[subagent-{task_index}] failed")
|
|
# Entry status "error" (contract), progress event status "failed" (UI vocabulary).
|
|
return run.finish_failed(
|
|
_fabricated_entry(task_index, "error", str(exc), child, run.elapsed()), _late_pending_steer,
|
|
preview=str(exc), summary=str(exc), status="failed",
|
|
)
|
|
finally:
|
|
run.cleanup(heartbeat=heartbeat, child_pool=child_pool, leased_cred_id=leased_cred_id, close_deferred=_child_close_deferred)
|
|
|
|
|
|
def _build_children(
|
|
task_list: List[Dict[str, Any]], task_schemas: List[Optional[Dict[str, Any]]], creds: Dict[str, Any], *,
|
|
top_role: str, max_iterations: int, parent_agent, routing_cfg: Dict[str, Any],
|
|
live_deleg_id: Optional[str], live_writers: list, task_images: Optional[List[Optional[List[str]]]] = None,
|
|
) -> tuple[List[tuple], Optional[str]]:
|
|
"""Build every child on the main thread (construction is not thread-safe);
|
|
``(children, None)`` or ``([], error)`` on an explicit-pin preflight failure."""
|
|
from tools.delegation_live_log import wrap_progress_callback
|
|
from tools.delegation_output_schema import append_output_contract
|
|
overrides = {
|
|
"override_provider": creds["provider"], "override_base_url": creds["base_url"],
|
|
"override_api_key": creds["api_key"], "override_api_mode": creds["api_mode"],
|
|
"override_request_overrides": creds.get("request_overrides"),
|
|
"override_acp_command": creds.get("command"),
|
|
"override_acp_args": creds.get("args"),
|
|
"routing_cfg": routing_cfg,
|
|
}
|
|
children = []
|
|
for i, t in enumerate(task_list):
|
|
_task_schema = task_schemas[i] if i < len(task_schemas) else None
|
|
_child_context = t.get("context")
|
|
if _task_schema is not None:
|
|
_child_context = append_output_contract(_child_context, _task_schema)
|
|
try:
|
|
child = _build_child_preserving_parent_tools(
|
|
task_index=i, goal=t["goal"], context=_child_context,
|
|
toolsets=None, # always inherit the parent's toolsets
|
|
model=creds["model"], max_iterations=max_iterations, task_count=len(task_list),
|
|
parent_agent=parent_agent, role=_normalize_role(t.get("role") or top_role), **overrides,
|
|
)
|
|
except ValueError as exc:
|
|
return [], str(exc)
|
|
if _task_schema is not None:
|
|
with _quiet("Could not attach output schema to child %d", i):
|
|
child._delegate_output_schema = _task_schema
|
|
# Validated per-task images; absent on image-less tasks, which keep the text-only goal turn.
|
|
_t_images = task_images[i] if task_images and i < len(task_images) else None
|
|
if _t_images:
|
|
with _quiet("Could not attach images to child %d", i):
|
|
child._delegate_images = _t_images
|
|
# Tee progress events into the live transcript (wrapper keeps the
|
|
# _flush contract and swallows writer failures).
|
|
_writer = live_writers[i] if i < len(live_writers) else None
|
|
if _writer is not None:
|
|
child.tool_progress_callback = wrap_progress_callback(getattr(child, "tool_progress_callback", None), _writer)
|
|
child._live_transcript_path = str(_writer.path)
|
|
if live_deleg_id:
|
|
setattr(child, "_delegation_id", live_deleg_id)
|
|
_ident_ref = getattr(child, "_progress_identity_ref", None)
|
|
if isinstance(_ident_ref, dict):
|
|
_ident_ref["delegation_id"] = live_deleg_id
|
|
children.append((i, t, child))
|
|
return children, None
|
|
|
|
|
|
def _oneshot_spawn_budget(parent_agent: Any, requested: int) -> Optional[str]:
|
|
"""Charge *requested* children against the finite one-shot session's total (delegation.oneshot_max_children);
|
|
the error text tells the model to do the work inline. Interactive and gateway sessions are never charged."""
|
|
from agent.oneshot_footprint import is_single_query_session
|
|
if not is_single_query_session():
|
|
return None
|
|
cap = _get_oneshot_max_children()
|
|
if cap <= 0:
|
|
return None
|
|
spent = getattr(parent_agent, "_oneshot_children_spawned", 0)
|
|
if spent + requested > cap:
|
|
return (
|
|
f"Delegation budget for this one-shot run is exhausted ({spent}/{cap} subagents used; "
|
|
f"delegation.oneshot_max_children). Do the remaining work yourself in this session — reviewing "
|
|
f"your own diff and running the tests inline is expected here, not a delegated review."
|
|
)
|
|
parent_agent._oneshot_children_spawned = spent + requested
|
|
return None
|
|
|
|
|
|
def delegate_task(
|
|
goal: Optional[str] = None, context: Optional[str] = None, tasks: Optional[List[Dict[str, Any]]] = None,
|
|
max_iterations: Optional[int] = None, role: Optional[str] = None, background: Optional[bool] = None,
|
|
output_schema: Optional[Dict[str, Any]] = None, images: Optional[List[str]] = None, action: Optional[str] = None,
|
|
subagent_id: Optional[str] = None, message: Optional[str] = None, parent_agent=None,
|
|
credentials_cfg: Optional[Dict[str, Any]] = None,
|
|
) -> str:
|
|
"""Spawn child agents (single ``goal`` or ``tasks=[...]`` batch) or control running ones. ``action``
|
|
list/steer/stop run synchronously and bypass the pause gate, depth limit and async dispatch. ``role`` is legacy
|
|
(per-task beats top-level; capability is depth-derived). Returns JSON with one results entry per task, or a
|
|
dispatch handle when running in the background."""
|
|
if parent_agent is None:
|
|
return tool_error("delegate_task requires a parent agent context.")
|
|
|
|
normalized_action = (action or "").strip().lower()
|
|
if normalized_action in _CONTROL_ACTIONS:
|
|
return _handle_control_action(normalized_action, subagent_id, message, parent_agent)
|
|
if normalized_action and normalized_action != "spawn":
|
|
return tool_error(f"Unknown action '{action}'. Use spawn (default), list, steer, or stop.")
|
|
|
|
# Operator kill switch (TUI / delegation.pause RPC): blocks NEW spawns only.
|
|
if is_spawn_paused():
|
|
return tool_error(
|
|
"Delegation spawning is paused. Clear the pause via the TUI "
|
|
"(`p` in /agents) or the `delegation.pause` RPC before retrying."
|
|
)
|
|
|
|
top_role = _normalize_role(role)
|
|
# background applies to single tasks AND batches: a batch is ONE async unit
|
|
# that joins on every child and re-enters as a single consolidated message.
|
|
background = is_truthy_value(background, default=False) if background is not None else False
|
|
|
|
depth = getattr(parent_agent, "_delegate_depth", 0)
|
|
max_spawn = _get_max_spawn_depth()
|
|
if depth >= max_spawn:
|
|
return tool_error(
|
|
f"Delegation depth limit reached (depth={depth}, max_spawn_depth={max_spawn}). Raise "
|
|
f"delegation.max_spawn_depth in config.yaml if deeper nesting is required (no hard ceiling, but each level "
|
|
f"multiplies API cost)."
|
|
)
|
|
|
|
cfg = _load_config()
|
|
default_max_iter = cfg.get("max_iterations", DEFAULT_MAX_ITERATIONS)
|
|
# Caller-supplied max_iterations is ignored: the config value is authoritative
|
|
# so budgets stay predictable (kwarg kept for internal callers/tests).
|
|
if max_iterations is not None and max_iterations != default_max_iter:
|
|
logger.debug(
|
|
"delegate_task: ignoring caller-supplied max_iterations=%s; using delegation.max_iterations=%s from config",
|
|
max_iterations, default_max_iter,
|
|
)
|
|
# credentials_cfg (internal callers only, e.g. /review → auxiliary.review) is
|
|
# a per-call routing owner shaped like the delegation config section. Keep
|
|
# the route and its fallback policy together through child construction.
|
|
routing_cfg = credentials_cfg if credentials_cfg is not None else cfg
|
|
try:
|
|
creds = _resolve_delegation_credentials(routing_cfg, parent_agent)
|
|
except ValueError as exc:
|
|
# Explicit-pin preflight failures (e.g. pinned delegation.command missing from PATH) refuse the
|
|
# spawn loudly (#80450).
|
|
return tool_error(str(exc))
|
|
max_children = _get_max_concurrent_children()
|
|
task_list, err = _normalize_task_list(goal, context, tasks, output_schema, top_role, max_children)
|
|
if not err:
|
|
task_schemas, err = _coerce_task_schemas(task_list, output_schema)
|
|
if not err:
|
|
task_images, err = _coerce_task_images(task_list, images)
|
|
if err:
|
|
return tool_error(err)
|
|
err = _oneshot_spawn_budget(parent_agent, len(task_list))
|
|
if err:
|
|
return tool_error(err)
|
|
|
|
overall_start = time.monotonic()
|
|
# Live transcripts: cache/delegation/live/<id>/task-<n>.log per task, a side channel with zero effect on message
|
|
# content or prompt caching. Best-effort: on failure live_paths is empty and delegation proceeds.
|
|
from tools.delegation_live_log import create_live_transcripts
|
|
live_deleg_id, live_writers, live_paths = create_live_transcripts(
|
|
task_list, context, model=creds.get("model"), provider=creds.get("provider")
|
|
)
|
|
_announce_batch(parent_agent, len(task_list), live_deleg_id)
|
|
origin = _capture_origin()
|
|
|
|
children, err = _build_children(
|
|
task_list, task_schemas, creds, top_role=top_role, max_iterations=default_max_iter, parent_agent=parent_agent,
|
|
routing_cfg=routing_cfg, live_deleg_id=live_deleg_id, live_writers=live_writers, task_images=task_images,
|
|
)
|
|
if err:
|
|
return tool_error(err)
|
|
batch = _Batch(
|
|
task_list, children, parent_agent, creds, context, top_role, max_children,
|
|
live_deleg_id, live_writers, live_paths, *origin, overall_start,
|
|
)
|
|
return _run_batch(batch, background)
|
|
|
|
|
|
# ── OpenAI function-calling schema ──────────────────────────────────────────
|
|
|
|
def _build_top_level_description(*, independent_completions=None) -> str:
|
|
"""delegate_task description: ONLY guidance stated nowhere else in the schema
|
|
(limits live in the 'tasks' parameter description, rebuilt per get_definitions())."""
|
|
try:
|
|
orchestration_available = _get_max_spawn_depth() >= 2 and _get_orchestrator_enabled()
|
|
except Exception:
|
|
orchestration_available = False
|
|
# Mention recursion only where it's actually available. send_message is deliberately not named (gateway-internal
|
|
# vocabulary); model_tools session-filters the list to tools the session has.
|
|
if orchestration_available:
|
|
restrictions_rule = (
|
|
"- Children cannot call clarify, memory, or cronjob.\n"
|
|
f"- Children can themselves delegate while depth remains (max_spawn_depth={_get_max_spawn_depth()}); the "
|
|
"runtime derives this from depth automatically.\n"
|
|
)
|
|
else:
|
|
restrictions_rule = "- Children cannot call delegate_task, clarify, memory, or cronjob.\n"
|
|
from tools.delegate_tool_config import _get_independent_completions
|
|
|
|
if independent_completions is None:
|
|
independent_completions = _get_independent_completions()
|
|
delivery = (
|
|
"each ungrouped task / `group` returns on its own"
|
|
if independent_completions else "one message per call"
|
|
)
|
|
return _DESCRIPTION_HEAD.format(delivery=delivery) + restrictions_rule + _DESCRIPTION_TAIL
|
|
|
|
_DESCRIPTION_HEAD = (
|
|
"Spawn subagents in isolated contexts; each gets its own conversation, terminal session, and toolset, and only its "
|
|
"final summary returns to you. Pass every task in `tasks` — one entry spawns one subagent, several run in parallel "
|
|
"(limit in the tasks description).\n\n"
|
|
"Sessions without a later-result consumer (including one-shot CLI and cron) join parallel children "
|
|
"and return results in this tool call. "
|
|
"Otherwise runs in the background: dispatch returns live transcript paths and results re-enter "
|
|
"as a new message when subagents finish ({delivery}). Background results are delivered only "
|
|
"BETWEEN your turns: finish whatever does not depend on them, then give a one-line status and END YOUR TURN. Never "
|
|
"wait or poll on transcripts, artifact files, or CI for a child. "
|
|
"While children run, `action` (list/steer/stop) controls them live.\n\n"
|
|
"USE FOR: reasoning-heavy subtasks, work that would flood your context, or independent parallel workstreams.\n"
|
|
"DO NOT USE FOR (use these instead):\n"
|
|
"- Mechanical multi-step work with no reasoning needed -> execute_code\n"
|
|
"- A single tool call -> call the tool directly\n"
|
|
"- Tasks needing user interaction -> subagents cannot ask questions\n"
|
|
"- Durable work that must survive this session -> cronjob or terminal(background=True, notify=True); /stop, /new, "
|
|
"or process exit halts running subagents (whole tree); each returns an 'interrupted' completion with partial output.\n\n"
|
|
"RULES:\n"
|
|
"- Children know nothing of this conversation: pass everything needed via 'context', including any required "
|
|
"output language, tone, or style (e.g. \"respond in Chinese\").\n"
|
|
"- Child summaries are SELF-REPORTS, not verified facts: a child claiming \"uploaded successfully\" or "
|
|
"\"file written\" may be wrong. For external side effects (uploads, remote writes, publishing), require a "
|
|
"verifiable handle (URL, ID, absolute path) and verify it yourself before telling the user the operation "
|
|
"succeeded.\n"
|
|
"- Children cannot close tracked work: a child asked to close it returns findings instead; "
|
|
"the parent applies the transition.\n"
|
|
)
|
|
_DESCRIPTION_TAIL = (
|
|
"- Children inherit the parent model unless pinned via delegation.provider / delegation.model in config.yaml."
|
|
)
|
|
|
|
def _build_tasks_param_description() -> str:
|
|
"""Compose the 'tasks' parameter description with current concurrency limit."""
|
|
try:
|
|
max_children = _get_max_concurrent_children()
|
|
except Exception:
|
|
max_children = _DEFAULT_MAX_CONCURRENT_CHILDREN
|
|
return (
|
|
f"The task(s), up to {max_children} in parallel for this user (set "
|
|
"via delegation.max_concurrent_children). Each entry spawns one "
|
|
"subagent with isolated context and terminal session; a single task "
|
|
"is a one-entry array. Required when spawning."
|
|
)
|
|
|
|
def _build_dynamic_schema_overrides() -> dict:
|
|
"""Per-call schema overrides (ToolEntry.dynamic_schema_overrides): every
|
|
get_definitions() pass rewrites the descriptions to the user's actual limits."""
|
|
from tools.delegate_tool_config import _get_independent_completions
|
|
|
|
independent_completions = _get_independent_completions()
|
|
overrides_params = {**DELEGATE_TASK_SCHEMA["parameters"]}
|
|
# Copy properties so the static schema dict is never mutated.
|
|
overrides_params["properties"] = {k: dict(v) for k, v in DELEGATE_TASK_SCHEMA["parameters"]["properties"].items()}
|
|
overrides_params["properties"]["tasks"]["description"] = _build_tasks_param_description()
|
|
|
|
if not independent_completions:
|
|
tasks = overrides_params["properties"]["tasks"]
|
|
tasks["items"] = {**tasks["items"], "properties": {
|
|
k: v for k, v in tasks["items"]["properties"].items() if k != "group"
|
|
}}
|
|
|
|
return {
|
|
"description": _build_top_level_description(independent_completions=independent_completions),
|
|
"parameters": overrides_params,
|
|
}
|
|
|
|
def _p(type_: str, description: str, **extra) -> dict:
|
|
return {"type": type_, **extra, "description": description}
|
|
|
|
DELEGATE_TASK_SCHEMA = {
|
|
"name": "delegate_task",
|
|
# description / tasks.description are placeholders: the real text is built per get_definitions() call by
|
|
# _build_dynamic_schema_overrides() so the model sees the user's actual max_concurrent_children / max_spawn_depth.
|
|
# Lazy (not at import) so cli.CLI_CONFIG isn't forced to load before the test conftest redirects HERMES_HOME.
|
|
"description": (
|
|
"Spawn one or more subagents in isolated contexts. "
|
|
"Description is rebuilt at every get_definitions() call to reflect the user's current delegation limits."
|
|
),
|
|
"parameters": {
|
|
"type": "object",
|
|
"properties": {
|
|
# The handler also accepts the legacy single-goal shape (top-level `goal`/`context`/`output_schema`),
|
|
# wrapped into a one-entry batch at dispatch, and a per-task `role` (legacy, ignored: capability is
|
|
# depth-derived). Both unadvertised on purpose (old transcripts only); do not re-add. No maxItems — the
|
|
# runtime limit (delegation.max_concurrent_children) is enforced with a clear error in delegate_task().
|
|
"tasks": {
|
|
"type": "array",
|
|
"minItems": 1,
|
|
"items": {
|
|
"type": "object",
|
|
"properties": {
|
|
"goal": _p(
|
|
"string",
|
|
"What this subagent should accomplish. Be specific and self-contained — it knows "
|
|
"nothing about your conversation history.",
|
|
),
|
|
"context": _p(
|
|
"string",
|
|
"Background THIS child needs: file paths, error messages, constraints. Each child "
|
|
"sees only its own context — repeat shared background in every task that needs it.",
|
|
),
|
|
"output_schema": _p(
|
|
"object",
|
|
"Optional JSON Schema this child's final answer must validate against (told to the "
|
|
"child up front; parent validates with one bounded correction retry; result gains "
|
|
"schema_valid, plus schema_errors on failure — the child's raw text is still returned "
|
|
"as summary, never discarded). Keep it forgiving — require only fields you will read.",
|
|
),
|
|
"images": _p(
|
|
"array",
|
|
"Optional images this child must SEE (max 8): local file paths or http(s) URLs — e.g. a "
|
|
"screenshot the user sent, a design mock, a chart. Vision-capable children receive the "
|
|
"pixels on their first turn; non-vision children get path hints for vision_analyze. Text "
|
|
"files do NOT belong here — put paths in 'context' instead.",
|
|
items={"type": "string"},
|
|
),
|
|
"group": _p(
|
|
"string",
|
|
"Optional result-delivery bucket within this call (only when delegation.independent_completions "
|
|
"is enabled; otherwise the whole call returns as one message). Tasks sharing a group return "
|
|
"together in ONE message; ungrouped tasks return individually as each finishes. This does not "
|
|
"order execution; if B needs A's output, dispatch B after A returns.",
|
|
),
|
|
},
|
|
"required": ["goal"],
|
|
},
|
|
"description": "(rebuilt at get_definitions() time)",
|
|
},
|
|
# `background` (bool) is also accepted — DEPRECATED, ignored: top-level
|
|
# delegations always run in the background. Unadvertised; do not re-add.
|
|
"action": _p(
|
|
"string",
|
|
"Default 'spawn'. Live control of running children: "
|
|
"'list' = ids/goals/status/transcripts; 'steer' = queue "
|
|
"course-correction text into one child (subagent_id + "
|
|
"message) without stopping it; 'stop' = end one child "
|
|
"early (subagent_id; partial result still returns). "
|
|
"Control actions return immediately; goal/tasks are ignored unless spawning.",
|
|
enum=["spawn", "list", "steer", "stop"],
|
|
),
|
|
"subagent_id": _p("string", "Target for action='steer'/'stop' (ids from the spawn response or action='list')."),
|
|
"message": _p(
|
|
"string",
|
|
"For action='steer': the course correction, appended to "
|
|
"the child's next tool result mid-run. Be directive and specific.",
|
|
),
|
|
},
|
|
"required": [],
|
|
},
|
|
}
|
|
|
|
|
|
# --- Registry ---
|
|
from tools.registry import registry, tool_error
|
|
|
|
def _model_background_value(args: dict, parent_agent=None) -> bool:
|
|
"""Background flag for the MODEL-facing dispatch path (registry fallback). Top-level delegations always run in the
|
|
background — the model does not choose — for single tasks and fan-out batches alike (one async unit, one
|
|
consolidated result); an orchestrator subagent (depth > 0) is the exception since it needs its workers' results
|
|
within its own turn. The live path is ``run_agent._dispatch_delegate_task``; this mirrors it for the rare case
|
|
the intercept is bypassed. Direct Python callers keep the synchronous default."""
|
|
return not getattr(parent_agent, "_delegate_depth", 0) > 0
|
|
|
|
_MODEL_HIDDEN_TASK_FIELDS = {"acp_command", "acp_args"}
|
|
|
|
def _strip_model_hidden_task_fields(tasks: Any) -> Any:
|
|
"""Drop trusted-config-only task fields from model-supplied tasks (same list object back when nothing changed)."""
|
|
if not isinstance(tasks, list) or not any(isinstance(t, dict) and _MODEL_HIDDEN_TASK_FIELDS & t.keys() for t in tasks):
|
|
return tasks
|
|
return [{k: v for k, v in t.items() if k not in _MODEL_HIDDEN_TASK_FIELDS} if isinstance(t, dict) else t for t in tasks]
|
|
|
|
|
|
registry.register(
|
|
name="delegate_task",
|
|
toolset="delegation",
|
|
schema=DELEGATE_TASK_SCHEMA,
|
|
handler=lambda args, **kw: delegate_task(
|
|
goal=args.get("goal"), context=args.get("context"), tasks=_strip_model_hidden_task_fields(args.get("tasks")),
|
|
max_iterations=args.get("max_iterations"), role=args.get("role"),
|
|
background=_model_background_value(args, kw.get("parent_agent")), output_schema=args.get("output_schema"),
|
|
images=args.get("images"), action=args.get("action"), subagent_id=args.get("subagent_id"), message=args.get("message"),
|
|
parent_agent=kw.get("parent_agent"),
|
|
),
|
|
check_fn=check_delegate_requirements,
|
|
emoji="🔀",
|
|
dynamic_schema_overrides=_build_dynamic_schema_overrides,
|
|
)
|
|
|
|
|
|
# ---- BEGIN PLUGIN-COMPAT (revert-scheduled; see COMPAT_MANIFEST.md) ----
|
|
# Names external plugins imported from this module before the Sep 2026 decomposition.
|
|
# Internal code MUST NOT use these (scripts/check_compat_pointers.py fails CI if it does).
|
|
# The whole block is removed by reverting the commit that added it.
|
|
from concurrent.futures import TimeoutError as FuturesTimeoutError # noqa: F401,E402
|
|
import contextvars # noqa: F401,E402
|
|
import enum # noqa: F401,E402
|
|
import json # noqa: F401,E402
|
|
import os # noqa: F401,E402
|
|
import re # noqa: F401,E402
|
|
import threading # noqa: F401,E402
|
|
from urllib.parse import urlsplit # noqa: F401,E402
|
|
from urllib.parse import urlunsplit # noqa: F401,E402
|
|
|
|
|
|
_PLUGIN_COMPAT_LAZY = {
|
|
'DEFAULT_CHILD_TIMEOUT': ('tools.delegate_tool_config', 'DEFAULT_CHILD_TIMEOUT'),
|
|
'DEFAULT_MAX_SUMMARY_CHARS': ('tools.delegate_tool_results', 'DEFAULT_MAX_SUMMARY_CHARS'),
|
|
'DEFAULT_TOOLSETS': ('tools.delegate_tool_toolsets', 'DEFAULT_TOOLSETS'),
|
|
'MAX_DEPTH': ('tools.delegate_tool_config', 'MAX_DEPTH'),
|
|
'TOOLSETS': ('toolsets', 'TOOLSETS'),
|
|
'base_url_hostname': ('utils', 'base_url_hostname'),
|
|
'file_state': ('tools', 'file_state'),
|
|
'request_hard_interrupt': ('agent.interrupt_compat', 'request_hard_interrupt'),
|
|
}
|
|
|
|
|
|
def __getattr__(name): # PEP 562 — lazy so no import cycles
|
|
target = _PLUGIN_COMPAT_LAZY.get(name)
|
|
if target is None:
|
|
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
|
import importlib
|
|
from hermes_cli.plugin_compat import warn_once
|
|
warn_once(__name__, name, *target)
|
|
return getattr(importlib.import_module(target[0]), target[1])
|
|
# ---- END PLUGIN-COMPAT ----
|