Commit Graph

262 Commits

Author SHA1 Message Date
Teknium
f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
Victor Kyriazakos
49ef015ca3 fix: join delegated work before finite chat exits 2026-09-10 05:02:41 -07:00
Teknium
aa83c6d614 fix: hide inactive grouping options from delegation schema 2026-09-08 13:38:18 -07:00
kshitijk4poor
1e24a8de39 refactor(delegation): resolve the child fallback chain through the canonical normalizer
_resolve_child_fallback_chain re-implemented hermes_cli.fallback_config's entry
validation to log per-index warnings, wrapped a pure function in try/except, and
carried two names (routing_cfg / fallback_cfg) for one argument. It now delegates
to get_fallback_chain() and keeps the single warning for "no usable routes";
the facade re-export is dropped (import from the defining module). Tests collapse
to the parametrized decision table plus the pin-derivation, real-config-loader,
review-ownership and activation invariants (22 -> 20 cases, 417 -> ~230 lines).
2026-09-08 02:26:05 +05:30
Ayush Nangia
c47bf78d68 fix(delegation): keep child routes and fallback policy together 2026-09-08 02:26:05 +05:30
Teknium
3c0d90e8ef feat(delegation): subagents hand background processes to the parent; leftovers are named, not trusted
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).

- process_manage(action="handoff", session_id, data="<purpose>"), children
  only: process_registry.transfer_ownership flips owner_task_id/task_id/
  session_key to the parent under the registry lock, so the completion is
  stamped with the parent's owner at exit, passes the parent's sa- filter,
  and is reaped by the parent, not the child. Cap 3 per child; an exited,
  foreign, or non-child request is a tool error. The purpose rides the
  event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
  notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
  orphaned_processes on the result before cleanup kills the leftovers; the
  parent's delegation block renders both.
2026-09-07 12:50:29 -07:00
Teknium
ae9cd7073e fix(delegation): keep the 'wait or poll' contract token in the tool description 2026-09-07 06:46:54 -07:00
Teknium
c89f3b8800 fix(delegation): one completion per call by default; queued units no longer stalled; tell the model results land between turns
Three orchestrator failures traced through the Sep 7 gpt-6-astra campaign sessions:

1. delegation.independent_completions (new, default false). #104299 made every
   ungrouped task its own completion message, so a 15-task call woke the
   orchestrator up to 15 times; one chain received 132 notices and answered
   130 of them with "already incorporated". A multi-task call now returns as
   ONE consolidated message unless the flag is on; `group` is inert until then.

2. Queued units were killed before they started. Units of one call share a
   pool slot but the executor was still sized by slots, so with 15 units live
   a new unit queued behind a full pool; the stale monitor's clock ran from
   dispatch, interrupted it at 450 s, and the child exited `interrupted 0.02s`
   when its thread finally came up (13 such lanes in one session). The
   executor now grows to the number of live units and the stall clock arms
   when the runner actually starts.

3. The tool text said "do not wait or poll — just continue" without saying
   that completions are delivered only BETWEEN turns. A model that never ends
   its turn (one 203-minute turn, 717 API calls) never received 40 finished
   results. Tool description, dispatch note and completion header now say to
   finish independent work, give a one-line status, and end the turn.
2026-09-07 06:46:54 -07:00
Teknium
fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium
a688e7d5ff docs: clarify when delegated results belong in one group 2026-09-07 01:23:34 -07:00
Teknium
dcdbc0093d fix(delegation): compression_threshold_tokens is opt-in (default 0); keep the value validation
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
  the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
  200K-400K caps within 5% of each other in cost once cache prefixes are intact,
  because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
  measured; at 500K a 1M child compacts roughly never.

What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
2026-09-06 10:36:08 -07:00
Teknium
b1403d521c Merge pull request #104168 from NousResearch/fix/subagent-cache-ttl-5m
fix(delegate): subagents never inherit the 1h prompt-cache tier (2x write price for retention they never use)
2026-09-06 07:17:43 -07:00
Teknium
028fe2c4c8 feat(delegation): per-task completion groups — ungrouped subagents return as they finish
A background delegate_task call used to be ONE async unit: the runner joined on
every child and a single consolidated message re-entered the conversation when
the SLOWEST finished. Fifteen independent PR reviews therefore waited on the
fifteenth before the parent could act on the first.

Each task now carries an optional `group`. `_units_of` partitions the call's
children into units — one per distinct group, one per ungrouped task — and each
unit is dispatched to the async registry on its own, so its results re-enter the
conversation as soon as THAT unit is done. Tasks that must be compared or merged
share a group and still return together.

Capacity is unchanged: every unit of one call joins the first unit's pool slot
(`slot_key` in `async_delegation._dispatch`), so splitting never consumes more of
`delegation.max_concurrent_children` than the call did. Unit ids suffix the
call's id (`deleg_xxxx-1`, `-2`, …) so live transcripts stay under one dir; the
completion block names the group and notes that sibling units report separately;
`active_task_count` counts a unit's own tasks.
2026-09-06 07:17:30 -07:00
Teknium
0edb6b928a fix(delegate): subagents never inherit the 1h prompt-cache tier
A delegated child copies the parent's prompt_caching.cache_ttl. The 1h tier
is priced for a person who steps away between turns (2x write vs 1.25x for
5m); a subagent calls every few seconds for minutes and is gone, so it paid
2x on every tool result and never collected the retention. Live measurement
(Sep 5, 40 concurrent Fable 5.1 children via OpenRouter, 1h markers on the
wire): cache writes billed at $20/M against $12.51/M for the same run at
5m, ~60% of a write-dominated bill.

_apply_child_cache_ttl runs right after child construction: 1h -> 5m,
5m stays, disabled stays disabled, parent untouched. Tests: unit (markers
on the wire drop ttl; parent's 1h layout still differs) and the real spawn
path with cache_ttl: 1h in a temp HERMES_HOME.
2026-09-06 02:20:49 -07:00
Teknium
ec4c1e0c98 fix(delegation): validate compression_threshold_tokens; state that it caps the trigger, not the payload
Independent review: a YAML `true` coerced to int 1 and gave every child a
one-token compression trigger; "200k" silently disabled the default cap.
Values are now validated: an int >= 16000 is used, 0/false/null disable on
purpose, anything else warns and falls back to the 200K default so a typo
never costs money. Config comment and docs say it caps the compaction
trigger, not a hard request-size limit.

Test: true / "200k" / 5 -> default; 0 / false / None -> disabled; valid ints
pass.
2026-09-05 06:29:28 -07:00
Teknium
40da0fd52f feat(delegation): subagents compress at an absolute context cap (delegation.compression_threshold_tokens, default 200K)
A delegate_task child inherits the compression threshold as a RATIO of the
model window. On a 1M-window model at the run's configured 0.85 that is an
850K-token trigger: in the 1,393-agent refactor run 1,373 of 1,375 children
never compressed once, 62% of all API calls carried >150K of context, and the
calls above 200K carried ~$10.9k of the $19.3k bill (58% of it cache WRITES,
i.e. re-sending a 300-800K prefix on every call). A sawtooth replay of the
logged calls with a 200K cap / 65K floor cuts context spend by ~49% (~$7k).

Children are brief-driven and disposable; they re-read their brief and the
files they touch, so a large window buys them little. New
delegation.compression_threshold_tokens (default 200000) is applied to the
child's ContextCompressor right after construction as the lower of it and
any global compression.threshold_tokens; the parent's own trigger is
untouched. 0 disables the subagent-specific cap. The compressor applies
threshold_tokens_cap on first window resolution, so this is byte-equivalent
to the user having set compression.threshold_tokens for the child.

Live through the real spawn path (_build_child_agent, real imports, temp
HERMES_HOME, 1M-window model): main child trigger 500,000 / branch 200,000;
parent 500,000 on both.

Tests (3): default caps a 1M child at 200K; the cap is the lower of the
delegation and global values and never raises a small-window child's
trigger; 0 disables and an already-resolved trigger is re-clamped.
Docs: delegation.md, configuration.md.
2026-09-05 01:07:09 -07:00
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
561b053f79 perf(agents): run per-child timers on one shared scheduler thread
A fan-out of N in-process subagents used to add one sleeping daemon
thread per delegated child (delegate heartbeat, 30s) and one or two per
active turn (durable turn-lease refresher; turn-liveness watchdog).  A
profiled session with ~130 children was carrying ~1000 threads.  All
of these timers now run on a single process-wide daemon thread.

- agent/periodic_scheduler.py (new): heap-ordered periodic scheduler on
  one Condition-driven daemon thread.  schedule(fn, interval) -> handle;
  handle.cancel(wait=) blocks for an in-flight run like the old join.
  A callback returning False stops itself; a raising callback is logged
  at debug and rescheduled, so one bad timer cannot kill the rest.
- tools/delegate_tool.py: _heartbeat_loop body -> _heartbeat_tick,
  scheduled at _HEARTBEAT_INTERVAL; stale-cycle closure state and
  idle/in-tool thresholds unchanged; cancel(wait=5) in finally where the
  stop-event + join(5) lived.
- run_agent.py: _refresh_durable_turn_lease body scheduled at
  _lease_refresh_interval; lease-lost / refresh-error interrupt paths
  and the stop-event fencing are unchanged; the join(timeout=1.0) is now
  cancel(wait=1.0) so the interrupt clear still runs after any in-flight
  tick.
- agent/turn_liveness.py: TurnLivenessWatchdog.make_thread/start ->
  schedule(); the poll body is _tick(), same sampling state machine.

Bench (evals/fanout_resource_bench.py, 30 children / 10 worktrees,
ok=30/30 both): peak threads 168 -> 132.  At peak the old tree held 30
"Thread-N (_heartbeat_loop)" threads; the new one holds zero plus one
"hermes-periodic-scheduler".
2026-09-03 02:44:24 -07:00
Teknium
798470930f refactor(delegate): tool description text as joined constants (byte-identical) 2026-09-02 19:50:30 -07:00
Teknium
defca3639a refactor(delegate): reflow comments/docstrings to 118 cols (word-identical, AST-identical) 2026-09-02 19:25:04 -07:00
Teknium
24ef289225 refactor(delegate): schema properties via _p() table (byte-identical schema) 2026-09-02 19:18:35 -07:00
Teknium
1283ccd855 refactor(delegate): docstring compaction by hand (every rule/why kept) 2026-09-02 19:16:07 -07:00
Teknium
aa31f83995 refactor(delegate): origin — packed AIAgent kwargs, override table in _build_children, tighter preamble 2026-09-02 19:11:48 -07:00
Teknium
eabec2ff50 refactor(delegate): AST-identical re-layout pass 2 2026-09-02 19:04:50 -07:00
Teknium
61f32ec330 refactor(delegate): drop unreferenced re-exports from origin (refs.py-verified) 2026-09-02 18:59:41 -07:00
Teknium
fe9c26ee5d refactor(delegate): _ChildRun dataclass replaces _WorktreeReporter/_ChildWorkspace/_ChildFailure plumbing 2026-09-02 18:27:59 -07:00
Teknium
378692be4f refactor(delegate): fold _construct_child_agent/_announce_child_spawn into _build_child_agent; _quiet best-effort sites 2026-09-02 18:04:43 -07:00
Teknium
b18228f377 refactor(delegate): AST-identical re-layout (pack signatures/call args, join split literals) 2026-09-02 17:59:28 -07:00
Teknium
901ba7db6f refactor(delegate): child_run best-effort blocks via _quiet; shared failure-entry tail 2026-09-02 17:54:08 -07:00
Teknium
64957b7e02 refactor(delegate): _Batch dataclass owns batch execution/dispatch; _quiet best-effort contextmanager 2026-09-02 17:50:16 -07:00
Teknium
d77505c007 refactor(delegate): extract toolset resolution + task validation modules, move runtime resolution to config 2026-09-02 17:44:45 -07:00
Teknium
3da2ff88e3 refactor(delegate): pack re-import block; single blank line between top-level defs (AST-identical) 2026-09-02 17:03:55 -07:00
Teknium
987bc82574 refactor(delegate): compact docstrings/comments by hand (every WHY kept); collapse orchestrator toolset branches 2026-09-02 17:03:05 -07:00
Teknium
016f31375d refactor(delegate): join wrapped logical lines that fit in 120 cols (AST-identical) 2026-09-02 16:59:26 -07:00
Teknium
e204008f10 refactor(delegate): _knob() unifies config>env>default int/float knobs; timeout diagnostic sections -> _diag_sizes/_diag_threads + attr table; drop dead _handle_child_wait_failure re-import 2026-09-02 16:57:22 -07:00
Teknium
338da30a45 refactor(delegate): _resolve_child_runtime returns AIAgent kwargs (drop _ChildRuntime; routing-filter table); dispatch reuses _detach_child/_signal_child_stop; control actions -> _CONTROL_OUTCOMES table; attribution lookup dedupe 2026-09-02 16:50:57 -07:00
Teknium
52419bda5d refactor(delegate): split delegate_task into _normalize_task_list/_coerce_task_schemas/_build_children/_execute_and_aggregate(_Batch); _build_child_agent -> _open_child_session_db/_construct_child_agent/_announce_child_spawn; _run_single_child -> _lease_child_credential/_await_child/_merge_late_steer 2026-09-02 16:46:39 -07:00
Teknium
05526b028a refactor(tools/delegate): split delegate_tool into child_run/config/dispatch/progress/registry/results; compact delegation helpers 2026-09-02 14:45:15 -07:00
Teknium
7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00
Teknium
a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium
bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
Teknium
9387bf929c fix(delegate): drain abandoned-worker transports FD-safely on child timeout
The #94248 native half. A delegation deadline abandons the child's daemon
worker while it is typically parked inside an in-flight OpenSSL read
(Codex Responses stream / httpx). PR #90889's deferred close (cherry-picked
here, authorship preserved) stops the timeout thread from closing the child
under the running future — but the deferred close only fires once the worker
unwinds, and a worker blocked in ssl.read never unwinds on its own: the
cooperative interrupt cannot reach a thread inside OpenSSL, so the child's
SessionDB, httpx pools, and subprocesses stayed pinned until process exit,
and any path that still hard-closed the transport released FDs under a live
SSL BIO (the #29507/#67142/#70773 native-corruption family; SIGSEGV 17-72ms
after "Subagent N timed out" on macOS arm64).

Fix — bounded drain after deferral:
- AIAgent._drain_transports_after_abandonment(): shutdown()-only sweep of
  the shared client's pooled sockets (force_close_tcp_sockets — FD release
  stays with the owning worker), abort+poison of the cached per-request
  openai/anthropic wire clients, Codex app-server request_interrupt(), and
  the inline _active_request_abort hook. Never client.close(), never
  socket.close().
- delegate timeout path: after registering the deferred-close callback,
  run one immediate drain plus one 5s re-sweep (covers a connection opened
  between the interrupt and the first sweep). The settled read (EOF/EPIPE)
  lets the worker unwind, which triggers the deferred close on the worker's
  own thread — the only safe FD-release boundary. A worker that still never
  settles retains its resources rather than risking a cross-thread close.

Live repro (Linux, real TLS server subprocess + real httpx client blocked
in OpenSSL read at the deadline + real SessionDB): before — child.close()
ran on the timeout thread with in_flight_ssl_read=True (client FDs released
under the live read; #94736 self-heal WARNING fired on the worker's unwind
flush); after — drain settles the read in ~1ms, worker unwinds, close runs
on the worker thread with in_flight_ssl_read=False.

Not live-tested on macOS arm64 (no macOS runner); the fix is
platform-neutral teardown ordering proven on Linux.

Closes #94248
2026-09-01 12:07:52 -07:00
Leandro Piccione
aa1d22670e fix(delegate): defer timed-out child teardown 2026-09-01 12:07:52 -07:00
kshitijk4poor
db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
itskaism
5ce8f71553 fix(delegation): report schema-invalid child results as failed, not completed
A delegate_task child dispatched with an output_schema whose final answer
still violates the schema after the one bounded retry (including the
common empty {} fallback) was reported status="completed" with a ✓ in
the batch report. Since the structured-output feature landed (d6ee58b58),
the result entry does carry schema_valid=false + schema_errors on
failure, but the status logic in _run_single_child only checked for a
non-empty summary and never consulted the validation outcome — so
consumers that read only status (orchestrators, the batch ✓/✗ icon,
subagent lifecycle state mapping) accepted a contract-violating verdict
as success.

Fix: in the status derivation, treat _schema_valid is False as a
failure ("failed"), between the interrupted and summary checks. The
failed entry names the schema violation in its error field instead of
the generic "Subagent did not produce a response.", and schema_errors
keep propagating verbatim. _schema_valid stays None on schema-less
delegations, so their entries remain byte-identical (wire-shape
pinning), and schema_valid=true children are untouched. Covers both
the single-goal and batch paths, which share _run_single_child.

Regression tests: schema-failing final ({} after retry) is failed with
a schema-specific error and the invalid text still in summary; retry-
exception path is failed; schema-valid and schema-less paths pinned
unchanged.
2026-08-31 01:02:42 -07:00
itskaism
1b6ea1a2c2 fix(delegation): report failed children as failed, not completed
A subagent whose loop gave up on a structured failure (e.g. "API call
failed after 3 retries: HTTP 524") returns that error message as
final_response together with completed=False / failed=True /
failure_reason. _run_single_child derived the batch-entry status from
the summary alone (`elif summary and not _empty_sentinel: status =
"completed"`), so the non-empty error text made the batch report show
the task as "✓ status=completed" — the `failed` flag was never
consulted anywhere in delegate_tool.py. Only the "(empty)" sentinel was
mapped to failed.

Fix, at the single status-determination choke point both the single-task
and batch paths share:

- `failed=True` on the child result now wins over a non-empty summary:
  status = "failed".
- The child's classified failure_reason (rate_limit / billing /
  server_error / ...) is propagated onto the batch entry so the parent
  can tell a quota wall from a real task error without parsing prose.
- exit_reason for a structured failure is "error" instead of falling
  through to "max_iterations" (which also wrongly set truncated=True).

Successful children (completed=True, no failed flag) are untouched —
covered by an explicit control test alongside the regression test,
which is red on the old code and green with the fix.
2026-08-30 22:19:10 -07:00