Commit Graph

124 Commits

Author SHA1 Message Date
Teknium
93af3db01d fix: checkpoint Kanban completion before tool access expires
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.

Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.

Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
2026-09-07 08:28:43 -07:00
Teknium
903b9bf187 fix(delegate): nested orchestrators get their workers' results back — no 420 s deadline on delegate_task, summary budget uses the current prompt not the session sum
Two defects in the same path, both measured on the 1,393-agent refactor run.

1. A nested orchestrator (depth > 0) runs delegate_task synchronously by
   design: it needs its workers' results inside its own turn. But the
   sequential tool runner put every tool call under the generic 420 s
   deadline, and delegate_task was not exempt, so every batch longer than
   seven minutes returned "timed out after 420.0s" while the children kept
   running as orphans. 332 such timeouts in 234 orchestrator sessions; only
   89 nested delegate_task calls in the whole run ever returned a real result.
   The orchestrators then spent 388 h of wall time polling: 1,526 reads of
   the live transcript files, 551 list actions, 242 h of explicit sleep,
   about $4k of API turns. delegate_task is now exempt from the sequential
   deadline (the batch owns its liveness: per-child heartbeats, the stale
   monitor, delegation.child_timeout_seconds).

   Live A/B, depth-1 orchestrator dispatching a 75 s leaf with the deadline
   set to 40 s (glm-5.3-flash via Nous): main -> "Error executing tool
   'delegate_task': timed out after 40.0s", leaf result lost; branch ->
   orchestrator blocked 161 s and returned the leaf's LEAF_DONE_MARKER.

2. _parent_summary_char_budget computed the parent's context headroom from
   session_prompt_tokens, which is the running SUM of prompt tokens over
   every API call in the session. After a few hundred calls it exceeds any
   window, headroom goes negative, and every child summary collapses to the
   2,000-char floor with the full text spilled to disk. All 1,393 child
   summaries in the run were truncated this way; the orchestrator planned
   from stubs. The budget now reads the last call's prompt_tokens from
   _last_turn_usage.

Tests: delegate_task is in the exempt set and the set is narrow; budget for
a long-lived parent equals the budget for a fresh parent with the same
current prompt, and exceeds the floor.
2026-09-05 00:10:26 -07:00
Teknium
b92308b1d2 simplify(compat): tools-A — repoint 4 stale docstring references (tools.approval.*, tools.transcription_tools.*) to the defining modules 2026-09-03 14:37:14 -07:00
Teknium
14791b4d4e simplify(compat): approval — drop 43 facade re-exports + _command_detection_variants late-bind seam, repoint 30 callers + 46 test files
tools/approval.py no longer re-exports sibling names (approval_context/prompt/floors/detection/
human_wait/smart/gateway_wait); it imports only what it uses. Siblings reference sibling-defined
names directly (module-attribute reads on tools.approval_context so patching the defining module
still works); only facade-owned state (_lock, _gateway_queues, _permanent_approved, _denied,
_denial_breaker_addendum, _gateway_notify_cb) is still read back through tools.approval.
approval_detection calls its own _command_detection_variants instead of late-binding through the facade.
2026-09-03 13:49:57 -07:00
Teknium
e43402fd83 simplify(compat): file_operations/file_tools — drop 47 re-exports/aliases, repoint 6 callers + 17 tests 2026-09-03 13:11:08 -07:00
Teknium
b818085298 simplify(compat): doctor/status — drop 13 re-exports + the doctor_* globals() facade (97 names), repoint 6 callers / 13 tests 2026-09-03 13:10:52 -07:00
Teknium
d98c48b410 review-fix(comments): restore lost rationale in the 7 files the sweep had to skip (docstring/comment-only) 2026-09-03 10:40:25 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
2252200b2a refactor(agent/tool_executor): contextlib.suppress for best-effort blocks; _preview helper; alias _pairing_tool_call_id 2026-09-02 19:29:07 -07:00
Teknium
a64889ef3d refactor(agent/tool_executor,tool_dispatch_helpers): compact spinner/dispatch helpers, planner closures, envelope helpers 2026-09-02 19:25:43 -07:00
Teknium
8efd6a682a refactor(agent/tool_executor): fold observe into _commit_tool_result; _ToolOutcome carries its ref; merge sequential abandon branches 2026-09-02 19:22:02 -07:00
Teknium
3ca2955138 refactor(agent/tool_executor): concurrent workers index parsed_calls directly; compact gate/auth helpers 2026-09-02 19:17:45 -07:00
Teknium
856d778df0 refactor(agent/tool_executor): flatten tool_search unwrap; compact docstrings/comments by hand 2026-09-02 19:11:43 -07:00
Teknium
f974353253 refactor(agent/tool_executor): _safe_callback guard replaces 5 try/except-log callback blocks 2026-09-02 19:08:27 -07:00
Teknium
15750e88ce refactor(agent/tool_executor): route begin/complete/middleware kwargs through _ToolCallRef; merge batch abandon branches 2026-09-02 19:04:35 -07:00
Jeffrey Quesnelle
48c0c3a873 Merge pull request #96633 from afourniernv/chore/relay-0.8.0
fix(relay): upgrade to 0.8.3
2026-09-02 22:03:27 -04:00
Teknium
12aebbfcbc refactor(agent/tool_executor): _ToolCallRef identity carrier + emit_post funnel replace 9 hand-rolled terminal-hook blocks 2026-09-02 18:59:32 -07:00
Teknium
b66776d594 refactor(agent/tool_executor): split executor god functions into phase helpers; unify observe/commit and start-gate handling 2026-09-02 18:49:11 -07:00
Teknium
3f93fdfc95 refactor(tool_executor): split execute_tool_calls_concurrent into _ConcurrentBatch + shared result helpers
Behavior-neutral extraction of the 686-LOC concurrent executor and the 429-LOC
sequential executor into focused helpers:

- _ConcurrentBatch (run_worker / submit_all / await_completion / run) and
  _StartOrderGate replace the nested closures + nonlocal counters; _ToolOutcome
  replaces the 7-tuple result slots; _ParsedCall/_parse_tool_call replaces the
  6-tuple parsed-call rows in both executors (-> 2 sites).
- _append_skipped_tool_results unifies the four cancelled/skipped-result loops
  (concurrent pre-flight, sequential pre-tool interrupt, sequential
  KeyboardInterrupt, sequential post-tool interrupt) -> 4 sites; absorbs
  _append_cancelled_tool_results.
- _observe_tool_result / _commit_tool_result / _finalize_tool_batch /
  _print_tool_completed / _tool_progress_enabled unify the post-execution
  guardrail-observe -> append+flush -> tool.completed -> budget -> /steer tail
  shared by the concurrent, sequential and segmented paths (-> 2-3 sites each).
- _unfinished_tool_result, _blocked_tool_result, _abandoned_sequential_result
  collapse the duplicated synthesize-result + terminal post_tool_call blocks.
- _registered_tool_worker / _interrupt_worker_tids unify worker tid tracking and
  interrupt fan-out between the two middleware runners (-> 2 / 3 sites).
- _run_with_activity_heartbeat extracts the heartbeat thread wrapper.
- _resolve_sequential_dispatch + _SequentialDispatch turn the 5-branch
  inline/delegate/context-engine/memory/registry if/elif into a resolver with
  per-branch spinner/error/KeyboardInterrupt policy, preserving branch order.
- _cancelled_tool_result and _managed_values inlined (single caller each).

Public signatures (execute_tool_calls_concurrent/sequential/segmented, both
middleware runners, every symbol imported by run_agent.py/agent/tests) are
unchanged; middleware/hook/persistence/progress-callback order is identical.
2026-09-02 16:51:42 -07:00
Teknium
4dcbbc84ab refactor(agent): table-driven inline tool executors shared by both tool paths; dedupe tool_executor
- agent/inline_tool_executors.py: INLINE_TOOL_EXECUTORS dispatch table (13 agent-level
  tools) replaces two drifted if/elif chains (invoke_tool + execute_tool_calls_sequential);
  resolve_invoke_tool_executor preserves the concurrent path's historical precedence.
- tool_hook_ids / emit_terminal_post_tool_call: single owners for hook identity kwargs
  and the terminal post_tool_call emit (was 8 + 3 hand-copied sites).
- tool_executor: shared quiet-spinner start/stop, tool-search unwrap + alias mapping,
  tool-result finalization tail and completion-callback fan-out between the concurrent
  and sequential executors; quiet/default registry branches merged.
2939 -> 2327 LOC.
2026-09-02 13:29:31 -07:00
DragonnZhang
9514d354ca fix(tool-search): validate deferred tool_call arguments against the concrete schema before dispatch
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.

Fixes #73175

Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
2026-09-01 23:30:33 -07:00
Alex Fournier
c3e86f7055 Merge branch main into chore/relay-0.8.0
Signed-off-by: Alex Fournier <afournier@nvidia.com>

# Conflicts:
#	uv.lock
2026-08-30 19:28:08 -04:00
Teknium
e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium
217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
Alex Fournier
31e7237adb fix(relay): upgrade to 0.8.1
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-28 08:28:21 -04:00
kshitij
f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30
Teknium
a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
kshitij
4341cf1df7 fix(approval): clamp approvals.timeout at the config-read chokepoint (#83220)
Widen the salvaged gate-site clamp to the bug class. The gate min() from
the contributor PR capped the gate bound at 360s, which would break the
#79719 contract (gate must extend while a legitimate >360s approval
prompt is answerable) and left the sibling overflow sites live: the CLI
prompt thread.join, the gateway poll deadline, and human_wait_ceiling
all consume the same config value.

Clamp once in _get_approval_timeout() via agent.deadline.MAX_SAFE_TIMEOUT_S
(1 year - semantically unbounded, platform-safe). The gate keeps its
approvals.timeout-tracking behavior above 360s; 7 regression tests pin
lock-acquire/thread-join safety and the gate-extension contract.

Bilateral E2E: with approvals.timeout=1e20 in a real config.yaml, main
crashes every consumer with OverflowError; this branch survives all 5
probes.
2026-08-24 16:11:03 +05:30
RelaxJonh
0dd0f6e648 fix(agent): clamp authorization gate lock timeout to prevent OverflowError on macOS
The _ConcurrentToolAuthorizationGate uses threading.Lock.acquire(timeout=...)
where timeout comes from human_wait_ceiling() (approvals.timeout + 60s).
When approvals.timeout is set very large (e.g. 999999999999 to effectively
disable timeouts), this overflows macOS timespec and raises:
  OverflowError: timestamp out of range for platform time_t

The gate only needs to serialize parallel dispatch — it should never wait
longer than the wedged-holder bound (_AUTHORIZATION_GATE_LOCK_TIMEOUT_S =
360s). Clamp the return value with min() so unbounded approvals.timeout
cannot overflow the lock acquire.

Fixes #83220
2026-08-24 16:11:03 +05:30
fangliquanflq
c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
joaomarcos
5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium
9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Teknium
e26d91dc11 feat(bot-mode): message_agent tool — structured, Bot-Chat-only agent-to-agent DMs
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.

Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).

The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
2026-08-21 15:23:51 -07:00
Brooklyn Nicholson
c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium
761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
Teknium
09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00
Brooklyn Nicholson
23d88c2b0e feat(tools): tour — let the agent walk a user through the UI
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.

Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
2026-08-19 00:52:54 -05:00
ethernet
879d6a4c78 feat(tui_gateway): batch clarify bridge with per-question locks
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).

Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.

Both agent-side clarify dispatch sites forward the questions arg.
2026-08-18 21:28:53 -04:00
Teknium
c257e9196b fix: make every tool interruptible — sequential executor abandons on user interrupt
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.

Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
  daemon worker (timeout None no longer means inline blocking) and polls
  the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
  synthesize a cancelled tool result (_ToolCancelledResult), emit the
  terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
  exactly like _ToolTimeoutResult, so an abandoned worker finishing late
  cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
  it owns its own human wait.

Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
  upscale) replaced with _wait_fal_result(), which polls is_interrupted()
  in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
  the "upscale failed, use original" fallback.

Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
2026-08-16 11:32:02 -07:00
Nikola Hristov
d083b85591 feat(hooks): pre_tool_call content transformation via modify directive
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.
2026-08-15 23:03:52 -07:00
blunkjamie-dev
163d7af310 fix(session-search): forward detail through agent paths
(cherry picked from commit 5f6de984f170ef470c7fbbd7662484bfaffc821d)
2026-08-15 00:34:29 -07:00
Ufonik
54aad0d670 fix(tools): refresh activity heartbeat while a tool call is in flight (#84491)
The gateway turn-inactivity watchdog (gateway/run.py::_watch_gateway_turn_inactivity)
abandons a turn once seconds_since_activity exceeds the inactivity timeout
(default 30 min). Activity was only stamped when a tool started and when it
completed, so a tool call that runs silently for 30+ minutes (quiet builds,
long pytest suites, large downloads, network waits with no output) froze the
clock and the watchdog hard-abandoned a turn that was still making progress,
reaping the tool's processes mid-execution (issue #84491).

Add a daemon-thread heartbeat inside _run_agent_tool_execution_middleware that
touches agent._touch_activity every 30s while the tool is in flight, until the
call returns. Both the sequential and concurrent execution paths funnel through
this single middleware, so one heartbeat covers every tool. The thread is
stopped in a try/finally so it always tears down even if execute() raises, and
is never started when a guardrail/authorization block short-circuits before
dispatch. A genuinely hung tool remains bounded by the tool layer's own
timeouts (terminal default 180s, concurrent batch deadline ~420s), so the
heartbeat only extends the turn's life while the call is legitimately running.

Verified by an independent reviewer (no security/logic defects); 5 unit/integration
tests pass on Python 3.12 (upstream CI). The 30-min gateway backstop remains
for turns whose agent loop itself stalls.
2026-08-14 21:20:03 -07:00
kshitij
367f0c21ed feat(agent): resolve sequential tool deadline via timeouts.tools.sequential_call (#85125 2a)
Follow-up on the #84795 salvage: the sequential deadline gets its own
resolver key. Unset, it inherits the concurrent batch deadline (same
value, same HERMES_CONCURRENT_TOOL_TIMEOUT_S bridge) so the two executor
paths cannot drift by default; set, it can be tuned or disabled
independently. Documented in cli-config.yaml.example; 5 contract tests.

Deliberately NOT on run_bounded_sync: the executors extend deadlines
dynamically during human approval waits (authorization-gate excluded
seconds) — the shared primitive is fixed-deadline. Noted in the docstring.
2026-08-15 02:21:39 +05:30
fangliquanflq
61645cde82 fix(agent): exempt clarify from sequential tool deadline
Clarify waits on a human for up to 3600s or unlimited. The generic sequential timeout was aborting that wait at 420s and leaving the prompt and worker active.
2026-08-15 02:21:39 +05:30
fangliquan
82a1b5a115 fix(agent): suppress late timeout observer events 2026-08-15 02:21:39 +05:30
fangliquanflq
ededa8c4f1 fix(agent): bound sequential tool calls 2026-08-15 02:21:39 +05:30
kshitij
d6a5cb9725 Merge pull request #85147 from kshitijk4poor/feat/unified-deadline-layer
feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
2026-08-15 01:10:40 +05:30
Teknium
2a26693e22 feat(delegation): live orchestration of running subagents via delegate_task action param
delegate_task gains a control plane: action='list' / 'steer' / 'stop'
let the parent agent see, redirect, and early-stop its own running
subagents mid-flight — the model-facing counterpart of the TUI's
delegation.pause / subagent.interrupt / subagent.steer RPCs.

- action='list': live children of this conversation's spawn tree
  (ids, goal, status, running_seconds, accepting_steer, live
  transcript path). Ownership is enforced via a _delegate_parent_ref
  weakref chain stamped at child build time, so a conversation can
  only control its own descendants, never a sibling tree.
- action='steer': queues text into a running child via the existing
  steer_subagent() registry path (delivered at the child's next tool
  boundary; missed steers surface as missed_steer in the completion).
- action='stop': interrupt_subagent() — child stops at its next
  iteration boundary, partial result still re-enters as a completion.
- Spawn dispatch response now includes subagent_ids + control hint.
- Control actions run synchronously (never backgrounded) and bypass
  the spawn pause gate and depth limit; they also never consume the
  per-turn subagent spawn cap, and remain usable once the cap is hit
  (that is when stop matters most).
- Small-model robustness (found live with gpt-5.4-mini on Nous
  Portal): tasks=[] alongside goal no longer trips the "Batch mode
  requires at least 2 tasks" gate — treated as single-goal.
- CLI display: control calls render as "steer sa-…" / "list" instead
  of an empty goal.

Live-tested E2E on Nous Portal (fable-5 + gpt-5.4-mini): full
spawn→list→steer→stop cycle, plus a steer-efficacy run where the
child acked the steer mid-essay and switched topics before finishing.
2026-08-13 09:34:36 -07:00
kshitij
083f8a6071 feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
One shared foundation for the timeout/hang backlog instead of per-incident
site-local fixes:

- agent/deadline.py: run_bounded_async (thread-timer deadline that survives
  a blocked event loop, generalizing the telegram adapter primitive),
  run_bounded_sync, clamp_timeout (kills the #83220 time_t OverflowError
  class at the boundary), resolve_timeout (config.yaml timeouts: section >
  legacy env bridge > default), kill_process_tree (whole-tree termination
  for the #71148 orphan class), DeadlineExpired (our deadline, mechanically
  distinct from provider timeouts).
- tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver;
  exact legacy env-var contract preserved (default 420, 0 disables).
- timeouts: accepted as a known config root; documented in
  cli-config.yaml.example.

Pure addition otherwise — no behavior change, no new env vars, no cache
impact. Later phases (#85125) migrate tool-execution, MCP, and subprocess
call sites onto these primitives.
2026-08-13 13:33:22 +05:30
Brooklyn Nicholson
adbc77eb50 feat(desktop): setup_mcp tool — inline MCP consent card over the clarify-style blocking bridge
New desktop_ui tool: the agent proposes an MCP server (install/enable/
authorize + a one-line reason) and blocks on mcp.setup.request until the
renderer's consent card answers mcp.setup.respond with the outcome
(installed/enabled/authorized/declined/unanswered/error). Same lifecycle
as clarify: 10-min timeout, allow_expired late answers, tool lifecycle
events forced on so the card mounts even with tool progress off. Desktop
prompt hint steers the model to the tool instead of hand-editing config;
every other surface keeps the schema out and is pointed at hermes mcp
install.
2026-08-13 01:06:51 -05:00