155 Commits

Author SHA1 Message Date
Teknium
4dcbbc84ab refactor(agent): table-driven inline tool executors shared by both tool paths; dedupe tool_executor
- agent/inline_tool_executors.py: INLINE_TOOL_EXECUTORS dispatch table (13 agent-level
  tools) replaces two drifted if/elif chains (invoke_tool + execute_tool_calls_sequential);
  resolve_invoke_tool_executor preserves the concurrent path's historical precedence.
- tool_hook_ids / emit_terminal_post_tool_call: single owners for hook identity kwargs
  and the terminal post_tool_call emit (was 8 + 3 hand-copied sites).
- tool_executor: shared quiet-spinner start/stop, tool-search unwrap + alias mapping,
  tool-result finalization tail and completion-callback fan-out between the concurrent
  and sequential executors; quiet/default registry branches merged.
2939 -> 2327 LOC.
2026-09-02 13:29:31 -07:00
DragonnZhang
9514d354ca fix(tool-search): validate deferred tool_call arguments against the concrete schema before dispatch
The generic tool_call(name, arguments: object) bridge hides a deferred tool's
real parameter schema from provider-native validation. Before this, only
top-level required-key absence was checked, so invalid enums, wrong types,
nested required fields and forbidden extra properties reached the handler
or MCP server. Now the call is coerced (same coerce_tool_args path normal
dispatch uses) and validated with the schema's declared JSON Schema draft;
failures return the path, constraint and parameters schema so the model
repairs the call in one round-trip. Fails open on missing/malformed schemas,
external $ref, or missing jsonschema.

Fixes #73175

Salvaged from #73179 onto current main (post core-tool deferral #97979).
Co-authored-by: teknium1 <teknium@nousresearch.com>
2026-09-01 23:30:33 -07:00
Alex Fournier
c3e86f7055 Merge branch main into chore/relay-0.8.0
Signed-off-by: Alex Fournier <afournier@nvidia.com>

# Conflicts:
#	uv.lock
2026-08-30 19:28:08 -04:00
Teknium
e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
Teknium
217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
Alex Fournier
31e7237adb fix(relay): upgrade to 0.8.1
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-28 08:28:21 -04:00
kshitij
f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30
Teknium
a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
kshitij
4341cf1df7 fix(approval): clamp approvals.timeout at the config-read chokepoint (#83220)
Widen the salvaged gate-site clamp to the bug class. The gate min() from
the contributor PR capped the gate bound at 360s, which would break the
#79719 contract (gate must extend while a legitimate >360s approval
prompt is answerable) and left the sibling overflow sites live: the CLI
prompt thread.join, the gateway poll deadline, and human_wait_ceiling
all consume the same config value.

Clamp once in _get_approval_timeout() via agent.deadline.MAX_SAFE_TIMEOUT_S
(1 year - semantically unbounded, platform-safe). The gate keeps its
approvals.timeout-tracking behavior above 360s; 7 regression tests pin
lock-acquire/thread-join safety and the gate-extension contract.

Bilateral E2E: with approvals.timeout=1e20 in a real config.yaml, main
crashes every consumer with OverflowError; this branch survives all 5
probes.
2026-08-24 16:11:03 +05:30
RelaxJonh
0dd0f6e648 fix(agent): clamp authorization gate lock timeout to prevent OverflowError on macOS
The _ConcurrentToolAuthorizationGate uses threading.Lock.acquire(timeout=...)
where timeout comes from human_wait_ceiling() (approvals.timeout + 60s).
When approvals.timeout is set very large (e.g. 999999999999 to effectively
disable timeouts), this overflows macOS timespec and raises:
  OverflowError: timestamp out of range for platform time_t

The gate only needs to serialize parallel dispatch — it should never wait
longer than the wedged-holder bound (_AUTHORIZATION_GATE_LOCK_TIMEOUT_S =
360s). Clamp the return value with min() so unbounded approvals.timeout
cannot overflow the lock acquire.

Fixes #83220
2026-08-24 16:11:03 +05:30
fangliquanflq
c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
joaomarcos
5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium
9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Teknium
e26d91dc11 feat(bot-mode): message_agent tool — structured, Bot-Chat-only agent-to-agent DMs
Bot Mode agents now DM teammates through a real tool instead of
hand-assembled shell commands. message_agent(target, message) validates
the target against the live roster, applies the sender's attribution
prefix server-side, and delivers over the existing proven transports
(hermes -p ... --query-file for local teammates, hermes peer dm for
peer gateways) as a tracked background process with notify-on-complete
— fire-and-forget, the reply wakes the sender on a later turn.

Containment: the schema is injected per-turn ONLY into a bot's
canonical 'Bot Chat' session on Bot-Mode-managed installs (same gate as
the protocol section); it is never registered in the tool registry or
any toolset, and dispatch re-gates on the session title so a forged
call from any other session refuses. The gate is session-stable, so the
tool list stays byte-identical across turns (prompt-cache safe).

The protocol section is rewritten to teach the tool and now carries the
teammate roster WITH ROLES (Bot Mode title + profile description), so
bots know who does what before picking a recipient. Roles and a
protocol version salt join the capability fingerprint: existing eternal
Bot Chats adopt the v2 protocol + tool with one epoch refresh, and a
rename/description edit refreshes the roster on the next message.
2026-08-21 15:23:51 -07:00
Brooklyn Nicholson
c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium
761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
Teknium
09e657793e feat: MCP tool results spill at 50K and carry upstream-elision warnings
Composio-style MCP servers return un-paginated 22-47K-char payloads that
sail under the generic 100K per-result spillover threshold, bloating
context and ballooning per-turn reasoning time on long conversations.
Competitors cap harder (OpenCode/pi 50KB, Claude Code 30K, Codex ~10K
tokens). Three changes:

- mcp_* tools spill at a tighter 50K default (BudgetConfig.mcp_result_size,
  config-overridable via tool_budget.mcp_result_size_chars; pinned and
  per-tool overrides still win; capped by the context-scaled default).
- The persisted-output preview now teaches recovery: page the saved file
  with read_file or process with execute_code instead of re-requesting the
  same data from the remote API.
- Untrusted/MCP string results are scanned (bounded, first 64KB) for
  provider-side elision markers ('...N more items', "has_more": true,
  'saved to sandbox', data_preview) and get ONE cache-safe incompleteness
  notice appended at result-construction time, before untrusted wrapping —
  so the model stops treating provider-elided enumerations as complete.
- Hard 2M-char allocation cap in mcp_tool.py (text, error, and
  structuredContent paths) so a pathological multi-MB server payload is
  bounded before it propagates, while ordinary large results reach
  spillover intact. Distilled from #56060/#56072/#56511 (issue #56059);
  supersedes their 50K lossy truncation with spillover-friendly semantics.

Docs: configuration.md spillover-budget section + cli-config.yaml.example.

Co-authored-by: Stoltemberg <215755014+Stoltemberg@users.noreply.github.com>
Co-authored-by: AlexFucuson9 <295703459+AlexFucuson9@users.noreply.github.com>
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-19 16:31:16 -07:00
Brooklyn Nicholson
23d88c2b0e feat(tools): tour — let the agent walk a user through the UI
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.

Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
2026-08-19 00:52:54 -05:00
ethernet
879d6a4c78 feat(tui_gateway): batch clarify bridge with per-question locks
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).

Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.

Both agent-side clarify dispatch sites forward the questions arg.
2026-08-18 21:28:53 -04:00
Teknium
c257e9196b fix: make every tool interruptible — sequential executor abandons on user interrupt
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.

Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
  daemon worker (timeout None no longer means inline blocking) and polls
  the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
  synthesize a cancelled tool result (_ToolCancelledResult), emit the
  terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
  exactly like _ToolTimeoutResult, so an abandoned worker finishing late
  cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
  it owns its own human wait.

Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
  upscale) replaced with _wait_fal_result(), which polls is_interrupted()
  in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
  the "upscale failed, use original" fallback.

Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
2026-08-16 11:32:02 -07:00
Nikola Hristov
d083b85591 feat(hooks): pre_tool_call content transformation via modify directive
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.

- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
  and returns (block_message, modified_args); modify directives
  shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
  {"action": "modify", "args": {...}} and Claude Code-compatible
  {"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
  dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).

Salvaged from PR #28953. Best fix for #18988.
2026-08-15 23:03:52 -07:00
blunkjamie-dev
163d7af310 fix(session-search): forward detail through agent paths
(cherry picked from commit 5f6de984f170ef470c7fbbd7662484bfaffc821d)
2026-08-15 00:34:29 -07:00
Ufonik
54aad0d670 fix(tools): refresh activity heartbeat while a tool call is in flight (#84491)
The gateway turn-inactivity watchdog (gateway/run.py::_watch_gateway_turn_inactivity)
abandons a turn once seconds_since_activity exceeds the inactivity timeout
(default 30 min). Activity was only stamped when a tool started and when it
completed, so a tool call that runs silently for 30+ minutes (quiet builds,
long pytest suites, large downloads, network waits with no output) froze the
clock and the watchdog hard-abandoned a turn that was still making progress,
reaping the tool's processes mid-execution (issue #84491).

Add a daemon-thread heartbeat inside _run_agent_tool_execution_middleware that
touches agent._touch_activity every 30s while the tool is in flight, until the
call returns. Both the sequential and concurrent execution paths funnel through
this single middleware, so one heartbeat covers every tool. The thread is
stopped in a try/finally so it always tears down even if execute() raises, and
is never started when a guardrail/authorization block short-circuits before
dispatch. A genuinely hung tool remains bounded by the tool layer's own
timeouts (terminal default 180s, concurrent batch deadline ~420s), so the
heartbeat only extends the turn's life while the call is legitimately running.

Verified by an independent reviewer (no security/logic defects); 5 unit/integration
tests pass on Python 3.12 (upstream CI). The 30-min gateway backstop remains
for turns whose agent loop itself stalls.
2026-08-14 21:20:03 -07:00
kshitij
367f0c21ed feat(agent): resolve sequential tool deadline via timeouts.tools.sequential_call (#85125 2a)
Follow-up on the #84795 salvage: the sequential deadline gets its own
resolver key. Unset, it inherits the concurrent batch deadline (same
value, same HERMES_CONCURRENT_TOOL_TIMEOUT_S bridge) so the two executor
paths cannot drift by default; set, it can be tuned or disabled
independently. Documented in cli-config.yaml.example; 5 contract tests.

Deliberately NOT on run_bounded_sync: the executors extend deadlines
dynamically during human approval waits (authorization-gate excluded
seconds) — the shared primitive is fixed-deadline. Noted in the docstring.
2026-08-15 02:21:39 +05:30
fangliquanflq
61645cde82 fix(agent): exempt clarify from sequential tool deadline
Clarify waits on a human for up to 3600s or unlimited. The generic sequential timeout was aborting that wait at 420s and leaving the prompt and worker active.
2026-08-15 02:21:39 +05:30
fangliquan
82a1b5a115 fix(agent): suppress late timeout observer events 2026-08-15 02:21:39 +05:30
fangliquanflq
ededa8c4f1 fix(agent): bound sequential tool calls 2026-08-15 02:21:39 +05:30
kshitij
d6a5cb9725 Merge pull request #85147 from kshitijk4poor/feat/unified-deadline-layer
feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
2026-08-15 01:10:40 +05:30
Teknium
2a26693e22 feat(delegation): live orchestration of running subagents via delegate_task action param
delegate_task gains a control plane: action='list' / 'steer' / 'stop'
let the parent agent see, redirect, and early-stop its own running
subagents mid-flight — the model-facing counterpart of the TUI's
delegation.pause / subagent.interrupt / subagent.steer RPCs.

- action='list': live children of this conversation's spawn tree
  (ids, goal, status, running_seconds, accepting_steer, live
  transcript path). Ownership is enforced via a _delegate_parent_ref
  weakref chain stamped at child build time, so a conversation can
  only control its own descendants, never a sibling tree.
- action='steer': queues text into a running child via the existing
  steer_subagent() registry path (delivered at the child's next tool
  boundary; missed steers surface as missed_steer in the completion).
- action='stop': interrupt_subagent() — child stops at its next
  iteration boundary, partial result still re-enters as a completion.
- Spawn dispatch response now includes subagent_ids + control hint.
- Control actions run synchronously (never backgrounded) and bypass
  the spawn pause gate and depth limit; they also never consume the
  per-turn subagent spawn cap, and remain usable once the cap is hit
  (that is when stop matters most).
- Small-model robustness (found live with gpt-5.4-mini on Nous
  Portal): tasks=[] alongside goal no longer trips the "Batch mode
  requires at least 2 tasks" gate — treated as single-goal.
- CLI display: control calls render as "steer sa-…" / "list" instead
  of an empty goal.

Live-tested E2E on Nous Portal (fable-5 + gpt-5.4-mini): full
spawn→list→steer→stop cycle, plus a steer-efficacy run where the
child acked the steer mid-essay and switched topics before finishing.
2026-08-13 09:34:36 -07:00
kshitij
083f8a6071 feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
One shared foundation for the timeout/hang backlog instead of per-incident
site-local fixes:

- agent/deadline.py: run_bounded_async (thread-timer deadline that survives
  a blocked event loop, generalizing the telegram adapter primitive),
  run_bounded_sync, clamp_timeout (kills the #83220 time_t OverflowError
  class at the boundary), resolve_timeout (config.yaml timeouts: section >
  legacy env bridge > default), kill_process_tree (whole-tree termination
  for the #71148 orphan class), DeadlineExpired (our deadline, mechanically
  distinct from provider timeouts).
- tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver;
  exact legacy env-var contract preserved (default 420, 0 disables).
- timeouts: accepted as a known config root; documented in
  cli-config.yaml.example.

Pure addition otherwise — no behavior change, no new env vars, no cache
impact. Later phases (#85125) migrate tool-execution, MCP, and subprocess
call sites onto these primitives.
2026-08-13 13:33:22 +05:30
Brooklyn Nicholson
adbc77eb50 feat(desktop): setup_mcp tool — inline MCP consent card over the clarify-style blocking bridge
New desktop_ui tool: the agent proposes an MCP server (install/enable/
authorize + a one-line reason) and blocks on mcp.setup.request until the
renderer's consent card answers mcp.setup.respond with the outcome
(installed/enabled/authorized/declined/unanswered/error). Same lifecycle
as clarify: 10-min timeout, allow_expired late answers, tool lifecycle
events forced on so the card mounts even with tool progress off. Desktop
prompt hint steers the model to the tool instead of hand-editing config;
every other surface keeps the schema out and is pointed at hermes mcp
install.
2026-08-13 01:06:51 -05:00
doncazper
85020f2238 fix(plugins): isolate ownership by profile 2026-08-12 19:13:32 -07:00
Brooklyn Nicholson
406501fd97 feat(agent): read_window_below tool — which OS window is underneath the desktop app
Desktop-gated (desktop_ui toolset) metadata-only window awareness: the agent
can ask which application window sits directly behind the Hermes window
(app, title, bounds — never pixels). Rides the same blocking bridge as
read_terminal: the gateway emits window.read.request and the renderer
answers window.read.respond.
2026-08-08 12:17:50 -05:00
kshitij
98e96e1a60 refactor(agent): drop the run_agent classify_persistence_error delegating wrapper
Post-merge simplify finding on #81613: the wrapper's docstring claimed it
existed for 'existing callers', but every caller was introduced by the same
PR - there was never a pre-existing import path to preserve. All callers
(conversation_loop, tool_executor, run_agent's own flush handler, tests)
now import the canonical hermes_state.classify_persistence_error directly,
matching how is_disk_full_error is consumed. No behavior change; imports
stay lazy inside the exception handlers.
2026-08-08 14:25:12 +05:30
Victor Kyriazakos
2a9f5b3476 fix(agent): classify session-persistence failures so lock contention is not misdiagnosed as disk-full
An enterprise deployment hit sustained SQLite write-lock contention on a
shared multi-gigabyte state.db (gateway + CLI processes writing
concurrently). Turns correctly failed closed with
session_persistence_failed, but the only user-facing wording claimed the
disk was full and the gateway rendered a generic failure.

The fast-fail semantics are deliberate and unchanged. This adds a pure
classifier (locked / disk / unknown) applied where the SQLite error is
still visible, threads the cause through the turn-completion explainer,
and stamps a machine-readable failure_reason
(session_persistence_failed:<cause>) plus a guaranteed non-empty error on
the result for downstream surfaces. The cron scheduler's explainer-text
suppression now matches every cause variant so refined wording cannot
leak into scheduled-job deliveries.
2026-08-08 14:18:26 +05:30
kshitij
ea0d54db1d refactor: fold /simplify-code findings
- Single source for the approval-derived bound: public human_wait_ceiling()
  in tools/approval.py; the gate's lock-timeout helper delegates to it
  instead of re-deriving timeout + margin (was duplicated in two modules
  and reached for a private _get_approval_timeout).
- Shared _clamped_window_seconds() for the close-time accrual and the
  open-window read, so the two clamps are identical by construction.
- Gate __init__ grows session_key kwarg; tests construct via the real
  constructor instead of mutating privates post-hoc.
- Gateway test resolves its pending approval via resolve_gateway_approval()
  (the production /deny path) instead of hand-rolling queue-entry internals.
- Docstring accuracy: human_wait_seconds monotonicity caveat under cap
  eviction; s/pre_tool_block/pre_tool_call/ hook name.
2026-08-06 17:03:10 +05:30
kshitij
10fb01e725 fix: harden human-wait tracker from review findings
Review-driven follow-up to the #79719 fix:

- Clamp the CLOSE-side accrual too: a wedged window that eventually closed
  used to inject its full unclamped overstay into completed_seconds,
  retroactively extending a running batch's deadline by hours. Both clamps
  now share one ceiling helper (_human_wait_ceiling = approvals.timeout +
  HUMAN_WAIT_MARGIN_S), and the gate's lock-timeout uses the same margin
  constant so the bounds cannot drift apart.

- Evict idle sessions until the table is under the cap (was: at most one
  per insert, so churn could outgrow _HUMAN_WAIT_MAX_SESSIONS). Entries
  with an open window are still never evicted.

- Log (debug) instead of silently swallowing a failed session-key snapshot
  in the gate constructor.

Tests: close-side clamp regression + table-cap assertion added; suite at
17 passed.
2026-08-06 17:03:10 +05:30
kshitij
3305cfd2bb fix(agent): measure batch-deadline exclusion at the human wait, not authorization-gate residency
A tool wedged inside _ConcurrentToolAuthorizationGate hung the whole turn
forever (#79719): excluded_seconds() measured residency in gate.run() —
arbitrary code — so an open window grew 1:1 with wall clock and the batch
deadline's remaining was constant (remaining = deadline - window_started;
now cancels out). A hanging pre_tool_call plugin or an approval round-trip
to a dead client defeated the deadline entirely. The serialization lock was
also an unbounded acquire, so every other worker needing authorization
parked behind the wedged holder forever.

Fix, in two halves:

- tools/approval.py grows per-session human-wait accounting
  (human_wait_window / human_wait_seconds). The two places that are
  verifiably blocked on a HUMAN — the CLI approval prompt and the gateway
  approval poll loop — mark their own windows. Both are intrinsically
  bounded by approvals.timeout; the open-window read is additionally
  clamped to that timeout plus a margin as belt-and-braces.

- _ConcurrentToolAuthorizationGate keeps only serialization, with a bounded
  acquire (approvals.timeout + 60s; on expiry the prompt runs unserialized —
  the same degradation the start-order gate accepted in #79705).
  excluded_seconds() becomes a baseline-delta read of the session's
  human-wait total.

A wedged plugin now contributes nothing to the exclusion, so the batch
times out at the normal deadline with correctly labeled results, while a
genuine approval wait — which can legitimately exceed any fixed bound —
still extends the deadline in full. E2E (real AIAgent, worktree imports):
wedged-plugin batch on main never ends (>30s observed, 3s deadline); with
the fix it ends at 3.0s. A 4s simulated approval over a 2s deadline
completes without a timeout label.

Closes #79719
2026-08-06 17:03:10 +05:30
kshitij
042a2cf3d7 fix(agent): keep the start-order gate under the batch deadline and abort abandoned workers
Follow-up to the salvaged start-order gate bound. Two gaps remained, both
reachable through the same knob.

1. The gate bound ignored the batch deadline it sits under. With
   HERMES_CONCURRENT_TOOL_TIMEOUT_S below 120s the deadline fired first, so
   the parked tools were still reported as "timed out" without ever running --
   the exact bug the bound exists to fix. The gate now clamps to
   min(120s, batch_timeout / 2), matching the sibling constant's documented
   habit of relating the two timeouts.

2. A gate-parked worker released purely by its own timeout could wake up after
   the batch was abandoned and dispatch its tool anyway: wasted work whose
   result nobody reads, a duplicate post_tool_call for a tool_call_id the turn
   already closed as timeout, and agent._current_tool left pointing at a dead
   tool for the rest of the session (the main thread's reset already ran).
   Abandonment is now a first-class wakeup: both abandon sites set an event and
   notify the condition, and a released worker raises _BatchAbandoned instead
   of dispatching. Parked threads are reclaimed in milliseconds rather than one
   full gate timeout plus a tool runtime.

Also names the tool in the gate-timeout warning. The closure's function_name
binds the last-parsed tool, so logging it directly would have printed the wrong
name; it is threaded through _begin_in_order instead.

Measured, 3-tool batch with the first tool wedged during dispatch:

                          main    PR as-is   with this commit
  dispatched in batch       0        0          tool_b, tool_c
  dispatched after return   0        2 (ghost)  0
  _current_tool leaked      no       "tool_b"   no

Adds tests/run_agent/test_start_order_gate.py (3 tests). Mutation-checked
against the parent commit: the starvation guard passes there (it binds the
salvaged fix), while the deadline-clamp and abandonment guards both fail,
reproducing the ghost dispatch as
"tool(s) dispatched after the batch was abandoned: [tool_a, tool_b]".
2026-08-06 03:53:38 +05:30
Sylvain Baeriswyl
5d83400f89 fix(agent): bound the concurrent start-order gate wait
_begin_in_order parks each concurrent tool worker on a timeout-less
Condition.wait_for until every earlier-ordered tool has advanced through
its dispatch. If one tool wedges during dispatch (observed in production:
a stuck skill_view; also reproducible via any blocking authorization),
three failures compound: every later-ordered worker is starved and never
starts; the batch deadline then falsely reports those never-started tools
as "timed out" (sub-second read_file/search_files calls get blamed while
having done zero work, and the model reasons against that false failure
info); and after the batch is abandoned the parked workers leak forever —
f.cancel() cannot cancel running threads, the per-thread interrupt flag
is never polled inside wait_for, and nothing notifies the condition
again. Confirmed with a faulthandler all-threads dump taken after batch
abandonment showing workers still parked at the gate.

Bound the wait at 120s; on expiry, log a warning and proceed out of
order (worst case: interleaved approval prompts — strictly better than
permanent starvation). The >= predicate lets one worker's timeout-jump
release every skipped worker immediately, and max() keeps the counter
monotonic for out-of-order advancement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 03:53:38 +05:30
brooklyn!
64646dda56 Hermes can read the in-app browser (#79482)
* feat(agent): read_preview — the desktop-gated tool that reads the in-app browser

The agent could open the preview pane (open_preview) and read the embedded
terminal (read_terminal), but the browser it had just opened was a black box —
'what does this page say?' had no answer. read_preview mirrors read_terminal
end to end: HERMES_DESKTOP-gated via check_fn (zero schema footprint outside
the GUI), dispatched through the same agent callback pattern, windowed with
start/count so a long page pages instead of flooding context.

* feat(gateway): preview.read blocking bridge

Same lifecycle as terminal.read: the tool blocks on preview.read.request, the
renderer answers preview.read.respond (allow_expired — a slow page extraction
losing the 45s race must not surface a raw 4009), and a timeout emits
preview.read.expire so late answers resolve quietly.

* feat(desktop): the renderer serializes the active preview tab for the agent

preview-reader.ts is the preview analog of the terminal's buffer registry: the
URL pane registers a page reader (webview executeJavaScript → title + visible
innerText) keyed by tab id; readActivePreview resolves the ACTIVE tab, windows
the text (24k cap per read), and answers file/artifact tabs with identity plus
a note pointing at the tool that reads that content directly. The gateway
event handler answers preview.read.request beside terminal.read.request.
2026-08-05 16:35:00 +00:00
Jeffrey Quesnelle
daf67f2e59 Merge branch 'main' into feat/hermes-relay-tool-metrics 2026-08-04 12:43:38 -04:00
kshitijk4poor
a7ad713f43 fix(tool-executor): unpack 5-tuple runnable_calls in _max_workers_for_tool_batch 2026-08-03 22:53:32 +05:30
EndeavorYen
c0b0cc3925 feat(image): parallelize image_generate batches 2026-08-03 22:53:32 +05:30
Alex Fournier
14c8bd646c Merge updated model metrics into tool metrics
Signed-off-by: Alex Fournier <afournier@nvidia.com>
2026-08-02 20:11:20 -07:00
annguyenNous
d5cf89f3b2 fix(tool-executor): use lazy logging format in debug calls
Re-applied from #50508 onto current main: the executor moved from
tools/ to agent/ and the JSON-error site was refactored away into
_parse_tool_arguments, but these 8 f-string debug calls survived the
move verbatim. Lazy %s formatting skips string interpolation when
DEBUG is off — the Tool-result line interpolates the FULL tool output
(can be 100KB+) on every tool call otherwise.
2026-08-02 23:17:40 +05:30
Teknium
38c09e5d73 fix(tool-executor): emit tool results on hard interrupt to keep alternation
The sequential executor's KeyboardInterrupt handlers emitted a cancelled
post-tool-call event for the current tool, called agent.interrupt(), then
re-raised — WITHOUT appending a tool result message for the interrupted call
or any remaining calls in the batch. The assistant tool-call turn was left
with no matching tool results, a message-role alternation violation that
malforms the next provider request (relying on downstream repair passes to
patch it, which don't run on every path).

The cooperative-interrupt block (_interrupt_requested) and the concurrent
executor already emit a result for every call_id; this brings the two hard-
interrupt handlers into line via a shared _append_cancelled_tool_results
helper that appends a cancelled result for the current + remaining calls
before re-raising.

Verified live before/after (0 tool results -> 3 for a 3-call batch
interrupted on the first tool) and with a sabotage-checked regression test.
52 interrupt/executor tests pass.
2026-08-01 16:42:57 -07:00
Xipong
0fd0db1a8c fix(agent): preserve /steer through turn-budget enforcement 2026-07-31 22:32:52 -07:00
Alex Fournier
aef76ba398 Merge updated model metrics into tool metrics
# Conflicts:
#	hermes_cli/observability/shared_metrics_contract.py
#	hermes_cli/observability/shared_metrics_subscriber.py
#	scripts/smoke_nemo_relay_shared_metrics.py
#	tests/hermes_cli/test_plugins.py
#	tests/hermes_cli/test_relay_shared_metrics.py
#	tests/hermes_cli/test_relay_shared_metrics_runtime.py
#	tests/run_agent/test_run_agent.py
#	tests/test_model_tools.py
#	tests/tools/test_approval.py
2026-07-31 07:18:02 -07:00
Soju06
ce9f6712ff refactor(agent): remove the inter-tool delay
The 1.0s sleep between sequential tool calls has been present since the
initial commit with no documented rationale. It sleeps between local
tool executions — the next LLM request goes out only after the whole
batch — so it rate-limits nothing, and the parallel read-only path
already runs with no delay. Every multi-tool turn pays (N-1) seconds
of dead time. Remove the sleep, the internal tool_delay plumbing, and
dead test assignments. AIAgent.__init__ keeps tool_delay as a
deprecated no-op keyword for one release so existing programmatic
callers construct cleanly; passing it emits a DeprecationWarning.
2026-07-29 15:21:26 -07:00