Commit Graph

3423 Commits

Author SHA1 Message Date
Teknium
18700c60d5 Merge branch 'simp/tools' into simp/integration 2026-09-02 15:01:45 -07:00
Teknium
aed2e156ad refactor(tools/files): finish file_operations wiring to extracted modules; restore WHY comments in file_tools; repoint test 2026-09-02 14:47:03 -07:00
Teknium
731a943ac4 refactor(tools/planning): compact kanban metadata merging and clarify tool/gateway helpers 2026-09-02 14:47:02 -07:00
Teknium
ae67178be1 refactor(tools/infra): split cronjob god-methods into per-action helpers; extract tool_search catalog/validation; compact registry, lazy_deps, tool_backend_helpers, desktop_ui 2026-09-02 14:47:02 -07:00
Teknium
6723628de9 refactor(tools/messaging): split send_message into senders/targets tables; dedupe discord/bot_mode/relay helpers; extract feishu_lark shared plumbing; compact graph client/auth 2026-09-02 14:45:42 -07:00
Teknium
20258aad9e refactor(tools/mcp_oauth): extract mcp_oauth_provider; compact oauth manager/dashboard bridge and schema_sanitizer; repoint tests 2026-09-02 14:45:42 -07:00
Teknium
8b3dde5dcc refactor(tools/web): split web_tools into truncate/extract siblings; dispatch tables for backend probes and stdlib extractors; compact url_safety/website_policy/result cache 2026-09-02 14:45:42 -07:00
Teknium
b6a3d3f440 refactor(tools): restore stdin=DEVNULL on subprocess probes dropped during compaction
voice_mode WSL playback probe, subagent_worktree._run_git, cua_backend
loginctl/xprop probes lost their explicit stdin= (and one line-wrapped past
the encoding= same-line rule); lazy_deps._run carries it via _SUBPROCESS_KW
so mark it for the guard. Fixes tests/tools/test_subprocess_stdin_guard.py
and tests/scripts/test_footgun_subprocess_encoding.py (green on base).
2026-09-02 14:45:42 -07:00
Teknium
8813eba345 refactor(tools): repoint tests to moved symbols; add voice_mode_transcript module 2026-09-02 14:45:15 -07:00
Teknium
9fe009f467 refactor(tools/media): extract vision_tools_image_prep; dedupe image/video generation providers; compact fal/xai helpers 2026-09-02 14:45:15 -07:00
Teknium
33ae442e94 refactor(tools/computer_use): split cua_backend into daemon/actions/... siblings; dispatch tables; compact doctor/tool 2026-09-02 14:45:15 -07:00
Teknium
3bbec90f23 refactor(tools/environments): split local/base into output/wait/session-env siblings; extract docker_egress + remote_common; dedupe remote backends; compact process_registry 2026-09-02 14:45:15 -07:00
Teknium
d4ded67503 refactor(tools/terminal): split terminal_tool god method into helpers; dispatch tables; compact env probe/passthrough/daemon pool 2026-09-02 14:45:15 -07:00
Teknium
2ef1e8e4e0 refactor(tools/voice): extract tts delivery + wake_word engines; dedupe transcription/voice_mode helpers; compact tts providers 2026-09-02 14:45:15 -07:00
Teknium
49d15faae6 refactor(tools/skills_sync,memory): split skills_sync client wire/org and bundled/optional ops; extract memory_tool_store and session_search common/discover 2026-09-02 14:45:15 -07:00
Teknium
7c1ec19d4a refactor(tools/skills): split skills_hub into per-source modules; extract skill_manager guards/batch, skills_tool dedup/plugin/setup; compact skill_usage, skills_guard 2026-09-02 14:45:15 -07:00
Teknium
05526b028a refactor(tools/delegate): split delegate_tool into child_run/config/dispatch/progress/registry/results; compact delegation helpers 2026-09-02 14:45:15 -07:00
Teknium
9f6335bc44 refactor(tools/browser): extract eval-policy/lightpanda-fallback/real-profile/snapshot modules from browser_tool; split supervisor dialogs/frames; dedupe camofox/cli 2026-09-02 14:44:15 -07:00
Teknium
c1f8af1e86 refactor(tools/mcp): split mcp_tool.py into transport/lifecycle/schema/handlers/... sibling modules; compact watchdog and schema cache 2026-09-02 14:44:15 -07:00
Teknium
606cb2de92 refactor(tools/code_exec): unify code_kernel local/remote helpers, split checkpoint_manager god methods, compact spill helpers 2026-09-02 14:44:15 -07:00
Teknium
5c919161d0 refactor(tools/approval): split approval.py into smart/human-wait/gateway-wait modules; dedupe guards 2026-09-02 14:44:14 -07:00
Teknium
ec49ae7c0f refactor(tools): wire file_operations to extracted common/lint/search modules; restore literal security pins in lazy_deps 2026-09-02 14:43:45 -07:00
Teknium
3f5564d727 refactor(tools): discovery scan sees loop-registered tools; fix lint module import 2026-09-02 14:43:45 -07:00
Teknium
d4cec15b47 refactor(tools): first-wave simplification of tools/ (file ops split, lazy_deps, code_exec, approval, browser, delegate, mcp, skills, terminal, voice, media)
Behavior-neutral structural pass over tools/*: god-file extractions into
sibling modules (file_operations_common/lint/search, file_tools_paths/
read_tracking/write, code_execution_env/rpc, tool_search_catalog/names/
validation, tts_command_provider, ...), duplicate helper unification,
if/elif -> dispatch tables, dead-code removal, docstring compaction.
Tool schemas (get_tool_definitions) verified byte-identical to base.
2026-09-02 14:43:45 -07:00
Teknium
fcb5c64101 refactor(agent): model_tools — extract argument type coercion into tools/arg_coercion.py (re-exported; logger name preserved) 2026-09-02 13:29:39 -07:00
kshitijk4poor
ff7233b815 fix(credential_files): apply the same exclusions to the symlink-safe mount copy
_safe_skills_path() is the sibling of iter_skills_files(): when a symlink in
skills/ forces a sanitized copy for mount-based backends (Docker/Singularity),
it rglob-copied the whole tree — .hub, .curator_backups, node_modules and all.
Prune EXCLUDED_SKILL_DIRS before descending, same rule as the sync generator,
so the mounted copy never carries (or walks) the bookkeeping trees either.
2026-09-03 01:33:18 +05:30
kshitijk4poor
1d06ef3a5d refactor(credential_files): prune excluded dirs before descending in the sync walk
Replaces the three hand-copied rglob loops + post-hoc parts check with one
os.walk generator that drops EXCLUDED_SKILL_DIRS from dirnames before
recursing. Same file set as the cherry-picked fix (the test binds it), but
the walk no longer stats every file under .hub/.curator_backups/node_modules
on each 5s FileSyncManager tick.

Bench (synthetic skills tree: 20 skills + 400 .hub files + 5x8MB curator
tarballs + 50 archived files): iter_skills_files() 35ms -> 2.4ms.
2026-09-03 01:33:18 +05:30
Carry00
edac49e473 fix(skills): stop syncing bookkeeping dirs to sandboxes
iter_skills_files() walked the skills tree with a bare rglob("*"), so the
.hub download cache, .archive, curator backups, and any node_modules/.git
under a skill package were uploaded to the sandbox on every sync. The
sandbox never reads them: skill content is resolved host-side.

EXCLUDED_SKILL_DIRS is already the canonical exclusion set, honoured by
discovery and backup. Apply it to the sync path too, across all three
roots iter_skills_files() walks (local, external, project-local), and add
.curator_backups to the set.

Measured on a local install: 900 files / 67.3 MB -> 771 files / 8.4 MB.

This is not just wasted bandwidth on the SSH backend, where the oversized
payload can exceed the 120s _ssh_bulk_upload deadline and surface as the
agent hanging on every tool call.

The filter intentionally does not reuse is_excluded_skill_path(), which
also prunes references/, templates/, assets/ and scripts/ -- those hold
support files and bundled scripts the sandbox does read and execute.
2026-09-03 01:33:18 +05:30
Teknium
f6234d00c5 fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session
directory automatically — the coding-workspace snapshot, gateway
project-tree build, /diff, @diff|@staged context refs, goal-gate
fingerprint, and -w startup worktree add — before any prompt, tool call,
approval, or trust gate. Those probes ran the system git without
stripping the repository's own config, so a repo delivered as files with
its .git directory intact (a shared zip, sync folder, or USB stick;
git clone never transfers .git/config) could set an execution-sink git
setting and get arbitrary host code execution as the user with nothing
on screen.

- core.fsmonitor / core.hooksPath / pager / editor / credential helper:
  neutralized by routing every automatic probe through
  noninteractive_git_env(), which pins those keys to inert values via
  GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the
  reported sink, coding_context._git + tui_gateway.git_probe) now defaults
  to that env; worktree-add, working_diff, web_git, context_references,
  goals, and subagent_worktree route through it too.
- Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker
  names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't
  enumerate them. Added harden_git_argv(), which inserts
  --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/
  log/blame) only — status et al reject the flags. Both flags required
  (verified empirically; each alone leaves the other live).

Builds on the noninteractive_git_env config-scrubbing from the
gemini-cli #28792 port. Real-git E2E regression suite arms a malicious
repo and asserts every automatic path neutralizes fsmonitor, hooks,
external-diff, and textconv; a baseline test proves the repo is armed.
2026-09-02 10:33:43 -07:00
Teknium
7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00
Victor Kyriazakos
fd35e1ec5a fix(cron): delivery bookkeeping reads the failure lane it actually routed through
Review findings (Salt, NS-788):

B1: delivery_outcome classification, unresolved_origin, and incident
'alerted' marking all read the deliver lane while the notice itself was
routed through failure_deliver — a silenced failure recorded
delivery_outcome='delivered' and marked its incident alerted (corrupting
the 'failure seen' vs 'operator was pinged' distinction the incident
store documents), and a failure delivered via failure_deliver over an
unresolvable deliver=origin recorded 'not_configured'. New
_delivery_lane_value() helper feeds the SAME lane to routing and
bookkeeping at all five sites (both classifiers, both unresolved_origin
computations, both zero-target checks). Three regression tests assert
outcome + alerted-marking; verified to bite on the pre-fix classifier.

S1: failure_deliver now goes through _resolve_cron_context_deliver on
tool create/update, matching deliver — a job created from inside a cron
run can no longer store literal 'origin' in its failure lane.

S2/T1: corrected the false 'same helper' comment in create_job; the
str/list flatten mirrors the tool layer for direct callers.

Full cron suite + interrupt tests: 87 files, 1112 passed, 0 failed.
2026-09-02 20:16:14 +05:30
Victor Kyriazakos
c9491e6a7d feat(cron): per-job failure_deliver — route or suppress failure notices (NS-788)
Coatue FR (Frank Long): jobs delivering into shared channels publish
engine failure notices ('⚠️ Cron X failed…') to those channels with no
opt-out. Adds an optional per-job failure_deliver field sharing
deliver's grammar: on failure, targets resolve from failure_deliver
when set (local = structural silence; state still recorded in
last_status/last_error/run history). Success delivery is unchanged;
absent field = today's behavior byte-for-byte.

Honored by every failure-category engine notice: the run_job failure
summary (+streak nudge), the escaped-failure retry path, drift-skip and
blocked-config alerts (composed into the same delivery), and the
gateway-shutdown interrupted-run notice (_notify_interrupted_cron_jobs).

Surfaces: cronjob tool create/update (same bot-chat validation as
deliver; '' clears on update), hermes cron create/edit
--failure-deliver, docs tip in automate-with-cron.

Existing fake_deliver test doubles gained **kwargs for the new
for_failure keyword — signature-compat only, no behavior change.
2026-09-02 20:16:14 +05:30
Teknium
8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
joaomarcos
65b0f00002 fix(agent): stop the between-turns tool refresh from forking the cached prefix
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:

* a tool whose `check_fn` merely flapped (headless browser probe, expired
  credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.

Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.

`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.

Refs #100336
2026-09-02 07:22:59 -07:00
Teknium
ee0e234a2c fix(gateway): discover and reload MCP servers per profile under multiplex
A multiplexed gateway ran `discover_mcp_tools()` once, unscoped, at boot
and again on `/reload-mcp`, so only the launch profile's `mcp_servers`
ever connected; secondary profiles' servers never registered, and a
`/reload-mcp` from any profile tore down every profile's connections.

- `_discover_gateway_mcp_tools()`: under multiplex, run discovery once per
  served profile inside `_profile_runtime_scope`, carried into the
  executor via `copy_context()` (same shape as
  `_run_in_executor_with_context`). Single-profile path unchanged.
- `_execute_mcp_reload()`: enter the requesting profile's scope when the
  caller (e.g. button-confirm callback) did not; shut down / rediscover /
  report only that profile's servers; refresh only that profile's cached
  agents.
- `shutdown_mcp_servers(scope=)`: scoped teardown keyed by the new
  `_server_scope_keys` ownership map; leaves the shared MCP loop running
  while other profiles' servers are live. Unscoped call keeps the full
  historical behavior.
- MCP tools register into the owning profile's registry overlay
  (`registry.register(scope=...)`), and `registry.deregister()` gains a
  matching `scope=` kwarg. Plugin callers still cannot name another
  profile's scope; the plugin-vs-global guard is unchanged for them.

Fixes #95518

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Kong <mgongzai@gmail.com>
Co-authored-by: roraag <232666910+roraag@users.noreply.github.com>
2026-09-02 07:00:13 -07:00
Teknium
4e7aa48716 fix(browser): reap idle multiplexed sessions under their owner profile scope
The inactivity janitor is one process-global thread started by whichever
profile first opens a browser, so under `gateway.multiplex_profiles` it runs
with no secret scope: `cleanup_browser` -> `is_camofox_mode` ->
`get_secret("CAMOFOX_URL")` raises UnscopedSecretError, the session entry is
never removed, and the same failure repeats every 30s while the Chromium
daemon leaks.

- `_update_session_activity` records the owning Hermes home per session;
  `_cleanup_inactive_browser_sessions` re-enters that owner's
  `set_hermes_home_override` + `build_profile_secret_scope` around each
  teardown (`_session_owner_scope`, mirroring `_profile_runtime_scope`).
  copy_context at thread spawn would pin the first profile's secrets onto
  every other profile's teardown; there is no os.environ fallthrough.
- 3 consecutive failures -> `_force_reap_browser_session`, which skips the
  failing `close` round-trips but still closes the cloud provider session
  and kills the local daemon via the shared `_release_session_resources`
  tail (extracted from `_cleanup_single_browser_session`, unchanged).
  An activity touch does not reset the failure budget.

Fixes #86402
Fixes #100738

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-02 07:00:13 -07:00
liuhao1024
7d509657a8 fix(tts): resolve default output dir from the active profile
DEFAULT_OUTPUT_DIR was resolved once at import time, so long-lived
multi-profile runtimes (dashboard console, TUI/Desktop backend, cron,
kanban workers) kept writing synthesized audio into the launch
profile's cache/audio instead of the requesting profile's (#98749).

Same bug class and fix as skills_tool (f8723c478) and skills_sync
(#65828): keep the legacy module attribute for tests and external
patchers, but re-resolve from the live profile-scoped HERMES_HOME on
every synthesis call.
2026-09-02 06:48:31 -07:00
liuhao1024
39fca697ae fix(tools): log boot-time UnscopedSecretError probes at debug, not warning
With multiplexing on, check_fns evaluated before any profile secret scope
exists fail closed by design: get_secret raises UnscopedSecretError and the
tool re-probes on the first scoped turn. _run_check_fn_uncached logged that
expected signal like a crashed check_fn (WARNING + exc_info), so every
multiplexed gateway start printed three full tracebacks that drowned real
check_fn failures.

Split the handler: an unscoped read reported while the profile cache scope
was unresolved logs one debug line without a traceback; the same error with
the scope resolved is a genuinely lost scope and keeps the loud
warning + traceback.

Fixes #100697
2026-09-02 06:48:31 -07:00
Teknium
d29a7936e4 fix(bot-mode): DMs to a Desktop-owned Bot Chat land in the live session instead of being dropped (#100523)
When the Desktop has a bot's "Bot Chat" open, that session holds the
single-owner lease, so the `hermes -p <bot> chat -c "Bot Chat"` subprocess
`bot_relay.deliver` spawns refuses with "already has a live owner" and the
DM payload is dropped — the sender was already acked.

bot_relay.deliver now looks up a live in-process session for the target
profile whose title resolves to "Bot Chat" (same profile_home match as
session.resume's _find_live_unpersisted, pending_title for lazy sessions,
otherwise the db title) and, when found, submits the message through the
existing prompt.submit handler — the composer's choke point — so it lands
as a normal user turn (role alternation preserved, streams to the open
window). No live owner → the subprocess path runs exactly as before.

On the local message_agent subprocess path, the lease refusal is surfaced
as a structured `target_busy` delivery failure telling the sender the
message was NOT delivered, instead of a raw exit-1 with the text buried
in stderr.

Closes #100523
Supersedes #100544, #100542

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: 686f6c61 <github@00b.tech>
2026-09-02 06:18:55 -07:00
Teknium
782dd635fe feat(tts): speech toggles warm/release plugin and command TTS providers
Extends the TTS lease from #100912 (4e3feb8bbb) beyond built-in local
engines: when the configured tts.provider is user-declared, acquiring
the first lease and releasing the last one now reach it, so a
self-hosted TTS server can preload its model when read-aloud / voice
conversation turns on and unload when it turns off (Discord request).

- agent/tts_provider.py: TTSProvider gains concrete no-op warm() /
  release() (not abstract — existing plugins are unaffected).
- tools/tts_tool.py: _signal_user_tts_provider() forwards the lease
  hook; plugin providers get warm()/release(), command providers run
  optional `warm_command` / `release_command` (config.yaml, under
  tts.providers.<name>) through the existing _run_command_tts helper
  on a daemon thread — best-effort, output discarded, failures at
  debug. warm_tts_provider() and release_tts_provider() call it.
- tests/tools/test_tts_lifecycle_leases.py: fake plugin provider and
  fake command provider observe warm/release through acquire/release
  lease (both fail on main with action == "noop").
- docs: features/tts.md — lease section, command-provider optional
  keys table, plugin optional hooks.
2026-09-02 05:35:14 -07:00
muhifni
1cd736ff63 fix(terminal): scope terminal config per turn under profile multiplexing
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.

Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:

- tools/terminal_scope.py: ContextVar holding the routed profile's
  COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
  TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
  resolves ONLY from it - an omitted key yields the defined default,
  never os.environ. Unreadable/malformed policy installs a refusal
  scope; terminal_tool / execute_code refuse instead of running under
  ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
  _profile_runtime_scope, tui_gateway session/build/turn scopes, cron
  per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
  (_get_env_config, _resolve_container_task_id shared key, orphan
  reaper lifetime, degraded mode), gateway/platforms/base.py docker
  media translation (volumes, shared key, persistence), runtime_cwd /
  agent_init / skill_utils / code_execution_tool / file_tools cwd
  anchors, prompt_builder / browser_tool / env_probe backend checks,
  gateway footer, @-refs and slash-command cwd. env_probe resolves the
  backend in the caller's context, since the probe worker thread does
  not inherit the ContextVar.

Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.

Fixes #68559
Fixes #94200
Fixes #101132
Fixes #95470

Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
2026-09-02 05:34:28 -07:00
Teknium
5c6cbbc1be fix(loops): pause /loop --until on a blocked verdict; trim redundant gate condition and duplicate test
The goal judge now returns 'blocked' for unachievable goals, but the
/loop --until gate only checked == 'done', so an impossible stop
condition would re-fire every tick until loops.max_ticks. Pause the
loop with the judge's reason instead. Also collapse the kanban gate
callers' 'gate_verdict == "continue" or rejection is not None' to
'rejection is not None' (rejection is None iff verdict == done), drop
the duplicate blocked-verdict goal test, and document the verdict.
2026-09-02 05:32:01 -07:00
itsflownium
1bd9fce6cb fix(kanban): judge unachievable goals as blocked, never done 2026-09-02 05:32:01 -07:00
Teknium
32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium
a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
Teknium
00a7115a02 fix(cron): make cron push-notify configurable (cron.delivery.notify) and surface UNVERIFIED live deliveries in cron list/doctor
De-risking for the notify=True UX change: the marker is now driven by
cron.delivery.notify (config.yaml, default true = current behaviour), read
once per delivery and applied to both the text and media routes; a missing or
malformed section keeps the default.

An evidence-free live-adapter ack (bare SendResult(success=True) from
Slack/Matrix/Mattermost) is still accepted, but the target is recorded on the
job as last_delivery_unverified (cleared by the next evidenced delivery) so
the state shows up in 'hermes cron list' (⚠ Delivery UNVERIFIED), 'hermes cron
doctor', and the cronjob tool listing — not only in a WARNING log line.

Live repro (real _deliver_result + real 'hermes cron list' against a temp
HERMES_HOME, Slack target, SendResult(success=True)): before — list showed
nothing beyond the Deliver line and route metadata always carried
notify=true; after — list prints the UNVERIFIED line, and
cron.delivery.notify: false yields notify=false in the route metadata.
2026-09-02 00:56:52 -07:00
Teknium
758114bb8d fix(cron): manual run reports delivery_failed as a failed run; docs for the distinct status
A manual cronjob(action='run') derived success from last_status == 'ok'
and read the error from last_error — so a run that now records
delivery_failed came back as success=False with error=None, an unexplained
failure. Surface last_delivery_error as the error in that case (the
#84006 direction, re-applied on the delivery_failed status), and pin the
manual-run completion summary to say 'Result: FAILED' over an undelivered
run. Document the status in the cron user guide.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-02 00:52:58 -07:00
赵桂雄
2f58cbfa7f fix(cron): adapt delivery-notice tests to the return_job claim API
Main grew claim_job_for_fire(job_id, return_job=True) — a claimed
snapshot dict instead of a bool — while this branch sat on an older
base. The merge-ref CI ran the hybrid: the wiring tests still mocked
return_value=True, which fails isinstance(claimed_job, dict) and fell
into the 'already being fired' branch, so every dispatch assert failed.

Mock the claim to return the job snapshot (the API's success shape),
read the summary's deliver from the claimed snapshot the run actually
executes, and keep the dispatch-result failure renderer. Rebased onto
current main; cron suite 710 passed.
2026-09-02 00:52:58 -07:00
赵桂雄
fd387c15eb fix(cron): treat falsy deliver as local in manual-run notice
Review follow-up on the #83993 fix: a stored falsy deliver ("", JSON
null) fell through the local check and produced 'output was delivered
there by the job itself' for a target that does not exist — the exact
false-delivery-claim class the PR removes. Fire time already normalizes
falsy deliver to local (no delivery, output persisted in last_output,
no delivery error), so the summary now canonicalizes with the
scheduler's own _normalize_deliver_value and reads saved-locally.

Whitespace-only deliver is deliberately not folded in: fire time
records 'no delivery target resolved' for it, and the error-driven
FAILED wording must stay visible.
2026-09-02 00:52:58 -07:00
赵桂雄
94e49b82b1 fix(cron): stop manual-run notice from asserting delivery that never happened
The _execute_job_now completion notice unconditionally claimed
"(output was delivered there by the job itself)" for non-local
delivery targets, even when the job record's last_delivery_error
showed the delivery failed (#83993). Derive the note from the
refreshed job record so a failed delivery is reported honestly to
the calling agent.
2026-09-02 00:52:58 -07:00