Commit Graph

999 Commits

Author SHA1 Message Date
Teknium
4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
teknium1
0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
teknium1
0c0875b746 chore: delete orphaned bench data, datagen examples and stale one-off docs
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.

- mcp-research-data/: 224K of July tool-search bench result rows. The
  harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
  gitignored dir; the rows were committed by hand once and the headline
  numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
  WebResearchEnv that no longer exists; the yaml paths point at a
  configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
  for a resolved bug, two RFCs whose work shipped, an unimplemented
  profile-builder proposal, the kanban dialog mock HTML and the kanban v1
  spec PDF. profile-routing.md duplicated the profile_routes section of
  website/docs/user-guide/multi-profile-gateways.md.

Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
2026-09-13 06:06:46 -07:00
teknium1
3e066dfedd fix(computer_use): approval goes through the shared gate; no callback now fails closed
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.

_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.

Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
  headless -q, plain library use): destructive actions are now REFUSED
  with a BLOCKED error and never reach the backend. Previously they
  silently ran. cron honors approvals.cron_mode, unattended platforms
  approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
  approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
  command_allowlist entry (`cua:click:background`) and is scoped to that
  action+mode — the old blanket "always_approve unlocks everything for
  the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
  "BLOCKED: Action timed out ..."); the error JSON keeps `action`.

Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
2026-09-13 05:21:02 -07:00
Teknium
d3202bbc8d feat(tui_gateway): fire agent_loop_stopped on session.interrupt too
Widens the new hook to the sibling interrupt surface: the TUI/desktop
session.interrupt path stops a live turn exactly like the gateway's
/stop, so plugins holding per-turn external resources get the same
signal there (platform='tui'). Gated on a genuinely running turn;
dispatch failures are swallowed so a plugin can never break the
interrupt. Docs updated to describe both surfaces.

Inspired by ChatGPT Work / Codex CLI 0.150.0 'Interrupt' hooks
(hooks that run when an active top-level turn is interrupted).
2026-09-12 22:19:44 -07:00
Franci Penov
f361971eed feat(gateway): fire agent_loop_stopped plugin hook on interrupt
Reapplied onto current main. The branch had drifted ~3348 commits and a trial
merge produced 48 conflict markers, so this is the same change re-landed rather
than a rebase of the old history.

_interrupt_and_clear_session interrupts the running agent without signalling
plugins, so a plugin holding a per-turn external resource — an outbound RPC
waiting on a tool result the loop will never consume — has no way to learn the
turn is gone. Dispatch agent_loop_stopped immediately after
running_agent.interrupt(), gated on a real running agent: the pending-sentinel
/stop path has no in-flight work, so firing there would be noise.

Per review on #27208, the current helper's behaviour is preserved untouched —
multiplex-aware _adapter_for_source() resolution and cached-agent eviction both
still run; the hook is additive and its dispatch failures are swallowed so a
misbehaving plugin cannot break an interrupt.

Tests fail without the change (hook registration and dispatch) and pass with
it. The three failures in tests/hermes_cli/test_plugins.py::TestPluginDiscovery
are pre-existing on this checkout and reproduce with the change stashed.
2026-09-12 22:19:44 -07:00
Teknium
ec58e08a35 feat(cron): automatic bounded re-runs when a fire never reached the model
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."

A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.

Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.

- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
  notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
2026-09-12 22:13:33 -07:00
Teknium
0c2e66ea8b test(auth): OpenRouter PKCE invariants, fake-authority A/B harness, docs, contributor map
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
2026-09-12 22:07:41 -07:00
liuhao1024
0e13fa98ec fix(web): cap web_extract provider dispatch with a wall-clock timeout (salvage #57180)
A provider whose backend keeps the response open without finishing (hanging
HTTP server, stuck SDK call) stalled the web_extract tool call — and with a
sync provider, the borrowed thread — indefinitely. The dispatch in
tools/web_tools_extract._dispatch_extract now runs under asyncio.wait_for with
web.extract_timeout (config.yaml, default 120s; 0 disables). On timeout the
tool returns structured per-URL error entries, and the one-shot keyless rescue
still gets its chance when eligible.

Salvaged from PR #57180 by @liuhao1024 (base predated the web_tools
decomposition; re-applied at the _dispatch_extract seam, env-var timeout
replaced with the web.* config section per the .env-is-for-secrets rule, and
the timeout path made rescue-aware).

Inspired by Claude Code 2.1.268: "Fixed WebFetch hanging indefinitely on a
server that keeps the response open without finishing; a fetch now fails
after 300 seconds."

Fixes #57155

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-12 21:30:38 -07:00
Teknium
101b986fdf docs(cron): document resnap alongside the drift guard
The drift-guard section only offered pin-or-disable; resnap is the middle
path (adopt the new default, stay unpinned). Same PR as the salvaged
feature per docs-in-same-PR policy.
2026-09-12 20:57:21 -07:00
Teknium
eae250534f docs(memory): explain that memory needs session boundaries on gateways
A community write-up showed a real usage pattern: running Hermes on a messaging gateway for weeks without ever issuing /new. The memory docs never said that the recall loop (MEMORY.md/USER.md snapshot + session_search) only fires at session boundaries, and that gateway chats are intentionally one continuous session across restarts. Add a section spelling out why boundaries matter and recommending /new at natural break points, cross-linked to the session-continuity docs.
2026-09-12 20:50:34 -07:00
teknium1
819988acb7 docs(kanban): per-platform opt-in is the documented path (hermes tools enable kanban --platform X) 2026-09-12 12:32:55 -07:00
Xipong
3d7f773bb4 fix(kanban): honor explicit platform tool opt-ins across configuration surfaces 2026-09-12 12:32:55 -07:00
kshitijk4poor
f003e449be refactor(cron): trim the scope-degrade dispatch to its invariants
Follow-up to the cherry-picked #102431 fix, addressing the review findings:

- The two real-helper scheduler tests ran the Linux-only helper unmarked
  and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
  scripts/check-windows-footguns.py --all (lint lane red). The remedy text
  is now built once in the helper and passed into the warning, so the
  "scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
  once (helper level); default config still Popens externally and
  `require_restart_safe_scope: true` raises (scheduler level, real helper).
  Dropped the stubbed duplicate, the standalone config-raise test and the
  in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
  with the same `except Exception` guard as the sibling
  `failure_nudge_threshold` read, so a config error no longer escapes
  the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
  instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
  `cron.require_restart_safe_scope` documented in the cron user guide.
2026-09-12 23:17:12 +05:30
teknium1
853bfec43a docs(browser): separate Chrome 136 flag refusal from the 144+ approval dialog
The salvaged paragraph described one mechanism that is actually two.
Per developer.chrome.com/blog/remote-debugging-port, Chrome 136+ simply
stops respecting --remote-debugging-port/--remote-debugging-pipe against
the default data dir — no dialog, nothing listens. The "Allow remote
debugging?" prompt the contributor saw is the separate Chrome 144+ opt-in
approval flow (chrome://inspect/#remote-debugging), which asks per
incoming connection, not per launch. Say so, cite both sources, and keep
the fix (non-default --user-data-dir, or use_real_profile) unchanged.
2026-09-12 08:30:27 -07:00
hankookdnt
d47df6ea00 docs(browser): the Chrome 136+ remote-debugging block is a consent dialog, not a silent failure
The `/browser connect` tip said Chrome 136+ "silently refuse[s] to open the
remote debugging port" on the default user-data-dir and that "there is no error
message". Chrome does surface it: on Chrome 152/macOS, launching the default
profile with --remote-debugging-port puts up an "Allow remote debugging?" dialog,
and the port opens once you press Allow. Port 9222 refusing connections is the
symptom while that consent is pending, not a permanent silent block.

Two details this cost time to rediscover, now written down:

- the consent is asked per launch, so it reappears on every browser restart
- the remote-debugging toggle in chrome://settings does not suppress it; that
  toggle only makes the feature available

Also cross-references browser.use_real_profile for the case the tip leaves
unanswered — wanting existing logins in the agent's browser. A dedicated
user-data-dir avoids the dialog but starts signed out; the real-profile snapshot
gets both, since the snapshot copy is itself a non-default user-data-dir.

Verified locally: default profile + --remote-debugging-port=9222 shows the
dialog and leaves 9222 closed, while a --user-data-dir launch answers
/json/version in under a second with no prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:30:27 -07:00
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
teknium1
44ce128a27 fix(personality): honour the top-level personalities: config block on every surface
DEFAULT_CONFIG ships a root-level `personalities: {}` (from #643) and the schema
whitelists it, but the single personality resolver read only
`agent.personalities`. A user who followed the generated config saw
"No personalities configured" from /personality on CLI, gateway and TUI.

`available_personalities()` now merges root `personalities` then
`agent.personalities` (later wins), so all three consumers pick both up.
Earlier attempt: PR #9657 (@flobo3) patched the CLI loader only.

Fixes #9636
2026-09-12 08:24:26 -07:00
Teknium
850c48cd84 feat(skills-hub): tap K-Dense and OpenScience scientific skills under one "science" bucket
~480 scientific research skills become searchable/installable through the Skills Hub
with nothing vendored: K-Dense-AI/scientific-agent-skills (165, MIT) and
synthetic-sciences/openscience (314 across 17 category paths, Apache-2.0).

A new optional tap-level `bucket` key stamps extra["category"] on every skill from a
tap whose repo ships no skills.sh.json grouping, so several repos surface as one hub
category; a sidecar grouping still wins when present. Both repos stay at community
trust (not in TRUSTED_REPOS) so the guard scans every install.

Re-grafted from #60559 onto the post-split tools/skills_hub_github.py.
2026-09-12 07:54:04 -07:00
Teknium
53c57871d6 feat(plugin-catalog): add snyk — Snyk MCP server + snyk-security-scan skill
Snyk lands as a standalone Agent Plugins v1 package
(NousResearch/hermes-plugin-snyk, pinned 2a41a07f) instead of an
optional-mcps entry: the package carries the pinned `npx -y snyk@1.1306.0
mcp` stdio launch with CLI analytics disabled AND the workflow skill that
tells the agent when to use which scanner, so a single `hermes plugins
install snyk` gives both the tools and the playbook. Third-party product
integrations ship outside the core tree per the contribution rubric.

Supersedes the optional-mcps manifest from #73860 (same pin, same
telemetry posture, tool-pruning rationale moved into the skill).
2026-09-12 06:46:17 -07:00
Teknium
be2f7e9c36 feat: curator prunes unused skills at 30 days (was 90), stale at 14
A skill nobody has loaded in a month is prompt weight, not knowledge, and
archival is recoverable (`hermes curator restore`). Defaults move
stale 30→14 / archive 90→30; config v44 rewrites only the OLD defaults so
an explicitly customized window is preserved. `hermes curator prune`
now defaults --days to curator.archive_after_days instead of a
hardcoded 90 so the manual and automatic paths agree.
2026-09-12 05:56:15 -07:00
Teknium
f364c19775 feat(plugins): touchdesigner ships as a catalog plugin (MCP server + skill), optional skill retired
The twozero TouchDesigner integration now lives in one installable unit:
plugin-catalog/touchdesigner.yaml points at NousResearch/hermes-plugin-touchdesigner
(portable Agent Plugins v1: mcp.json registers the twozero Streamable HTTP hub, skills/
carries touchdesigner-mcp). `hermes plugins install touchdesigner` + `hermes plugins
enable td` replaces the optional skill whose setup.sh hand-wrote an mcp_servers block.

The manifest name is `td` because Hermes names portable MCP tools
mcp__agent_plugin_<name>_<hash>__<server>__<tool>; twozero's longest tool under a
`touchdesigner` namespace is 71 chars, past the 64-char provider function-name cap.

optional-skills/creative/touchdesigner-mcp and its generated docs (bundled, optional,
zh-Hans) are removed; catalog tables, sidebar and kanban-video-orchestrator references
are updated to point at the plugin. Supersedes #68607 (MCP-catalog-only approach).
2026-09-12 05:15:22 -07:00
Siddharth Balyan
b7b35a84b7 docs: remove the Nous guest free-tier guide (#108991) 2026-09-12 10:06:04 +00:00
Teknium
6cd4fbd640 docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
2026-09-12 01:51:59 -07:00
Teknium
df51797e2e feat(cron): let planned downtime skip missed recurring runs 2026-09-12 01:51:59 -07:00
Teknium
f923faa0b8 feat(voice): GPT-Live voice chat mode — a full-duplex voice frontend that delegates to Hermes (Desktop)
`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS
loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and
has no tools of its own. Every real request it hears becomes a normal Hermes
turn on the open chat — any model/provider the session selected, full toolset,
memory, approvals — and the voice paraphrases the reply aloud.

Backend
- tools/voice_live.py: mode/credential/persona resolution and the one server-side
  step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer,
  pinned to client delegation; the OpenAI key never reaches the renderer. Voice
  persona follows the vendor prompting guide (role, style, labelled delegation
  policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn
  model-input note (transcript in, speakable prose out).
- REST: GET /api/audio/voice-live/status (mode + readiness, non-secret),
  POST /api/audio/voice-live/session (SDP exchange). The offer is passed
  byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF".
- prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the
  note rides the model input via the existing _prepend_note seam, the persisted
  user row stays the user's words, the system prompt stays byte-stable.
- config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*.

Desktop
- lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript
  accumulation, session.commentary/thinking/instructions appends (500-token
  chunking), mute, graceful close waiting for session.closed.
- hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation;
  delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking
  appends; reply streamed back per sentence; spoken stop phrase ends the chat;
  a newer delegation interrupts an in-flight turn.
- use-composer-voice mounts both engines and latches one at conversation start
  from the backend-resolved status; gpt-live without a key falls back to chained
  with a notice. Settings → Voice gets the mode dropdown, voice picker, persona.

Live-verified on the worktree desktop build (headless Electron, CDP, synthetic
mic): "what is 17 times 23 and which model are you on" → delegation → Hermes
(Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through
OpenRouter"; follow-up "double that" resolved from the spoken context → 782;
"run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as
quiet context → spoken kernel version; "stop" closed the session
(reason=close_requested). Chained mode creates no RTCPeerConnection.
2026-09-11 19:14:24 -07:00
Teknium
fafb27ee5c feat(plugins): on_room_member_activity hook projects Group Chat member runtime events to plugins
A hosted room member runs on a hidden room_plumbing session with no client
transport, so the tool.start/complete, approval.request, message.delta and
reasoning.delta frames its turn already emits bottom out at stdio and vanish.
Between turn.started and turn.settled in the durable room log a client sees a
black box, and community clients (Hermes Crew) cannot render tool cards,
approvals or live member status without inferring them from text.

One seam in write_json (plus the connector bypass in tool_progress) re-routes
those frames, stamped with the session's _hosted_room_task coordinates
(room_id, thread_id, member_id, turn_id, task_id, execution_generation), to a
new observer hook through the bounded per-consumer queues on_stream_* already
use, so plugin code never runs on the token path. Nothing is written to the
room log: deltas would exhaust a room's byte budget in minutes and checkpoint
replay must stay a pure function of the durable events. The task stamp gains
member_id (the driver already knows it; the proof did not carry it).

Group Chat keeps execution, scheduling and persistence; plugins own
presentation.
2026-09-11 19:04:58 -07:00
Teknium
819517fbac fix(gateway): routed profiles get their own max_turns, fallback chain, hooks, aux auth and media policy
One multiplexed gateway process serves every profile, but several per-turn
reads still went through state frozen from the LAUNCH profile:

- `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the
  module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS,
  so every secondary ran with the default profile's turn budget. A routed turn
  (HERMES_HOME override) now resolves `agent.max_turns` from its own config.
- `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide
  slot, so secondaries fell back through the default's provider/model with their
  own keys. It now reads the active gateway home and keeps a last-known-good
  chain per home.
- `_load_prefill_messages` resolved relative paths against the launch home.
- `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a
  secondary's compression/title/vision calls authenticated to Nous with the
  default profile's token when it had no pool entry. Resolved per call via
  `hermes_cli.auth._auth_file_path()` (patched constant still wins in tests).
- `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was
  loaded outside any profile scope, so secondaries' `hooks/` never ran and the
  default profile's handlers received every profile's messages, responses and
  user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the
  runner holds one registry per served home, picked from the active scope at
  emit time and front-loaded under each secondary's startup scope.
- Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME
  and the default profile's secrets). They now get the routed HERMES_HOME via
  `build_subprocess_env`, scrubbed under multiplexing, and the stdin payload
  carries `profile` so one script can tell which profile fired it.
- Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`,
  `trust_recent_files*`) was bridged once into env at startup and read from env
  per delivery; under a HERMES_HOME override the validator now reads the routed
  profile's config. Single-profile runs keep the env-bridge contract.

Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro
(temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7,
fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values.
2026-09-11 15:44:00 -07:00
Teknium
759024bdff fix(deepseek): honor Flash 1M window leftovers and native vision
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.

Two catalog gaps still produced the reported symptoms:

- A leftover context_length_cache.yaml entry of 128K (the old
  ``deepseek`` catch-all) outlived the 1M catalog keys because
  deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
  image routing falls through to lossy text. Vendor docs (2026-09-10)
  mark deepseek-flash as vision-capable and deepseek-v4-pro as not.

Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
2026-09-11 12:09:41 -07:00
Teknium
338bf9ea9a fix(cli): banner/TUI/dashboard update checks go through the GitHub API, cached 24h
The Python passive check (`banner.check_for_updates`, used by the CLI
banner, `hermes --tui`, every `tui_gateway` spawn and the dashboard's
/api/hermes/update/check) also ran `git fetch origin main` on every cache
miss, and never cached an inconclusive result so a flaky line retried on
every start. Same GitHub complaint, same fix:

- remote tip via GET /repos/{slug}/commits/main (vnd.github.sha), local tip
  via rev-parse, exact count + changelog via the compare API when they
  differ. HTTPS `ls-remote` remains only as the fallback when the API is
  unreachable or the origin isn't on GitHub.
- cache TTL 6h -> 24h, failures cached 1h; the cache is keyed on HEAD so
  `hermes update` invalidates it immediately.
- the dashboard's "what's changed" list comes from the memoized compare
  payload (`upstream_commits_behind`) instead of `git log HEAD..origin/main`,
  which was stale without a fetch.

Tests rewritten to the new contract: passive checks must not run
`git fetch`/`ls-remote` for a GitHub origin; the daily cache invalidates
when HEAD moves and re-asks after the failure window.
2026-09-10 18:15:54 -07:00
Teknium
d9ca9c974d feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped
the agent cold: the login classifier excludes one-time-code fields on
purpose (a password must never land in an OTP box) and there was no tool
for the second step, so the only move was to ask in chat.

browser_vault_enter_code
  Fills the one-time code the current page asks for. Two sources, same
  invariant as passwords (the code goes to the page over the supervisor
  socket and never enters model context):
  - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib,
    verified against the RFC test vectors), 1Password `op item get --otp`,
    Bitwarden `bw get totp`. Nobody is asked.
  - no seed: the surface prompts "Verification code for {site}"; the user
    types what their phone/email/app shows. Enter on empty / Skip declines
    and the tool returns code_declined ("do not ask again this turn").
  no_code_field tells the model the site wants a passkey / hardware key /
  app approval: hand it to the user's device and wait for navigation.
  Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order.

Surfaces
  CLI: sudo-style panel, code shown as typed (not a secret worth masking,
  typos must be visible), Enter submits, ESC/empty skips.
  Desktop: "Verification code for {site}" card via vault.code.request /
  vault.code.respond (gateway), owner-routed like the other vault prompts.
  Settings → Passwords & Logins: optional "Authenticator key" field on the
  add form (base32 or otpauth:// link); items with one show a "2FA auto"
  badge. `hermes vault add` asks for the same optional key.
  browser_vault_fill's result now says what to do next ("if the site asks
  for a verification code, call browser_vault_enter_code with this handle").
  Six locales.

Verified live (real model, local 2FA site that checks the TOTP; CLI PTY):
  A. login saved with authenticator key → signed in through 2FA, zero
     prompts, code/password absent from the transcript
  B. login without key → code panel → user types code → signed in
  C. panel dismissed → agent stops and explains, never asks in chat
Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit
spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip).
2026-09-10 11:48:01 -07:00
Teknium
438a313500 docs(honcho): document a2aSessions and per-author writes on the site page 2026-09-10 10:45:57 -07:00
Teknium
98d11c95f4 feat(vault): zero-setup UX — save a login on the page that needs it, managers auto-detected, one "Passwords & Logins" surface
Nobody should have to learn `hermes vault add` or find a toggle before "log into GitHub" works.

- browser_vault_save_login: when the agent reaches a sign-in page with no saved login it asks the user
  on THEIR surface (CLI two-step panel on the sudo modal: identifier shown, password masked; Desktop
  card with labelled Email/username + Password fields). The answer goes to the encrypted vault bound to
  the page origin and is filled at once; the model gets back only the handle and identifier. Declining
  returns save_declined; headless sessions get prompt_unavailable. Never a password in chat.
- Vault tools ride with the browser toolset (check_browser_requirements) instead of appearing only once
  the vault has items — an empty vault is exactly when save_login is needed. browser_vault_list hints
  at it when empty.
- 1Password / Bitwarden are login sources as soon as their CLI is installed; `vault.<name>.enabled`
  is opt-OUT only. Settings shows Detected/Locked/Unlocked/Off/Not detected with a switch only for
  installed managers; `hermes vault sources` reports detection, `--disable`/`--enable` flip the opt-out.
- Desktop nav/page renamed "Passwords & Logins"; empty state tells the user they do not need to add
  anything; all five locales updated. Docs rewritten from "how it works" to "say log into X".
- New per-thread SaveLoginPrompt callback (agent/vault_backends/unlock.py) installed beside the unlock
  prompt on every CLI site and the gateway bridge (vault.save_login.request/respond/expire), propagated
  to worker threads via tools.thread_context.

Live: CLI PTY (real model, packaged Chromium, local login server) — panel shown, identifier + masked
password typed, server received the correct password, password absent from terminal transcript and
from every file under HERMES_HOME outside vault/. Native Electron (headless, isolated HOME/HERMES_HOME,
own Vite + CDP port) — card shown, "Save & sign in", server received the password, Settings lists the
saved item, password absent from the rendered UI.
2026-09-10 10:35:07 -07:00
Teknium
c8998c3957 feat(vault): fill payment cards and addresses at checkout, cards behind a confirm prompt
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could
fill them: a dead surface holding real card numbers. browser_vault_fill now handles all
three kinds through the same origin-bound, supervisor-only, redacted path:

- classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens
  (cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code,
  country-name) with label/name heuristics as backup; a combined "MM/YY" control gets
  exp_month+exp_year and suppresses the split fills; inspection now covers <select>
  (country, state, expiry month) and the fill script picks an option by value or text.
- Every payment fill goes through request_elicitation_consent (gateway button round-trip
  or CLI panel) before a byte is written; declined → payment_declined, headless sessions
  are refused. A prompt injection that reaches a checkout can ask, not spend. Card values
  join the redaction registry like passwords; the result lists targeted field tokens only.
- Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the
  field) because a card without a bound origin is unfillable.
- The tool descriptions, docs and CLI copy drop "Phase 1 / login only".

Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes
nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the
country <select> untouched, and neither the card number nor the CVC appears in any result.

browser_vault_tool also: focuses the tab on the bound origin holding the right form before the
origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's
input tool" (rewritten per session by model_tools); _check_vault_available is registered
uncached because its answer is per profile (vault dir + config) and the probe is a file stat.
2026-09-10 10:35:07 -07:00
Teknium
dedc99ec6c fix(vault): ownership and race findings from the second independent review
Manager tokens: lock generation fence (a Lock acknowledged while `bw unlock`
/ `op signin` is still running discards the late token); tokens record the
unlocking gateway session and are released when THAT session ends, not when
any sibling session in the profile is torn down.

1Password: OP_CONNECT_HOST/TOKEN come from the profile's scoped secret store
like the service token (Connect outranks a service token inside op), never
from the launch environment.

Vault RPCs bind params.profile (home + secret scope) so a shared remote
backend serving several profiles locks/lists/unlocks the requested one;
unknown profile → RPC error, not a crash.

Fill target: inspection stamps are `<nonce>:<index>`; a fill resolves only
its own inspection's stamps, so an interleaved second inspection can no
longer redirect A's password into a newly mounted field (real Chrome: 0
filled, both fields empty).

Desktop Settings: every RPC goes through the owner profile's socket
(requestGatewayForProfile), query keys carry (connection, profile), an owner
change closes dialogs and wipes drafts (a master password typed for A is
never submitted to B; a late list from A never paints under B), and vault.add
secrets travel in a ref consumed by the mutationFn instead of mutation
variables. Three owner-routing invariant tests on the real component.

Docs/PR body: session-scoped release, lock-race semantics, bw --passwordenv.
2026-09-10 10:35:07 -07:00
Teknium
5aa9e7f033 fix(vault): independent-review findings — vendor contracts, profile scope, transport, target binding
Bitwarden unlock now uses the CLI's documented non-interactive channel:
`bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child
environment only (bw 2026.x rejects a piped password with "Master password
is required"). Verified against the real published binary.

Manager session tokens are keyed by (profile home, backend): a Desktop
gateway hosting several profiles can no longer reuse or lock another
profile's session. Status probes (`vault.sources`, is_unlocked) no longer
refresh the idle TTL; only real manager calls do. Gateway session teardown
locks the profile's managers (a per-session unlock ends with the session).

1Password service-account token comes from the profile-scoped secret store
(get_secret), not ambient os.environ.

`vault.source.set` no longer references a module constant (bind_module
rebinding dropped it → NameError on every Settings toggle).

Fill target binding: inspection stamps each input with a per-inspection
slot attribute; the fill resolves by stamp and requires type=password, then
strips every stamp. A DOM reflow between inspect and fill can no longer
redirect the password into a text field (reproduced in real Chrome before,
0 filled after).

Redaction boundary: no 4-char floor, CR/LF-normalized form registered
(what a text input actually stores), JSON object KEYS scrubbed in both
browser redactors; longest value first. Docs now state the real trust
model: accidental-disclosure protection, not an execution sandbox.

Desktop: the mid-turn card sends the master password through the owning
session's socket (requestForOwnedSession), never the ambient foreground
gateway; `vault.unlock.expire` clears a stale card; Settings keeps the
master password out of react-query mutation variables (ref consumed by the
mutationFn). One renderer invariant test for the routing.
2026-09-10 10:35:07 -07:00
Teknium
92e0de0ac4 feat(vault): Desktop, TUI and CLI surfaces for password-manager unlock
Desktop
- Settings → Credential Vault gains a "Password managers" section: per-manager
  toggle (disabled with a hint when the CLI isn't installed), Locked/Unlocked
  pill, Unlock (masked master-password dialog → vault.unlock) and Lock.
  Items from a manager show a source badge instead of a delete button.
- Mid-turn vault.unlock.request renders a masked card in the chat (same
  contract as the secret/sudo cards: dismiss = keep locked, late answers
  tolerated, blocks the composer, badges background sessions).
- i18n parity en/ar/ja/zh/zh-hant.

Ink TUI (hermes --tui): vault.unlock.request/expire overlay via MaskedPrompt;
Esc keeps the manager locked.

CLI: `hermes vault sources [--enable|--disable NAME]`; `hermes vault list`
shows the source column and names enabled-but-locked managers.

Docs: credential-vault.md covers managers, per-session unlock, and the
headless (cron/webhook/API/-q) no-prompt posture.
2026-09-10 10:35:07 -07:00
Teknium
cde0ad0edd feat: agent signs into sites from an encrypted local vault (CLI, browser fill, Desktop Settings)
Consolidated re-apply of #96988 onto current main. Ported from
Merit-Systems/OpenInstinct (MIT) opaque-handle autofill design: the model
sees vault handles + login metadata, the password is resolved and filled
server-side over the supervised CDP socket, and filled values are scrubbed
from every browser tool result by an unconditional redaction registry.

Rebase adaptations to the Sep-2026 facade/sibling layout:
- toolsets: one _HERMES_CORE_TOOLS entry (the browser toolset derives from it)
- hermes_cli/main.py: vault parser registered via the subcommand owner table
- file_safety: vault/ joins the _READ_DENIED_DIRS credential-dir table
- redact: registry scrub runs before the redact_secrets early-return
- browser_vault_tool: _run_browser_command now lives in browser_tool_session
2026-09-10 10:35:07 -07:00
Teknium
7e0b5cd235 fix(auxiliary): auto never bills a provider the user did not select
With a main provider selected, an unusable main route (expired xAI/Codex
OAuth token, 401/402/429 mid-session) fell through the built-in discovery
chain (OpenRouter -> Nous -> custom -> api-key) and quietly ran every
compression, title and memory-flush call on whichever OTHER account was
still logged in. Reported as "using Grok on my Premium+ sub, my Nous
Portal balance kept draining" — the chat visibly stayed on Grok while the
side tasks were billed elsewhere, and re-logging into X did not help
because the aux side never consulted the selected provider.

The discovery chain is now reserved for installs with no selected main
provider (`model.provider: auto` / unset). Otherwise the ladder is
main -> auxiliary.<task>.fallback_chain -> fallback_providers -> refuse
with a warning naming the dead provider and the fix. Both entry points
gate on the same predicate: the resolve-time route and the mid-request
payment/auth hop (_try_payment_fallback).

Existing chain tests that asserted the hop now pin `provider=auto`, the
one case where discovery is still the contract.
2026-09-10 09:11:00 -07:00
Teknium
6d89c1afde feat(plugins): pin a plugin to a commit from Desktop, and show pins in every list
Only the CLI could install a plugin at an exact commit (`--ref <sha>`); the
Desktop "Install from Git" dialog, the TUI gateway `plugins.manage install`
method and the dashboard `/agent-plugins/install` endpoint all called
`_install_plugin_core` without a ref, so a team that wanted everyone on the
same private-plugin commit had to leave the app for a terminal.

- `dashboard_install_plugin(ref=)` threads the pin to `_install_plugin_core`;
  the 40-hex validation and HEAD verification are unchanged. The gateway
  method and dashboard body accept `ref`.
- Desktop dialog gains an optional "Pin to commit" field (custom sources only,
  client-side 40-hex check disables Install on anything shorter).
- `plugins.manage list` rows carry `pinned_sha`; the Desktop plugins tab shows
  a `pinned @ <sha8>` badge and `hermes plugins list` prints `git pinned@<sha8>`
  in Source, so a team can eyeball that everyone runs the same commit.
2026-09-10 05:09:15 -07:00
kshitijk4poor
4a76e99f87 docs(cron): drop stale mirror-scope wording left by the home lane
- `_target_matches_origin` docstring still claimed fan-out targets are
  "deliberately NOT mirrored" and mirroring is origin-only; that has been
  false since origin_fallback landed and is doubly so with the home lane.
  Point to `_target_mirror_eligible` for the policy instead.
- cron.md DM-only bullet said "origin DM session"; the config comment in
  the same change already says "target DM". Align.
- Remove two `_expand_routing_tokens` unit assertions bolted into the
  eligibility test; the `" ALL , slack "` parametrize covers expansion
  end-to-end through `_resolve_delivery_targets`.
2026-09-10 15:32:47 +05:30
Victor Kyriazakos
22356075f1 fix(cron): make user-written bare-platform home targets continuable
A managed cron with `deliver: "slack"` (no captured origin) delivers to the
Slack home channel, but the brief was never mirrored into that session and
the in_channel seed never fired: bare-platform targets resolved with no
`_resolved_from` provenance, so `_target_mirror_eligible` treated them like
`all` broadcast expansions.

- `_resolve_delivery_targets` derives `from_broadcast` from the raw token
  and passes it to `_resolve_single_delivery_target`; a user-written bare
  platform token is tagged `_resolved_from: "home"`, `all` expansions stay
  untagged (fan-out is never continuable).
- `_target_mirror_eligible`: `home` eligible under the same flags as
  `origin_fallback` (per-job attach_to_session wins, else cron.mirror_delivery).
- `_MIRROR_PROVENANCE_RANK`: `home` == `origin_fallback` so token order
  through dedup cannot strip eligibility.
- Docs, tool schema text, cron/gateway AGENTS.md updated; existing
  exact-dict pins carry the new key.

Squashed from PR #101819 (commits 5fe612c9, 65e36926, 10e0af1d, f2cc4639):
they predate the cron/scheduler.py -> scheduler_delivery.py split and do
not cherry-pick individually onto current main.
2026-09-10 15:32:47 +05:30
Teknium
61afcde8f9 refactor(desktop): one Plugins surface — Capabilities → Plugins owns agent + desktop plugins, install, and the catalog
Plugins were split across two pages that each showed half the picture:
Settings → Plugins listed desktop plugins plus "Install from Git" and a
pointer saying agent plugins live elsewhere; Capabilities → Plugins listed
agent plugins plus the catalog picker but knew nothing about desktop
plugins. A user asking "what extends my Hermes and where do I add more?"
had to visit both and still could not see the whole set in one place.

Capabilities → Plugins is now THE plugins page:

- Agent plugins section (scoped to the profile selector) with the
  "Install from Git" button in its header — installs target the scoped
  profile, not whichever one is active.
- Desktop plugins section beneath it (same for every profile), with the
  folder/rescan controls and the "agent half missing here" drift chip,
  whose repair also lands in the SCOPED profile.
- The catalog picker underneath, unchanged.

Settings → Plugins is removed. `/settings?tab=plugins[&plugin=…]` and the
existing `?tab=mcp` redirect share one table (`settings/moved-tabs.ts`) so
old bookmarks and palette links land on the same row on the new page.
Command palette: plugins moved from the Settings group to the Capabilities
group; installed-plugin rows deep-link to `/skills?tab=plugins&plugin=…`.
Dead `settings.plugins.agent.*` and `settings.nav.plugins` i18n keys dropped;
docs and in-code pointers say Capabilities → Plugins.
2026-09-10 02:30:50 -07:00
Teknium
9886f6e53b feat(plugins): install plugins from private git repos using the user's stored credentials
`hermes plugins install owner/private-repo` failed for every private repository
even when the same user could `git clone` it from their shell: the hardened
`noninteractive_git_env` disables credential helpers, askpass and global config
(so a hostile repo can't make our plumbing prompt or hang), which also blocks
the user's own stored credential. The result was "could not read Username" or a
60s hang on a GUI askpass, with no hint about how to authenticate.

New `hermes_cli/git_credentials.py` resolves a credential up front from sources
the user already owns — GITHUB_TOKEN/GH_TOKEN (profile-scoped), `gh auth token`,
then `git credential fill` against their configured helpers with prompting
disabled (any host) — and hands it to git as a one-shot
`http.<origin>/.extraheader` via the GIT_CONFIG_* env block. Nothing lands in
the URL, `.git/config` or install metadata. The same path covers
`plugins update`, catalog MCP git installs and profile-distribution staging,
which share the same hardened env and the same failure.

A private-repo clone with no credential now fails fast with an actionable hint.
2026-09-10 01:00:19 -07:00
Siddharth Balyan
b4d04eb8fd Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding

* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg

The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.

On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.

* docs(tool-search): connectors section — remote tools through the bridge

The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.

* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots

dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.

The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.

The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).

Live, 311 local tools + gateway, before -> after:
  "send gmail email":           5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
  "read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
  "linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.

* refactor(tool-search): connector leg into tools/connector_search.py

tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.

No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.

* fix(tool-search): at most 7 queries per call, the gateway's search limit

One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.

The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.

* fix(tool-search): the model is told that connectors__ names are manage_connections accounts

tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.

The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.

Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.

Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.

* fix(connectors): /stop halts a connector batch before the next remote call

dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.

The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.

Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.

* test(connections): schema assertions become dispatch contracts

test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.

Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.

Test count in the file goes from 26 to 25.

* docs(tool-search): connector batches are one gateway request per entry

The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.

Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.

Docs only, no test.

* fix(tools): the between-turns refresh never rewrites the bridge tools

The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.

The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.

Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.

* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation

The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.

* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote

bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.

The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.

Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.

* fix(connectors): search keeps the twin a colliding name reaches, and says so

format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.

Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
2026-09-10 02:21:16 +05:30
Siddharth Balyan
cf4b78e91f fix(tool-search): a query no tool answers returns nothing, not five tools sharing one word (#106676)
search_catalog admitted every document with BM25 score > 0 and then padded
to `limit`. BM25 sums over the tokens a document shares with the query, so
on a 300-tool catalog "send gmail email" returned five incident tools that
shared only "email", and the discriminating word ("gmail", in no document)
had no say. The model read those as the answer.

Admission is now the query's rarest token: a document is a result only if it
contains the query token with the highest IDF, the one that names the intent.
Common verbs ("send", "read", "create") sit in dozens of documents and never
gate; vendor and object words ("gmail", "github", "incident") do. A token no
document carries admits nothing, and the existing empty-group hint tells the
model to retry without it. The name-substring fallback is deleted: it admitted
tools that matched no query token at all.

Result descriptions are clipped at 500 characters instead of 400. Over 353
vendor tool descriptions, 500 keeps 91% whole and every first sentence
(first-sentence max 329); 400 kept 82%.

Measured on the live 311-tool catalog with 25 hand-labelled queries:
precision@5 0.18 -> 0.43, wrong names returned 102 -> 66, false positives on
absent intents 17 -> 13. Live before/after: "send gmail email" went from five
betterstack tools to an empty group with the retry hint; "linear create issue"
and "betterstack incident" are unchanged.
2026-09-10 02:21:15 +05:30
Teknium
d5926b2494 fix: persist API delegation units once without waking the model 2026-09-09 10:55:32 -07:00
teknium1
322905e91b feat(memory): make the mem0 sync char cap configurable via mem0.json
A flat 450-char cap fits 512-token embedders (bge-small-zh-v1.5,
all-minilm) but stores only ~5% of the window on 8192-token models
(text-embedding-3-small, jina-embeddings-v3, bge-m3), degrading memory
quality for users those models served fine before truncation existed.

Read `sync_max_chars` from mem0.json once in initialize() (450 default)
and pass it to _truncate_for_sync(). Config over auto-detection: the
Ollama /api/show probe + known-model table proposed in #37427 adds a
network call and a curated list for a number the operator already knows
from their embedder choice; the setup wizard's mem0.json is the plugin's
behavioral-settings surface (no new HERMES_* env var). Documented in the
plugin README and the memory-providers docs page.

Dynamic-cap requirement and measurements (450 OK / 600 -> HTTP 500 on
bge-small-zh-v1.5:f16) by @szicely in #106235.

Refs #37421 #106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-09 10:55:02 -07:00
Teknium
e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Teknium
41b4555ed9 feat(plugins): catalog is the sole discovery system — re-port onto main's layout
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
  .hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
  update, and the dashboard/TUI payload builders. plugins_cmd.py only
  gains the hooks (cmd_install catalog branch, cmd_update / dashboard
  update re-pin, dashboard_install_plugin catalog_name + kill list,
  dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
  test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
  published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
  fallback) instead of the unauthenticated GitHub contents API (60 req/h,
  1 request per entry); in-tree and live removals are unioned so a stale
  cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
  limits); _plugin_runtime_status shared from web_server_dashboard.py;
  hub rows carry removed_reason. TUI plugins.manage gains catalog_name
  install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
  (real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
  first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
  plugin-catalog/** so entry merges republish it.
2026-09-09 04:38:01 -07:00