Commit Graph

2593 Commits

Author SHA1 Message Date
teknium1
4748caff76 fix(gateway): explicit tool_progress new/all keeps text progress in un-cardable Slack chats
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.

Test proven red on the salvaged head (adapter.sent == [] with `all`).
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
e9c037f65a docs(slack): clarify progress resolution and fallback lifetime 2026-09-14 07:46:51 -07:00
Victor Kyriazakos
bc125d59d5 fix(gateway): null tool_progress inherits; name the task-card suppression latch
Review findings (Salt, adversarial pass on the two preceding commits):

- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
  counted as an explicit mode because the gate tested key presence, while
  the display resolver skips None and inherits. Null resolved to Slack's
  tier default `off` and disabled cards, which is the default-off trap the
  change exists to avoid. Explicit intent is now a non-None value (or the
  env bridge). Tests cover null at each level plus null-over-global-all;
  mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
  destinations, so the name no longer described the field. Renamed to
  `publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
  described the opt-in as independent of tool_progress. Rewritten: cards
  follow an operator-written off (including /verbose), null inherits, an
  un-threaded chat with the card lane active shows no tool progress, other
  native failures keep the editable fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
3412490ad1 fix(gateway): no text tool progress when a Slack chat cannot host a task card
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.

Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
ed25a40917 fix(gateway): explicit tool_progress off disables Slack task cards
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.

Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.

Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
2026-09-14 07:46:51 -07:00
teknium1
9f7f2f28c0 feat(gateway): server→client JSON-RPC requests replace the *.request/*.respond event pairs (#110521)
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.

`tui_gateway/server_requests.py` owns the one mechanism:

  send()          block the agent thread until the response frame
                  (`srq-<n>` ids; ints belong to the client)
  send_async()    fire-and-callback variant (bot relay)
  cancel*()       withdraw with ONE `request.cancel {id, method, reason}`
                  event (timeout / interrupt / process exit /
                  answered elsewhere) instead of per-kind *.expire
  open_requests() the still-open frames, replayed by session.resume,
                  session.activate and session.events.since so a
                  reconnecting client re-renders every kind, not two
  clarify.lock    stays a real client→server RPC (locks one batch
                  answer early); locked answers merge into the final
                  set even when the closing response carries only the
                  tail the user answered last

A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.

Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.

Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
2026-09-14 06:02:05 -07:00
teknium1
7d368c7d2c docs: background_review.reasoning_effort applies on the routed path; same-model warns once 2026-09-14 05:25:01 -07:00
teknium1
ee51e8bf8b chore(webhook): drop external-product attribution from code, tests and docs 2026-09-13 21:30:02 -07:00
Teknium
45ab5e3b5d Inspired by ChatGPT Work: event-triggered cron jobs via webhook routes
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.

- gateway/platforms/webhook.py: cron_job route mode — after the same
  HMAC auth / rate limit / filters / script / idempotency as agent
  routes, the rendered prompt becomes transient per-run context and the
  job fires through execute_job_for_event on a worker thread (202
  Accepted immediately). Startup validation rejects cron_job +
  deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
  the shared claimed-run body (_execute_job_now), so event fires share
  at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
  handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
  subscriptions; job ref validated (and canonicalized to the job ID) at
  create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
  cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
  CLI tests in test_webhook_cli.py.
2026-09-13 21:30:02 -07:00
teknium1
c8e155dcbf fix: ai-presenter-video resolves its dir via ${HERMES_SKILL_DIR}, regen docs page
The shell `find ~/.hermes/skills ~/.hermes/hermes-agent/optional-skills ...`
snippet hardcoded a dev-clone path that does not exist on user installs; the
loader already substitutes ${HERMES_SKILL_DIR} (skills.template_vars), which
is what every other bundled skill uses. Drop the upstream-agent path mention.
Regenerated the docs page so it matches SKILL.md (also fixes the comfyui
related-skill link, which now lives under optional).
2026-09-13 21:12:23 -07:00
Teknium
a79ff58d65 feat(skills): ai-presenter-video optional skill (port of lanshu, 955★ MIT)
Ports cclank/lanshu-create-ai-presenter-video (MIT, 955 stars in 7 days)
into optional-skills/creative/ai-presenter-video. Provider-neutral
presenter-video production: locked narration as master clock, avatar
generation with pilot-first cost discipline, lip-sync/identity QA,
captions, deterministic ffmpeg finalization with loudness normalization
and contact-sheet verification.

Hermes adaptations in the hub SKILL.md: SKILL_DIR resolution (upstream
hardcoded ~/.codex/skills), capability mapping to text_to_speech / FAL
video families / vision_analyze / hyperframes, consent-flag JSON paths
(input.* vs root), preflight error-vs-remote-blocker semantics.
References kept substantively verbatim (all-English upstream). Scripts
unmodified. LICENSE carried.

Validated hands-on: init_job -> preflight gating (blocked until manual
review booleans + input.remote_upload_approved) -> finalize_delivery on
a synthetic 1080x1920 render (master+share decode-verified, delivery
report, 9-frame contact sheet). Cold-subagent live test: SHIP; 3
friction fixes folded in (resolution guard, boolean-flip example,
preflight-writes-job note).
2026-09-13 21:12:23 -07:00
Teknium
56cc2bd814 feat(skills): scrollcraft — premium scroll-driven landing pages (port of nateherkai/scroll-craft, 1.2k★ MIT)
Optional skill: scroll-as-timeline landing pages on a deterministic
CSS/JS engine, with interview → page grammar → signature move workflow
and screenshot-based scroll verification. Engine and scripts vendored
verbatim; asset generation re-anchored on image_generate with the
upstream kie.ai flow kept as an optional path.
2026-09-13 21:11:25 -07:00
teknium1
230ca004a7 fix(delegate): forward inline data-URL images to vision children; trim tests to invariants
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.

Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
2026-09-13 21:05:42 -07:00
Teknium
f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
Teknium
c63de5a231 feat(tool_search): long hunts for nonexistent tools now return no results instead of incidental matches
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).

A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.

Docs: relevance-floor bullet added to tool-search.md implementation
details.
2026-09-13 21:04:08 -07:00
Teknium
3304d205be feat(video-gen): Kling 3.0 Standard + Pro families on the FAL backend
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
2026-09-13 21:01:32 -07:00
Teknium
65a6b6831e Port from code-yeongyu/oh-my-openagent#7151: worktree audit gains --json, --older-than, and external-tree visibility
omo's omo-agent-toolkit worktree-sweep (their PR #7151) added three
capabilities our hermes worktree command lacked:

- --json on list and prune: machine-readable audit/result payloads so
  scripts and agents can consume verdicts without scraping table output.
- --older-than DAYS: an age floor that only ever RESTRICTS reaping
  (young-but-reapable trees are kept); it never widens eligibility, so
  the existing safety invariants are untouched.
- External-tree visibility: linked worktrees registered outside
  .worktrees/ are now reported read-only in the audit (branch, locked,
  missing) instead of being invisible, and registrations whose
  directory has vanished are dropped via git worktree prune (metadata
  only, no files touched) during prune.

Not ported: omo's ancestor-of-default-branch merge test (our git cherry
patch-equivalence is strictly stronger under rebase/squash merges), and
their hardcoded external-root exclusion list (we exclude by location:
everything outside .worktrees/ is hands-off).

Tests: 9 new contracts in tests/hermes_cli/test_worktree_gc.py (age gate
restrict-only, external trees never reaped, stale-registration prune
dry-run/real, JSON shapes, negative --older-than rejected). Live E2E on
a scratch repo verified all three flags end to end.
2026-09-13 20:55:29 -07:00
teknium1
cb3447b139 test: trim the context-cache guard tests to invariants; docs: name the surfaces that actually confirm
Tests collapse 13 change-detectors into 7 invariants (silent below threshold /
without context, fires above, same-model re-select silent, config override and
0-disables, registry threading, agent context derivation). The legacy 5-arg
guard test goes with the TypeError fallback it covered: that fallback was
defence for a case nobody has (every in-tree guard and test double is
*args-tolerant) and would re-run a guard whose real TypeError it masked, so the
rebased port passes the context positionally like every other argument.

Docs no longer claim the confirm fires on the Telegram/Discord pickers or the
dashboard: those surfaces call combined_selection_warning() without a live
agent, so the context-cache guard is (correctly) silent there.
2026-09-13 20:54:50 -07:00
Teknium
8b6931393e Port from langchain-ai/deepagents#5829: confirm mid-session model switches that abandon a large cached context
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).

- hermes_cli/model_selection_guards.py: new context_cache guard +
  SelectionContext carrier + selection_context_for_agent() helper;
  registry threads live-session facts to guards (6-arg signature with a
  TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
  live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
2026-09-13 20:54:50 -07:00
Teknium
d9e88e19e2 feat(mcp): bind stored OAuth refresh tokens to their issuer
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.

- HermesTokenStorage records hermes_issuer alongside cached tokens
  (stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
  tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
  on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
  unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
  (no forced re-login for existing installs — deliberate divergence
  from Codex, which requires reauth).

Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
2026-09-13 20:43:33 -07:00
Teknium
ce318290bd Inspired by Factory Droid: /queue prompts are now listable, editable, and reorderable before they run
Droid v0.203 (Aug 25 2026) added 'edit queued messages' — a queued steering
message can be pulled back and changed before it is sent. Hermes /queue could
only append blindly: no way to see, fix, drop, or reorder queued prompts.

/queue now supports management subcommands in the CLI:
- /queue            — list pending prompts (bare prompt still enqueues)
- /queue list       — same
- /queue edit N <p> — replace item N (keeps voice sentinel, #65827)
- /queue rm N       — remove item N
- /queue move A B   — reorder
- /queue clear      — drop everything
- /queue add <p>    — force-enqueue prompts starting with a management word

Queue mutations hold queue.Queue's mutex and rebuild unfinished_tasks so
join()/task_done bookkeeping stays consistent. Paste references expand on
enqueue and edit, matching the old inline path.

Reimplementation of PR #18833 by @abhinav11082001-stack (commit was authored
under a fabricated 'Hermes Agent' noreply identity that cannot be carried
into history; engineering credit is theirs), hardened for current main:
voice-sentinel-aware previews/edit, paste-reference expansion, queue
bookkeeping asserts, out-of-range no-op tests, and docs.
2026-09-13 20:41:41 -07:00
Teknium
48bd70b586 Port from nearai/ironclaw#7378: doc-fact contract test keeps slash-commands.md in sync with the command registry
Two-direction contract test (tests/website/test_slash_commands_doc_parity.py):
every CommandDef must be documented under its name or an alias, and every
doc table row must resolve to a registered command. Ported from IronClaw's
doc-fact contract tests (nearai/ironclaw#7378), adapted from their clap
--help parser to our COMMAND_REGISTRY single source of truth.

Real drift it caught, fixed here: /loop (alias /proactive) shipped with a
full feature page (user-guide/features/loops.md) and CLI+gateway handlers
but never got a row in the slash-commands reference. Added to both the CLI
Session table and the messaging table, plus the both-surfaces note.
2026-09-13 20:40:45 -07:00
xxxigm
1468e98e48 test(google-chat): pin hosted install wiring and document the lazy target 2026-09-13 19:19:36 -07:00
teknium1
1d33a4fee6 feat(desktop): persist auxiliary reasoning_effort through the models router
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.

The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.

Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.

Closes #89259. Salvages #90649 by @higgs1729.
2026-09-13 19:19:32 -07:00
teknium1
2c0bec33f9 feat(model-pickers): reasoning effort selection on every model picker
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.

One request now carries a model pick AND its effort on every surface:

- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
  <level>` (validated against `parse_reasoning_effort`; unknown level ->
  `MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
  `ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
  high` applies the effort AFTER the agent swap (`switch_model` re-resolves
  `reasoning_config` from config.yaml, so an earlier write is clobbered) with the
  pick's scope (session; config on `--global`; `--once` snapshots and restores it).
  The `/model` picker gains a third stage, "Reasoning effort for <model>", built
  from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
  inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
  `config.set model "X --reasoning high"` applies after the swap; session pin
  (`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
  one-turn restore carries `reasoning_config`; re-emits `session_info` so the
  status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
  `<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
  label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
  `_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
  the Copilot-only inline prompt; Copilot keeps its per-model level set via
  `github_model_reasoning_efforts`, other routes get the ladder, catalog
  `supports_reasoning=False` skips it) plus a "Reasoning effort for the current
  model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
  with the same step (+ "Provider default"), stored as
  `auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
  the task list ("openrouter · model · high"), cleared by "Reset all to auto";
  tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
  it.

Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
  "Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
  and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
  writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
  errors "Model names cannot contain spaces"; after switches and `config.get
  reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
  effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
  "reasoning: high", status bar "fable 5.1 high".
2026-09-13 16:43:50 -07:00
teknium1
2dfd831d3b fix(webhook): bind a subscription to a profile with --route-profile, not --profile
The salvaged flag was spelled --profile, which collides with the global
-p/--profile that hermes_cli.main scans BEFORE argparse: `hermes webhook
subscribe x --profile compta` would switch this CLI process to compta's
HERMES_HOME and write the subscription into compta's webhook_subscriptions.json
— a file the default gateway's webhook adapter never reads — while the route
still lacked the profile key. #109020 special-cased the scanner for the webhook
subcommand; naming the flag --route-profile removes the ambiguity without
touching _scan_profile_flag: -p picks the gateway whose subscriptions file is
written, --route-profile picks which /p/<profile>/ prefix may hit the route.

Docs: cli-commands reference row, multi-profile-gateways webhook section, the
route `profile` field. Builds on #109020 (fangliquanflq). Fixes #109016.
2026-09-13 15:45:57 -07:00
teknium1
6636b0896c docs(multiplex): OAuth and mTLS servers are never shared; trust and parallel policy are per profile
Extends the multi-profile MCP paragraph with the identity rules landed in
this branch: OAuth tokens live per profile so each profile's calls run as
its own account; client_cert/client_key are part of the connection
identity; trust and supports_parallel_tool_calls are the consuming
profile's policy even when it shares another profile's connection.

Wording of the per-profile OAuth account guarantee follows the docs draft
in PR #109574 (its token-file fingerprint code path was not taken).

Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
2026-09-13 15:41:01 -07:00
teknium1
ebe11403c6 fix(gateway): a profile named 'main' gets its own session namespace
`main` is a valid profile name (only hermes/default/test/tmp/root/sudo are
reserved), but _session_key_namespace mapped it to `agent:main` — the default
profile's namespace. Both profiles then built byte-identical keys: one routing
entry, one cached agent, and, since 75ae2859b9 pinned default-namespace
keys to the launch store, profiles/main's scoped sessions were written into
the ROOT state.db instead of profiles/main/state.db.

Key the `main` profile as `agent:main~` (`~` is outside the profile-id
alphabet, so the marked form cannot be any other profile's id) and give the
namespace slot one inverse, profile_from_session_key_namespace, used by the
store's key parser, _parse_session_key, the update-marker profile reader and
the profile-delete eviction prefix. Default keys stay byte-identical.
2026-09-13 15:41:01 -07:00
teknium1
8c06594c40 docs: state the per-profile setting precedence rule (env → own YAML → default)
Multi-profile guide gains the rule and the consumers it covers; the adapter
authoring guide and the Slack allow_bots page no longer claim YAML wins.
2026-09-13 15:39:11 -07:00
Tim Smykov
bfbf31d576 fix(gateway): polish background process notifications
The raw-output watcher modes (all/result/error) and the interim running
update sent the bracketed debug wrapper with the internal process id
(`[Background process proc_… finished with exit code N~ Here's the final
output: …]`) to Telegram/Discord/Slack chats. Reuse the concise one-line
status header for every mode and append the bounded, ANSI-stripped output
tail in a code block; the running update gets the same shape.

Salvaged from #54266 (rebased onto the post-#102117 run_notifications
sibling; the concise mode had landed in between, so the header is shared
rather than reimplemented). Also covers #13122 (ANSI stripping).
2026-09-13 15:05:11 -07:00
teknium1
0abfd1105c fix(migrate): hermes update refuses to fold cross-user / cross-scope gateways; auto_multiplex_migration opt-out
Reshape the two salvaged commits onto current main (#109954):

- Move the boundary guard out of the gateway_migrate facade into a new sibling
  hermes_cli/gateway_migrate_guards.py as a table of guard functions
  (_AUTO_MIGRATION_GUARDS: service domain, UNIX user, HERMES_HOME tree) plus the
  identity resolver. The facade grows by ~20 lines only (uid/runtime_home on
  ProfileGateway, one seam, the hook wiring).
- Compare uids, not strings: live pid owner via /proc (ps fallback only on
  macOS, where /proc does not exist), else the system unit's User= via
  _read_systemd_user_from_unit (root when absent), else the home directory's
  owner. None means unknown and never blocks.
- The home-tree guard reads the HERMES_HOME the installed unit pins, not the
  directory the plan enumerated: that is where the gateway really runs and is
  exactly the "stale copies under profiles/" shape from the report.
- When the default is detached, a service-managed secondary is a different
  domain for the AUTO path (it must not elect the secondary's manager); the
  explicit command keeps electing it as before.
- The explicit command surfaces the same findings as notices (dry run shows
  them) and is never blocked by them; only the update hook refuses.
- Rename the opt-out key to gateway.auto_multiplex_migration (nested only, no
  top-level alias) and read it before a plan is built, so false prints nothing
  and touches nothing. The explicit command ignores it.
- Tests trimmed to the invariants: one parametrized boundary test that exercises
  the real hook end to end (refuses, touches nothing, dry run shows the notice),
  one "same user / same scope still migrates" control, one opt-out test.
- Docs: boundary table + renamed opt-out section in multi-profile-gateways.md;
  one line in hermes_cli/AGENTS.md.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Athena <athena@olympus.local>
2026-09-13 14:48:09 -07:00
Athena
8ed2fec94c feat(gateway): let an install opt out of the automatic multiplex migration
`hermes update` folds an eligible multi-profile install onto one multiplexed
gateway on its own, and there is currently no way to say no. The only lever,
`gateway.multiplex_profiles: false`, is also the default: `_read_multiplex_flag`
returns `False` for "absent" and for an explicit `false` alike, so an operator
who has already decided to stay on per-profile gateways has no way to record
that decision. The migration runs again on the next update.

Add `gateway.auto_migrate` (bool, default `true`). Read from the default
profile's config, it gates the automatic path only:

- absent or `true`   -> today's behaviour exactly, no change
- `false`            -> `maybe_auto_migrate_after_update()` returns before
                        building a plan; no output, no changes

`hermes gateway migrate --multiplex` is an explicit request and still migrates
regardless of the flag, so it stays the supported way to opt back in.

One early return, one schema entry with the reasoning inline, one invariant
test (opt-out blocks the hook, absent/true do not, explicit command still
applies), one section in the multi-profile gateways guide.
2026-09-13 14:48:09 -07:00
teknium1
e30d0639bb fix(cli): sessions export accepts a directory for single-file formats
`hermes sessions export --session-id X <dir>/` crashed with IsADirectoryError
because jsonl/html/trace opened the positional as a file while --help called it
an "output path" and md/qmd really do take a directory. An existing directory
(or one spelled with a trailing separator) now receives a default-named file
(`hermes_session_<id>.<fmt>`), and the help text spells out per-format what
OUTPUT means.
2026-09-13 14:45:10 -07:00
teknium1
8abe6ab8ff docs(a2a): say the orphan sweep follows A2A_REPLY_TIMEOUT and live waiters
The troubleshooting entry told users to raise A2A_REPLY_TIMEOUT for long tasks,
which did nothing against the hardcoded 300s orphan sweep (#106972). Now that the
sweep derives its grace from the reply window and skips tasks with a live waiter,
state that contract next to the variable.
2026-09-13 14:43:35 -07:00
fangliquanflq
2710d85914 fix(kanban): gate gateway notifier polling 2026-09-13 14:43:19 -07:00
teknium1
1e76efbe28 fix: scan plugin test trees again, cap their criticals at caution
Skipping `tests/`, `spec/`, ... in EXCLUDED_DIRS made those trees
invisible to the guard, but `plugins_loader._load_directory_module`
sets `submodule_search_locations=[plugin_dir]`, so a plugin
`__init__.py` doing `from .tests import evil` imports and runs whatever
lives there: a `tests/evil.py` with a destructive root remove scanned
`dangerous` on main and `safe` on this branch. `_walk` also matched the
names at any depth, so `src/spec/handler.py` — plain runtime code — went
unscanned.

Keep scanning everything; instead cap a critical finding located under a
ROOT-level test dir at `high`, so the verdict is `caution` (confirmation
required, `--force` overridable) rather than the un-overridable
`dangerous`. Fixture strings still cannot brick an install, which was
the reported problem, while a critical in any runtime file (`setup.sh`,
`src/spec/...`) still yields `dangerous`. Trade-off stated in the PR
body: hostile code deliberately placed under `tests/` is now
force-installable rather than blocked outright.

Docs no longer claim test code never runs.
2026-09-13 14:43:04 -07:00
teknium1
713270d3d0 docs(plugins): document skipped test trees and the critical-finding block reason
User-visible scanner behaviour changed in this PR (test trees skipped, the block
reason names the critical rule ids), so the plugin docs say so in the same PR.
2026-09-13 14:43:04 -07:00
teknium1
6de2dde61f docs: rename under a live multiplexer unroutes the old name
The served-set paragraph documented create and delete as live operations; rename now
follows the same unroute-before-mutate protocol, so say so where operators look for it.
2026-09-13 14:40:00 -07:00
teknium1
630a4eb3a1 docs(mcp): mTLS credentials count toward connection sharing; OAuth token path is per profile
The multiplex guide now says client_cert/client_key are part of the
"same credentials" test and states the OAuth rule as its own sentence;
the MCP config reference names the per-profile token directory and the
never-shared-across-profiles rule next to the OAuth behaviour list.

Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
2026-09-13 14:38:20 -07:00
kshitijk4poor
1b27f8ac4e docs(mcp): OAuth servers are never shared across profiles 2026-09-13 14:38:20 -07:00
teknium1
e10d30f8a8 docs(kanban): review handoffs also stage declared artifacts
`kanban_request_review(artifacts=[...])` now preserves scratch deliverables
the same way `kanban_complete` does; say so where the scratch-workspace
lifecycle is documented.
2026-09-13 14:37:17 -07:00
teknium1
b47fb2eba3 fix(profiles): clone builds in a hidden staging dir, never writes through symlinks, refuses --clone-channels in core; channel inventory is ownership-based
Post-merge review of #109502 (gaoanze888) on current main, findings 1-3 and 5-11
(finding 4, the migrate manifest ordering, was already fixed by f9e47aa6fe).

Why:
- `--clone-all` used copytree(symlinks=True); a symlinked source `.env` was then
  edited THROUGH the link by the channel strip, deleting the SOURCE's bot token.
  Root files the clone edits (.env, config.yaml, auth.json, SOUL.md) are now
  materialized as private copies before any write.
- The final profiles/<name> existed during the copy; the multiplexer rescans
  profiles/ on create (#109239) and could adopt the half-copied tree and start
  adapters on credentials not yet stripped. Clones are built in
  profiles/.<name>.staging-<pid> (a leading dot never matches _PROFILE_ID_RE, so
  profiles_to_serve never lists it) and published with one os.rename after the
  strip; a failed create removes the staging tree.
- The live-multiplexer refusal for --clone-channels lived only in the CLI; REST
  (POST /api/profiles) and the TUI (profiles.create) bypassed it. It now lives in
  create_profile as profile_channels.clone_channels_refusal, raising ValueError
  which every surface already maps to a 400 / 4062. --clone-channels without a
  clone flag is an error instead of a silent no-op.
- Channel inventory is ownership-based, evaluated in the SOURCE profile's plugin
  scope: private platform plugins under <source>/plugins/ contribute their keys
  (previously discovered under the ambient HERMES_HOME); GATEWAY_ALLOW_ALL_USERS /
  GATEWAY_ALLOWED_USERS and GATEWAY_RELAY_ID/SECRET/DELIVERY_KEY are channel
  settings; alias prefixes SUPPLEMENT the canonical <PLATFORM>_ prefix
  (WECOM_DM_POLICY, SMS_WEBHOOK_PORT now stripped). Prefixes shared with tools
  (HASS_, TWILIO_, EMAIL_) are stripped only when the source's gateway would run
  that adapter (enabled in config, or complete credentials and not explicitly
  disabled); their allowlist/port keys are always channel-only.
- --clone-all state removal handled files only; Google Chat's directory-shaped
  google_chat_user_tokens/ survived. Directories are removed too, without
  following a copied symlink.
2026-09-13 14:33:11 -07:00
teknium1
d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
kshitijk4poor
8dd0e80f29 fix(discord): drop the unreachable frame-silence dimension, keep the config warning
The cherry-picked commit added an `event_silence` probe dimension stamped from
`on_socket_raw_receive`. Two verified problems make it a regression rather than a fix:

- discord.py 2.7.1 dispatches `socket_raw_receive` only when the client is built with
  `enable_debug_events=True` (client.py:330, gateway.py:410-412; the default
  `log_receive` is a no-op). The adapter never sets it, so the stamp only ever moves at
  `on_ready` and every healthy connection reads `event_silence` 300s later — a forced
  reconnect every ~5 min. Live-verified against a real `commands.Bot` +
  `DiscordWebSocket.received_message`: 6 frames delivered, stamp unchanged, probe unhealthy.
- discord.py already keeps a per-frame clock (`KeepAliveHandler._last_recv`) and closes the
  socket itself after `heartbeat_timeout` without frames; and because ACKs are frames,
  `ack_stale` (60s) always trips before `event_silence` (300s). A raw-frame stamp cannot
  detect the "ESTAB + ACKing + zero events" incident by construction.

Kept and tightened the warning half: bool values (`float(True) == 1.0` silently enabled a
knob at 1s), negative ints, and unparsable strings now warn; an explicit `0` is the documented
opt-out and stays silent. Tests trimmed to the two invariant contracts (warn / don't warn),
proven red on origin/main. Docs updated to match.
2026-09-13 20:03:40 +05:30
salch-cred
cef499fbb5 fix(discord): liveness probe gains a dispatch-side dimension (#109521)
The Gateway WS health probe sampled only transport state — ready, open,
heartbeat-ACK age, latency. A socket that stays ESTAB and keeps ACKing while
zero gateway frames arrive (the #109521 "connected-but-deaf" incident) read
healthy indefinitely, and the adapter went silent for hours with no log line
and no watchdog firing.

Two defects fixed:

1. Dispatch-side dimension. `on_socket_raw_receive` now stamps
   `_last_gateway_frame_at` for every inbound raw gateway frame — heartbeats
   and ACKs included, so a legitimately quiet server is not flagged. The
   health check gains an `event_silence` reason with its own bound,
   `websocket_event_max_silence_seconds` (default 300s, 0 disables the
   dimension alone). The stamp resets on `on_ready` so a reconnect never
   inherits pre-restart silence. Trip path is unchanged: consecutive
   failures -> retryable `discord_websocket_health_stale` -> the existing
   reconnect watcher builds a fresh adapter.

2. Silent probe disable. `_finite_positive_config_float` / `_config_int`
   mapped anything `float()` rejects ("15s", "nan", "true") to 0.0 with no
   log line, permanently disabling the watchdog invisibly. Unparsable and
   non-positive values now log one WARNING naming the knob and raw value.

The new key rides the existing `_YAML_WEBSOCKET_LIVENESS_KEYS` seeding and is
documented in the Discord guide's liveness section.
2026-09-13 20:03:40 +05:30
teknium1
c764d6d354 fix(shared): ensureContrast keeps the desktop's 0.2-step ladder; TUI chain opts into 0.05
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.

Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
2026-09-13 06:50:57 -07:00
Teknium
4beb7e29a2 fix(mcp): unattended paths never open browser OAuth; gateway follows mcp_servers edits
A gateway with an OAuth MCP server whose refresh token expired opened a new
authorize tab every 300s, all night (92 tabs). Four defects stacked:

- The parked-server self-probe re-entered the SDK's authorization-code flow
  with interactive OAuth enabled. The timed wake is unattended by definition:
  `_wait_for_reconnect_or_shutdown` now distinguishes "self-probe" from an
  explicit "reconnect", and `_park` flips the task-local
  `_oauth_interactive_enabled` off before a self-probe revival.
- Gateway MCP discovery (startup, `/reload-mcp`, hot-added multiplex
  profiles) ran interactive, unlike the CLI's background discovery. All three
  now run under `suppress_interactive_oauth()`; an expired token parks with
  the `hermes mcp login` hint instead of a browser.
- `_is_interactive()` trusted `sys.stdin.isatty()`, which the Windows CRT
  reports True for a DEVNULL/detached stdin. `_stdin_is_console()` confirms
  with `GetConsoleMode` on Windows.
- Removing an `mcp_servers` entry (or `enabled: false`) never reached a
  running gateway; the parked server probed forever. New
  `reconcile_mcp_servers_with_config()` tears down dropped/disabled servers
  (via `shutdown_mcp_servers(names=...)`) and connects new ones; a
  housekeeping chore runs it when config.yaml's (mtime, size) changes.
  `_select_new_servers` also stops nudging disabled parked servers.

Fixes #81830. Fixes the browser-storm item of #96320.
2026-09-13 06:29:12 -07:00
teknium1
b05a47b9d2 feat(desktop): reasoning effort gets its own composer pill
The composer showed "<model> · Med" in one truncating pill and the only way
to change the effort was to open the model menu, find the active model's
row, and hover it for the per-row options submenu. Users read the pill as
"this model is medium only" and never found the submenu.

- New `ReasoningPill` next to the model pill: shows the active model's live
  effort (session value, else the profile default) and opens the same
  Thinking / Fast / Effort rows the catalog submenu offers, for the active
  model only. Hidden when the catalog reports `reasoning: false`; stays
  while capabilities are unknown so it never flickers during the fetch.
  Folds away with the model pill in the compact composer stages.
- `useModelMenuController` (shell sibling) now owns the session write /
  preset / optimistic-store / rollback logic that lived inside
  `ModelMenuPanel`; the model menu and the new `ReasoningMenuPanel` share it
  so an edit from either surface is one code path. Tiles get their own
  pill bound to their SessionView, primary or tile — never the globals.
- `ModelOptionsContent` (the submenu body) is exported container-free so
  the pill's top-level menu renders it without a Radix Sub wrapper.
- The model pill drops the effort suffix (`formatModelPillLabel`: name +
  Fast); `formatModelStatusLabel` had no other caller and is removed.
- `currentModelCapabilities()` in lib/model-options resolves the active
  pick's caps through `catalogProviderMatches` (aliases, custom slugs).

Live (headless Electron + worktree `hermes serve`, CDP): before — one
pill "Deepseek V4 Flash · Low", no effort control; after — "Deepseek V4
Flash" + "Low" pill; pick High → `config.get reasoning` on the live
session returns high; a `reasoning:false` cap unmounts the pill; the
catalog row submenu still writes through and the pill mirrors it.

Credit: the dedicated-pill direction was proposed independently in
composer selector on current main with the shared-controller shape.
2026-09-13 06:11:49 -07:00
teknium1
0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
teknium1
0c0875b746 chore: delete orphaned bench data, datagen examples and stale one-off docs
Nothing in the tree reads any of these; they landed with feature PRs and
were never routed to their proper home.

- mcp-research-data/: 224K of July tool-search bench result rows. The
  harnesses (scripts/tool_search_livetest_ue*.py) write their output to a
  gitignored dir; the rows were committed by hand once and the headline
  numbers already live in the bench commit messages.
- datagen-config-examples/: Feb 2026 RL datagen configs for a
  WebResearchEnv that no longer exists; the yaml paths point at a
  configs/ dir that was never created.
- docs/: ADR log with one entry, an implemented cron-doctor spec, an RCA
  for a resolved bug, two RFCs whose work shipped, an unimplemented
  profile-builder proposal, the kanban dialog mock HTML and the kanban v1
  spec PDF. profile-routing.md duplicated the profile_routes section of
  website/docs/user-guide/multi-profile-gateways.md.

Kanban docs and the `hermes kanban` parser description pointed readers at
the PDF; those now point at the user guide (the patterns table it was
citing is on that same page).
2026-09-13 06:06:46 -07:00