Commit Graph

2616 Commits

Author SHA1 Message Date
salch-cred
101861c7f7 fix(discord): dispatch-side liveness dimension detects an ACKing-but-deaf gateway socket (#109521)
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.

This adds the dispatch-side signal that was requested instead:

- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
  every parsed DISPATCH frame with no debug gate (verified live against
  the real received_message path: 4/4 frames fired with
  enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
  incident report's field-proven operator bound); 0 opts out of this
  dimension ONLY — the #109782 review failure put the knob in
  _start_liveness_probe's all-or-nothing guard, killing the whole
  watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
  stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out

Fixes #109521

(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
2026-09-15 10:48:58 +05:30
teknium1
2de17e5d40 feat(plugin-catalog): default shelf is Desktop, catch-all is General; categorise today's six entries
Teknium's call: most community submissions are Desktop panes, so an entry
without a category lands on the Desktop shelf; "other" becomes "general" for
plugins that genuinely span areas. Shelf order puts Desktop first. The six
entries merged today (pets-all, newswire, auto-titler, live-voice,
metamask-wallet, web-octen) get explicit categories.
2026-09-14 21:00:29 -07:00
teknium1
55dbd7f6e1 feat(plugin-catalog): shelve the catalog by category (Memory, Desktop, Platforms, …)
The catalog page was one undifferentiated grid filtered only by tier, so a
memory provider sat between two Desktop panes. Entries now carry an optional
``category`` (memory | desktop | platform | web | tools | voice | automation |
models | other, default other) that the loader, the admission validator and
the site extractor all understand.

/docs/plugins renders one shelf per category in browse mode, a category pill
row under the tier pills, a clickable category chip on every card, and a
results bar (active category, count, clear) when a filter or search flattens
the view. ``hermes plugins catalog`` gains a Category column and groups by it.
All 18 shipped entries are categorised. Unknown categories fail admission
(same contract as tier) so a typo cannot create a phantom shelf.
2026-09-14 21:00:29 -07:00
teknium1
e80642df60 fix(bot-mode): keep group follow-ups ordered and late answers visible
Serialize room drives through their actual member completion, freeze input
watermarks by retained entry identity, and share the same completion path
with handoff continuations. Stop discards queued work without releasing an
active owner early; rename follows the existing room binding.

Observe stranded replies for the hard-cap duration plus grace after the
foreground wait, and retain unresolved failures in collapsed Activity.
Never automatically retry an ambiguous failed submit within the same drive.

Adapted from the queue and boundary approach in #92041 by @enwaiax and
harvest-budget approach in #107193 by @Finn763; #106502 by @wadib identified
failed-submit watermark consumption. The implementation retains current
numeric watermark storage, room lifecycle bindings and serial round limits.

Related: #92003, #105247, #100026
2026-09-14 17:35:04 -07:00
Teknium
9975fd56be Merge pull request #111320 from NousResearch/docs/plugin-catalog-no-self-update
Plugin catalog: no self-updaters instead of a 2-week pin age, enforced in CI
2026-09-14 17:34:43 -07:00
teknium1
42602bd12a docs(computer-use): skill matches the driver's current vocabulary
The skill described a SOM overlay burned into the screenshot, a driver-side
`capture` tool, and manual symlinking of the cua-driver skill pack; users
who read the driver's own docs then called raw MCP tools (`capture`,
bare `element_index`) and hit "no reviewed risk classification" and
`snapshot_id_required`. State plainly that `computer_use(action=...)` is a
wrapper vocabulary the driver never sees, that `element=N` is translated to
the snapshot token, what a `stale` refusal means, how text-only models get
vision (auxiliary.vision routing / mode=ax), and the Windows WindowsApps
doctor failure. `cua-driver skills install` links into ~/.hermes/skills now.
2026-09-14 17:34:31 -07:00
teknium1
8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1
002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1
3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1
5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
teknium1
48763a4d01 docs(plugin-catalog): replace the 2-week pin-maturity rule with a no-self-updater rule, enforced in CI
Teknium's ruling: catalog plugins do not need the 2-week maturity window,
but they may not ship an in-app updater that downloads and replaces their
own files, because that makes the reviewed SHA pin decorative. Rule 3 in
the README and item 5 on the docs page now say so, and plugin-catalog-ci
fails an entry whose catalog build both fetches from GitHub releases/raw
and writes or renames plugin files (either half alone is allowed).
2026-09-14 17:17:19 -07:00
teknium1
de5ff9ba85 test(bot-mode): prove native profile fallback and nested delivery ownership
Retain executable Electron evidence for the named-unowned and live-owner
paths; stale PATH fails on base and passes with the contributor fix.
Clarify that --in selects cwd rather than the profile database.
2026-09-14 17:04:56 -07:00
teknium1
1f1f02e354 fix(desktop): preserve qualified bot identity in primary handoffs
Keep the salvaged botHandle normalization, but remove the unconditional
hermes alias on remote default profiles: the parser's last-wins map
otherwise retargets a local @hermes handoff by roster order.

Consolidate regression coverage into two invariants for persisted primary
handles, both handoff directions and three-source qualified identity.
Capture real Electron screenshots, durable logs and source receipts;
exercise the reverse live handoff too. Document the repair and bump the
bundled Desktop patch version.
2026-09-14 17:04:27 -07:00
teknium1
8c82e93132 fix(gateway): migrate --multiplex rolls back or resumes when the default cannot come up; every installed unit counts
`hermes gateway migrate --multiplex` ran its one fallible step LAST (install +
start the default gateway) with nothing around it. On a fleet whose secondary
ran a system unit as root (#110850) that step raised, leaving the flag on, the
secondary's unit removed and no gateway anywhere, and the re-run hit the
"already multiplexing (flag on)" short-circuit over an empty fleet.

- apply_migration(): the default bring-up runs inside a rollback. On failure the
  manifest written before the first destructive step restores the flag and
  reinstalls every recorded per-profile gateway with its recorded User=.
- MigrationPlan.interrupted: flag on + manifest present + no live default
  gateway is a half-applied migration, not "already multiplexed"; the re-run
  resumes from the manifest (target manager and User= read from it, since the
  units themselves are gone) instead of refusing. Flag off + leftover manifest
  refuses to overwrite it and points at --standalone.
- ProfileGateway.services records EVERY installed unit (user and system) and the
  manifest carries them; apply stops/uninstalls all of them and rollback
  reinstalls all of them, so a second owner is never left live beside the
  multiplexer. The unattended hook treats a two-unit profile as an ambiguous
  topology and refuses (review finding on #110205).
- gateway_identity(): an unresolvable User= on a system unit stays None instead
  of borrowing the profile directory's owner; the unattended hook treats the
  unknown principal as a boundary (review finding on #110205).
- auto_migration_opted_out(): reads the effective config (load_config_readonly
  under the default home), so a managed `false` wins over a user `true` and a
  YAML string "false" is an opt-out, not a truthy value (review finding on
  #110205).

Builds on KoNit-K's #110854 (run_as_user threaded through install, preserved
from the removed system unit).
2026-09-14 16:16:06 -07:00
teknium1
a21747fe4e fix(kanban): an anchorless thread subscription warns once instead of vanishing (#110919)
Follow-up on the salvaged #110928 (CLI `--parent-chat-id` / `--guild-id`):

- `_claim_for_sub` skipped a thread-shaped row that matched no `profile_routes`
  entry at DEBUG on every tick. Legacy rows written before the flags existed can
  never match a channel-level route (no `parent_chat_id`), so the notifier now
  logs ONE WARNING per row naming the task, the thread and the re-subscribe
  command. Still fail-closed: the events stay unclaimed.
- Docs: the kanban user guide explains the anchors and shows the Discord-thread
  subscribe command under `profile_routes`.
- Test (red on origin/main): two collects → exactly one WARNING, events unseen.
2026-09-14 16:14:33 -07:00
teknium1
40f2702b22 feat(desktop): chat/UI font picker (desktop.font_family) for readability faces
Settings → Appearance gains a Chat Font row next to Terminal Font. The value
lands in config.yaml as desktop.font_family and is layered in front of the
active theme's fontSans when the theme paints --dt-font-sans, so an empty value
is exactly the theme and a chosen family keeps the theme's CJK/emoji fallbacks.

Why: #72485 asked for OpenDyslexic in the app; #76395 only made the terminal
pane configurable, and chat/chrome typography had no user-facing knob at all
(theme presets set fontSans, imported VS Code themes carry no font opinion).

Refs #72485, #37566
2026-09-14 15:21:02 -07:00
warmheartrobot
bbea5b98b9 docs: note macOS installer is Apple Silicon only
platform-support.md already lists "macOS on x86 (Intel) processors" under
Unsupported, but the two pages users actually land on don't reflect it:
installation.md recommends the macOS installer without qualification, and
desktop.md says the app "runs on macOS, Windows, and Linux".

Cross-reference the existing policy from both pages so Intel users find it
before downloading rather than after "Bad CPU type in executable".

No change in platform support is proposed or implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 15:15:31 -07:00
Siddharth Balyan
cf35e7351e fix(connectors): manage_connections is absent for accounts the portal has not enabled (#111238)
A signed-in, paid Nous account that the portal had not enabled for
connectors got `manage_connections` in its schema and a raw "tool gateway
request failed with status 404" back from every call. The gateway answers
404 for any such account by design, and Hermes gated the tool on paid access
or a free tool pool, which says nothing about that.

The gate now reads the portal's own answer: a `managed_tools` token claim,
plus the existing free-tier leg. The gate is also the tool's check_fn, so a
session without the claim never sees the tool and the model has no 404 to
narrate. A token without the claim reads as not enabled.
2026-09-14 22:01:32 +00:00
Siddharth Balyan
ee2f5629b8 Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code

Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.

* feat(connectors): managed connect runs on the connection operation

Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).

Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.

What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
  transition table. `operation.transition()` enforces it; a card cannot claim a managed
  target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
  reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
  settle -> result). The managed `observe` hook polls the gateway list once per tick for the
  whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
  shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
  are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
  the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
  docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
  added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
  is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
  `connect_card_available` instead of the link when the session platform is `desktop`.

Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.

* feat(desktop): connector card subscribes to the connection operation

The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.

- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
  `applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
  entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
  is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
  snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
  `Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
  failed / expired reissues through `connectors.connect` on the open operation, Not now is a
  per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
  with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
  `HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
  backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
  and compiles unchanged; PR3 moves it onto the operation.

anti-slop: no net-new findings (17 touched files vs 11d1a12472).

* fix(connectors): the card never parks the tool thread; every update carries the snapshot

Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).

- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
  the tool thread on a private request-id Event until a `_respond` that no longer exists for
  this event. `connection.respond` settled the operation but the tool waited its full deadline
  before the watcher loop even started. The callback now only emits the card; the operation's
  own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
  through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
  (state, link, detail). The initial mint happened before the card existed, so the renderer
  never saw the links and Connect stayed disabled; a Continue settlement stamped
  `not_connected` on the backend while the card still showed `initiated`. The store now
  overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
  once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
  `_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
  and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
  `interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
  (the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.

tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.

* docs(connectors): prompts and docs describe the operation, not the deleted wait verb

The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.

* fix(connectors): the panel re-mints only a dead link

Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.

* test(connectors): the local-batch test answers the operation the way the card does

The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.

* ci: retrigger

* fix(connectors): the desktop card appears outside guided onboarding

Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.

The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").

The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.

ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.

message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.

* style(connectors): shorter comments, no module mock in the card router test

The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.

anti-slop: no net-new findings (25 touched files)

* fix(connectors): Connect on a waiting row opens the stored link

ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.

The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.

* fix(connectors): a settled card stays dead; the card binds to its tool call only

A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.

The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.

`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.

`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.

Sid's rule of record: a resolved card is fully dead; no path brings it back.

* fix(connectors): the watch loop settles once, on time, and never raises into the result

Three findings from the live review, one loop.

Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.

Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.

Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.

Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.

* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking

run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.

The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.

* fix(connectors): a failed Try again shows the failure, not the old dead link

The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.

One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.

* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected

`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.

A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.

* fix(connectors): the operation registers under the gateway session key

The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.

The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.

* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed

Three follow-ups from the verification of the fix pass.

The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.

Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.

`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.

`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
2026-09-15 00:41:14 +05:30
Siddharth Balyan
e0ef0eb9c3 manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema

One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.

MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.

Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.

`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.

Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.

`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.

The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.

* wip(desktop): connection.request store, resume restore, card routing for MCP targets

Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.

* fix(config): hermes update turns on the connections toolset for saved toolset lists

`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.

Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.

`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.

* refactor: anti-slop pass on the desktop slice; shorten added comments

Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.

* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed

The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.

session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.

lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.

vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.

* chore: drop __pycache__ files swept in by an over-broad git add

* fix(desktop): correlate the connection.request row with the model's tool call by reason

The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.

* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds

* fix(connections): settle reason derives from target state, never from the renderer

A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.

* fix(desktop): a pending connection card re-arms on resume and activate

The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.

* style: literal wording in added comments, docstrings and docs

* fix: shared gateway-event contract and config-schema category for the connection events

connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.

* style: import order (perfectionist) in the desktop and shared files this PR touches

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-15 00:41:13 +05:30
teknium1
3275ca88ec fix(image_gen): Codex-auth images use the native images endpoints, no chat host model
The openai-codex image provider rode a Responses call with a hosted
image_generation tool on a pinned chat model (gpt-5.5). Two failure
classes came with that shape: when OpenAI withdrew gpt-5.5 from an
account cohort every image call 404'd while chat kept working
(#105398, #107076), and the host model was free to answer in text
instead of calling the tool, so we streamed SSE, kept partial frames
and retried on empty streams.

Post to chatgpt.com/backend-api/codex/images/generations and
images/edits instead - the route the official Codex client uses
(codex-rs/ext/image-generation). No host model, no SSE, no
partial-frame handling; the response is a plain JSON body with
b64_json. Remote source URLs are fetched client-side and inlined as
data URLs because the backend's own downloader 400s on ordinary
public images.

The backend treats model/quality/size as advisory (#107233), so the
result now reports reported_quality/reported_size next to the
requested values plus the x-codex-imagegen-request-id for support.
GPT Image 2.5 is deliberately not added to this catalog: the backend
accepts any model id, including nonexistent ones, and generates with
its server-managed engine (C2PA reports gpt-image 2.0), so a 2.5 tier
here would be a label with no effect (#106708).
2026-09-14 10:18:13 -07:00
teknium1
498abb677e fix(cron): a successful run resolves the job's open incidents; a repeat re-opens them
The incident ledger only ever grew. A one-off failure (a drift skip after a
global model bump, a provider outage) stayed `detected`/`alerted` forever
after the job recovered, so `hermes cron incidents` listed 32 "open"
incidents on an install where all 32 jobs had since run OK, and the list
stopped saying anything about current health.

A successful run now marks that job's `detected`/`alerted` incidents
`resolved` (new state). `resolved` is distinct from the operator's `closed`
ack on purpose: `upsert_incident` re-opens a resolved incident as
`detected` when the same error signature recurs, so the operator is alerted
again for a job that broke a second time, while `closed` keeps the
signature silent as before. Wired from `_compose_run_delivery` next to the
failure-side upsert; best-effort, store errors never affect delivery.

CLI: `--state resolved` filter and a green `resolved` colour; `closed` is
now dim. Docs updated in the same change.
2026-09-14 09:19:32 -07:00
teknium1
e383c28d2f fix(approval): judge shell quoting on the raw command, not the escape-stripped one
The hardline floor tracked quote state on text that
_normalize_command_for_detection had already rewritten (`\"` -> `"`).
That flips quote parity and broke both ways:

- false positive: a shell-valid `grep -o "[^\"]*" f` lexed as an
  unterminated quote and hit the unconditional "malformed executable
  payload" block. 118 of the 125 hardline blocks in one week of real
  agent use on this install were this shape, every one a benign grep,
  and the block cannot be bypassed by --yolo or approvals.mode=off.
- bypass: `cat "f\"n.txt"; rm -rf --no-preserve-root /` put the `; rm`
  start "inside" a phantom quote, so no command start was marked and the
  floor let the root wipe through with approved=True.

Fix: the malformed-quoting verdict reads the raw command (only quoted
newlines masked, which keeps quoting intact), and
_command_detection_variants adds a variant whose command starts were
marked on the raw command before normalization. The marker is " \n" so a
preceding literal backslash cannot eat it as a line continuation.
_iter_shell_command_starts no longer treats the `{` of `${...}` as a
brace-group opener, so the new variant does not split `${IFS}` and defeat
the IFS collapse.

Direction from #85922 by @Soju06, re-implemented onto the decomposed
tools/approval_detection.py; the parameter-expansion scanner from that PR
is replaced by the one-character `${` check above.

Co-authored-by: Soju06 <qlskssk@gmail.com>
2026-09-14 09:19:25 -07:00
teknium1
4748caff76 fix(gateway): explicit tool_progress new/all keeps text progress in un-cardable Slack chats
The destination preflight / refusal path suppressed the whole progress lane
for a flat DM regardless of mode, so an operator who WROTE `tool_progress:
all` got nothing there (before #108668 they got text bubbles via the
fallback). Silence is right only for Slack's tier default, where no text
lane was asked for; explicit new/all now routes through the editable text
fallback instead. Also hoists resolve_tool_progress into the existing
display_config import in _run_agent_display_settings.

Test proven red on the salvaged head (adapter.sent == [] with `all`).
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
e9c037f65a docs(slack): clarify progress resolution and fallback lifetime 2026-09-14 07:46:51 -07:00
Victor Kyriazakos
bc125d59d5 fix(gateway): null tool_progress inherits; name the task-card suppression latch
Review findings (Salt, adversarial pass on the two preceding commits):

- BLOCKING: a `tool_progress: null` (global, platform, or legacy overrides)
  counted as an explicit mode because the gate tested key presence, while
  the display resolver skips None and inherits. Null resolved to Slack's
  tier default `off` and disabled cards, which is the default-off trap the
  change exists to avoid. Explicit intent is now a non-None value (or the
  env bridge). Tests cover null at each level plus null-over-global-all;
  mutation to key-presence turns the three null cases red.
- TASTE: `_TaskCardState.egress_declined` now also latched on unsupported
  destinations, so the name no longer described the field. Renamed to
  `publication_suppressed` with both causes documented; readers unchanged.
- SHOULD-FIX: slack.md still promised an unconditional text fallback and
  described the opt-in as independent of tool_progress. Rewritten: cards
  follow an operator-written off (including /verbose), null inherits, an
  un-threaded chat with the card lane active shows no tool progress, other
  native failures keep the editable fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
3412490ad1 fix(gateway): no text tool progress when a Slack chat cannot host a task card
In flat Slack DMs (reply_in_thread false) the connector refuses task cards
("slack task_card requires a thread anchor"; native Slack: "No Slack thread
target"). The card lane treated that like a transient native failure and
fell back to an editable text message, so every tool event re-rendered
"Hermes is working / - tool - running" in the DM: text tool progress on a
platform whose default is off, for an operator who never enabled it.

Treat unsupported-destination refusals as terminal for the turn (same
latch as an egress decline) and log at info; transient native failures
keep the text fallback.
2026-09-14 07:46:51 -07:00
Victor Kyriazakos
ed25a40917 fix(gateway): explicit tool_progress off disables Slack task cards
Slack task cards are tool progress rendered natively, but the card lane
ignored the operator's tool_progress mode. Slack's built-in display tier
sets tool_progress off, so the lane was decoupled on purpose (#29483) to
keep cards on for unconfigured installs. The side effect: an operator who
wrote `display.platforms.slack.tool_progress: off` to silence tool updates
still got cards, and on relay-fronted Slack (where the connector always
advertises task_card) there was no setting that could turn them off.

Gate the card lane on operator intent, not the tier default: cards stay on
when nothing is configured, and go off only when tool_progress was written
as `off` (global, platform override, legacy overrides, or the env bridge).
`new`/`all` keep cards.

Tests assert the wire contract: no native card send, no stop, no fallback
text for an explicit off; card lane engaged for `new` and for the
unconfigured tier default (regression guard for #29483). The duplicate-tools
fixture now mirrors production's _safe_callback null-guard.
2026-09-14 07:46:51 -07:00
teknium1
9f7f2f28c0 feat(gateway): server→client JSON-RPC requests replace the *.request/*.respond event pairs (#110521)
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.

`tui_gateway/server_requests.py` owns the one mechanism:

  send()          block the agent thread until the response frame
                  (`srq-<n>` ids; ints belong to the client)
  send_async()    fire-and-callback variant (bot relay)
  cancel*()       withdraw with ONE `request.cancel {id, method, reason}`
                  event (timeout / interrupt / process exit /
                  answered elsewhere) instead of per-kind *.expire
  open_requests() the still-open frames, replayed by session.resume,
                  session.activate and session.events.since so a
                  reconnecting client re-renders every kind, not two
  clarify.lock    stays a real client→server RPC (locks one batch
                  answer early); locked answers merge into the final
                  set even when the closing response carries only the
                  tail the user answered last

A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.

Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.

Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
2026-09-14 06:02:05 -07:00
teknium1
7d368c7d2c docs: background_review.reasoning_effort applies on the routed path; same-model warns once 2026-09-14 05:25:01 -07:00
teknium1
ee51e8bf8b chore(webhook): drop external-product attribution from code, tests and docs 2026-09-13 21:30:02 -07:00
Teknium
45ab5e3b5d Inspired by ChatGPT Work: event-triggered cron jobs via webhook routes
ChatGPT Work's Aug 25 2026 release lets scheduled tasks fire from app
events (new Gmail message, Slack activity, GitHub PR feedback) instead
of polling on a cadence. This ports the pattern by composing two
existing Hermes subsystems: a webhook route can now set cron_job to
fire an existing cron job on each inbound event.

- gateway/platforms/webhook.py: cron_job route mode — after the same
  HMAC auth / rate limit / filters / script / idempotency as agent
  routes, the rendered prompt becomes transient per-run context and the
  job fires through execute_job_for_event on a worker thread (202
  Accepted immediately). Startup validation rejects cron_job +
  deliver_only.
- tools/cronjob_tools.py: execute_job_for_event() — public wrapper over
  the shared claimed-run body (_execute_job_now), so event fires share
  at-most-once claiming, in-flight dedupe, delivery, and [SILENT]
  handling with scheduler and manual runs.
- hermes webhook subscribe --cron-job: creates event-trigger
  subscriptions; job ref validated (and canonicalized to the job ID) at
  create time.
- Docs: webhooks.md route table + Event-Triggered Cron Jobs section,
  cron.md capability list, zh-Hans mirrors.
- Tests: tests/gateway/test_webhook_cron_trigger.py (adapter + unit),
  CLI tests in test_webhook_cli.py.
2026-09-13 21:30:02 -07:00
teknium1
c8e155dcbf fix: ai-presenter-video resolves its dir via ${HERMES_SKILL_DIR}, regen docs page
The shell `find ~/.hermes/skills ~/.hermes/hermes-agent/optional-skills ...`
snippet hardcoded a dev-clone path that does not exist on user installs; the
loader already substitutes ${HERMES_SKILL_DIR} (skills.template_vars), which
is what every other bundled skill uses. Drop the upstream-agent path mention.
Regenerated the docs page so it matches SKILL.md (also fixes the comfyui
related-skill link, which now lives under optional).
2026-09-13 21:12:23 -07:00
Teknium
a79ff58d65 feat(skills): ai-presenter-video optional skill (port of lanshu, 955★ MIT)
Ports cclank/lanshu-create-ai-presenter-video (MIT, 955 stars in 7 days)
into optional-skills/creative/ai-presenter-video. Provider-neutral
presenter-video production: locked narration as master clock, avatar
generation with pilot-first cost discipline, lip-sync/identity QA,
captions, deterministic ffmpeg finalization with loudness normalization
and contact-sheet verification.

Hermes adaptations in the hub SKILL.md: SKILL_DIR resolution (upstream
hardcoded ~/.codex/skills), capability mapping to text_to_speech / FAL
video families / vision_analyze / hyperframes, consent-flag JSON paths
(input.* vs root), preflight error-vs-remote-blocker semantics.
References kept substantively verbatim (all-English upstream). Scripts
unmodified. LICENSE carried.

Validated hands-on: init_job -> preflight gating (blocked until manual
review booleans + input.remote_upload_approved) -> finalize_delivery on
a synthetic 1080x1920 render (master+share decode-verified, delivery
report, 9-frame contact sheet). Cold-subagent live test: SHIP; 3
friction fixes folded in (resolution guard, boolean-flip example,
preflight-writes-job note).
2026-09-13 21:12:23 -07:00
Teknium
56cc2bd814 feat(skills): scrollcraft — premium scroll-driven landing pages (port of nateherkai/scroll-craft, 1.2k★ MIT)
Optional skill: scroll-as-timeline landing pages on a deterministic
CSS/JS engine, with interview → page grammar → signature move workflow
and screenshot-based scroll verification. Engine and scripts vendored
verbatim; asset generation re-anchored on image_generate with the
upstream kie.ai flow kept as an optional path.
2026-09-13 21:11:25 -07:00
teknium1
230ca004a7 fix(delegate): forward inline data-URL images to vision children; trim tests to invariants
data:image/... entries were treated as local paths and silently skipped as
"unreadable". They now ride as image_url parts only (never pasted into the
text hint, never appended to a text-mode goal). Skips and forwarding
failures log at warning since the caller explicitly asked for the images;
decide_image_input_mode gets the child's requested_provider like the CLI
and gateway callers.

Tests collapse to five invariants, including one that drives _ChildRun
and asserts the multimodal content list reaches run_conversation as the
first user turn. Docs mention data: URLs and the read guard.
2026-09-13 21:05:42 -07:00
Teknium
f3f5c4f7c7 Port from RooCodeInc/Roomote#1796: per-task image forwarding on delegate_task
Subagents can now SEE images. Each delegate_task task accepts an optional
images list (max 8; local paths or http(s) URLs). Vision-capable children
receive native image_url content parts on their goal turn (local files as
data URLs, remote URLs verbatim); non-vision children get
[Image attached at: ...] hints plus a vision_analyze pointer. Routing
reuses agent.image_routing (decide_image_input_mode /
build_native_content_parts), so agent.image_input_mode governs delegation
exactly like inbound gateway images.

Best-effort by contract: malformed images arrays fail the call loudly
before any child spawns; unreadable paths are skipped with a log line;
any exception in the forwarding path degrades to the text-only goal.

Adapted from RooCodeInc/Roomote#1796 / #1767 (Fast agent forwards bounded
current-turn attachments to delegated coding tasks).
2026-09-13 21:05:42 -07:00
Teknium
c63de5a231 feat(tool_search): long hunts for nonexistent tools now return no results instead of incidental matches
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).

A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.

Docs: relevance-floor bullet added to tool-search.md implementation
details.
2026-09-13 21:04:08 -07:00
Teknium
3304d205be feat(video-gen): Kling 3.0 Standard + Pro families on the FAL backend
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
2026-09-13 21:01:32 -07:00
Teknium
65a6b6831e Port from code-yeongyu/oh-my-openagent#7151: worktree audit gains --json, --older-than, and external-tree visibility
omo's omo-agent-toolkit worktree-sweep (their PR #7151) added three
capabilities our hermes worktree command lacked:

- --json on list and prune: machine-readable audit/result payloads so
  scripts and agents can consume verdicts without scraping table output.
- --older-than DAYS: an age floor that only ever RESTRICTS reaping
  (young-but-reapable trees are kept); it never widens eligibility, so
  the existing safety invariants are untouched.
- External-tree visibility: linked worktrees registered outside
  .worktrees/ are now reported read-only in the audit (branch, locked,
  missing) instead of being invisible, and registrations whose
  directory has vanished are dropped via git worktree prune (metadata
  only, no files touched) during prune.

Not ported: omo's ancestor-of-default-branch merge test (our git cherry
patch-equivalence is strictly stronger under rebase/squash merges), and
their hardcoded external-root exclusion list (we exclude by location:
everything outside .worktrees/ is hands-off).

Tests: 9 new contracts in tests/hermes_cli/test_worktree_gc.py (age gate
restrict-only, external trees never reaped, stale-registration prune
dry-run/real, JSON shapes, negative --older-than rejected). Live E2E on
a scratch repo verified all three flags end to end.
2026-09-13 20:55:29 -07:00
teknium1
cb3447b139 test: trim the context-cache guard tests to invariants; docs: name the surfaces that actually confirm
Tests collapse 13 change-detectors into 7 invariants (silent below threshold /
without context, fires above, same-model re-select silent, config override and
0-disables, registry threading, agent context derivation). The legacy 5-arg
guard test goes with the TypeError fallback it covered: that fallback was
defence for a case nobody has (every in-tree guard and test double is
*args-tolerant) and would re-run a guard whose real TypeError it masked, so the
rebased port passes the context positionally like every other argument.

Docs no longer claim the confirm fires on the Telegram/Discord pickers or the
dashboard: those surfaces call combined_selection_warning() without a live
agent, so the context-cache guard is (correctly) silent there.
2026-09-13 20:54:50 -07:00
Teknium
8b6931393e Port from langchain-ai/deepagents#5829: confirm mid-session model switches that abandon a large cached context
Providers key prompt caches per model, so a mid-session /model switch makes
the next reply re-read the entire conversation at full input price. deepagents
gates user-initiated switches behind a confirmation once the active thread
exceeds a configurable token threshold; this ports the same protection into
Hermes' unified selection-guard registry so it renders on every surface at
once (CLI/TUI picker, gateway /model, Telegram/Discord pickers, dashboard).

- hermes_cli/model_selection_guards.py: new context_cache guard +
  SelectionContext carrier + selection_context_for_agent() helper;
  registry threads live-session facts to guards (6-arg signature with a
  TypeError fallback for externally patched 5-arg guards).
- config: model.switch_context_confirm_tokens (default 100000, 0 disables).
- cli.py / gateway/slash_commands.py / tui_gateway/server.py: thread the
  live agent's measured context into the guard call.
- docs: configuring-models.md mid-session switch section.
- tests: tests/hermes_cli/test_context_cache_switch_guard.py (13 cases).
2026-09-13 20:54:50 -07:00
Teknium
d9e88e19e2 feat(mcp): bind stored OAuth refresh tokens to their issuer
Port from openai/codex#39615: the authorization server discovered for an
MCP server can change (protected-resource metadata edit, server
migration, DNS takeover). Without binding, Hermes would send the stored
refresh token to whatever issuer the server now advertises — handing a
long-lived credential to a different authorization server.

- HermesTokenStorage records hermes_issuer alongside cached tokens
  (stripped before OAuthToken.model_validate; never sent on the wire).
- Both provider classes (tools/mcp_oauth.py legacy path and
  tools/mcp_oauth_manager.py managed path) stamp the discovered issuer
  on every token save and enforce the binding on _initialize.
- On mismatch: refresh token is stripped from memory and disk; the
  unexpired access token keeps working; full re-auth happens at expiry.
- Legacy token files without an issuer adopt the current one once
  (no forced re-login for existing installs — deliberate divergence
  from Codex, which requires reauth).

Validated: 12 new tests + 163 existing MCP OAuth tests green; sabotage
run confirms the new tests fail without the enforcement; E2E against
the real manager provider class with a temp HERMES_HOME confirms
mismatch strips and match preserves.
2026-09-13 20:43:33 -07:00
Teknium
ce318290bd Inspired by Factory Droid: /queue prompts are now listable, editable, and reorderable before they run
Droid v0.203 (Aug 25 2026) added 'edit queued messages' — a queued steering
message can be pulled back and changed before it is sent. Hermes /queue could
only append blindly: no way to see, fix, drop, or reorder queued prompts.

/queue now supports management subcommands in the CLI:
- /queue            — list pending prompts (bare prompt still enqueues)
- /queue list       — same
- /queue edit N <p> — replace item N (keeps voice sentinel, #65827)
- /queue rm N       — remove item N
- /queue move A B   — reorder
- /queue clear      — drop everything
- /queue add <p>    — force-enqueue prompts starting with a management word

Queue mutations hold queue.Queue's mutex and rebuild unfinished_tasks so
join()/task_done bookkeeping stays consistent. Paste references expand on
enqueue and edit, matching the old inline path.

Reimplementation of PR #18833 by @abhinav11082001-stack (commit was authored
under a fabricated 'Hermes Agent' noreply identity that cannot be carried
into history; engineering credit is theirs), hardened for current main:
voice-sentinel-aware previews/edit, paste-reference expansion, queue
bookkeeping asserts, out-of-range no-op tests, and docs.
2026-09-13 20:41:41 -07:00
Teknium
48bd70b586 Port from nearai/ironclaw#7378: doc-fact contract test keeps slash-commands.md in sync with the command registry
Two-direction contract test (tests/website/test_slash_commands_doc_parity.py):
every CommandDef must be documented under its name or an alias, and every
doc table row must resolve to a registered command. Ported from IronClaw's
doc-fact contract tests (nearai/ironclaw#7378), adapted from their clap
--help parser to our COMMAND_REGISTRY single source of truth.

Real drift it caught, fixed here: /loop (alias /proactive) shipped with a
full feature page (user-guide/features/loops.md) and CLI+gateway handlers
but never got a row in the slash-commands reference. Added to both the CLI
Session table and the messaging table, plus the both-surfaces note.
2026-09-13 20:40:45 -07:00
xxxigm
1468e98e48 test(google-chat): pin hosted install wiring and document the lazy target 2026-09-13 19:19:36 -07:00
teknium1
1d33a4fee6 feat(desktop): persist auxiliary reasoning_effort through the models router
Backend half of the per-task effort control, on today's layout: POST /api/model/set
distinguishes omitted (leave the task's override alone) from explicit null (clear →
inherit) via model_fields_set, canonicalises a level through parse_reasoning_effort
(400 on an unknown one), and "Reset all to main" also drops every override. GET
/api/model/auxiliary returns reasoning_effort per task and the row summary shows it.

The inherit row reads "inherit · main model effort" (own i18n key in all six locales)
rather than reusing the provider's "auto · use main model" copy — the two mean
different things and the reused string read as "use the main model" for the effort.

Runtime already consumes auxiliary.<task>.reasoning_effort (agent/auxiliary_client.py)
and hermes model writes the same key (#110346), so Desktop and CLI now edit one value.

Closes #89259. Salvages #90649 by @higgs1729.
2026-09-13 19:19:32 -07:00
teknium1
2c0bec33f9 feat(model-pickers): reasoning effort selection on every model picker
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.

One request now carries a model pick AND its effort on every surface:

- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
  <level>` (validated against `parse_reasoning_effort`; unknown level ->
  `MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
  `ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
  high` applies the effort AFTER the agent swap (`switch_model` re-resolves
  `reasoning_config` from config.yaml, so an earlier write is clobbered) with the
  pick's scope (session; config on `--global`; `--once` snapshots and restores it).
  The `/model` picker gains a third stage, "Reasoning effort for <model>", built
  from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
  inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
  `config.set model "X --reasoning high"` applies after the swap; session pin
  (`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
  one-turn restore carries `reasoning_config`; re-emits `session_info` so the
  status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
  `<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
  label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
  `_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
  the Copilot-only inline prompt; Copilot keeps its per-model level set via
  `github_model_reasoning_efforts`, other routes get the ladder, catalog
  `supports_reasoning=False` skips it) plus a "Reasoning effort for the current
  model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
  with the same step (+ "Provider default"), stored as
  `auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
  the task list ("openrouter · model · high"), cleared by "Reset all to auto";
  tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
  it.

Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
  "Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
  and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
  writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
  errors "Model names cannot contain spaces"; after switches and `config.get
  reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
  effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
  "reasoning: high", status bar "fable 5.1 high".
2026-09-13 16:43:50 -07:00
teknium1
2dfd831d3b fix(webhook): bind a subscription to a profile with --route-profile, not --profile
The salvaged flag was spelled --profile, which collides with the global
-p/--profile that hermes_cli.main scans BEFORE argparse: `hermes webhook
subscribe x --profile compta` would switch this CLI process to compta's
HERMES_HOME and write the subscription into compta's webhook_subscriptions.json
— a file the default gateway's webhook adapter never reads — while the route
still lacked the profile key. #109020 special-cased the scanner for the webhook
subcommand; naming the flag --route-profile removes the ambiguity without
touching _scan_profile_flag: -p picks the gateway whose subscriptions file is
written, --route-profile picks which /p/<profile>/ prefix may hit the route.

Docs: cli-commands reference row, multi-profile-gateways webhook section, the
route `profile` field. Builds on #109020 (fangliquanflq). Fixes #109016.
2026-09-13 15:45:57 -07:00
teknium1
6636b0896c docs(multiplex): OAuth and mTLS servers are never shared; trust and parallel policy are per profile
Extends the multi-profile MCP paragraph with the identity rules landed in
this branch: OAuth tokens live per profile so each profile's calls run as
its own account; client_cert/client_key are part of the connection
identity; trust and supports_parallel_tool_calls are the consuming
profile's policy even when it shares another profile's connection.

Wording of the per-profile OAuth account guarantee follows the docs draft
in PR #109574 (its token-file fingerprint code path was not taken).

Co-authored-by: ly6751 <99090550+ly6751@users.noreply.github.com>
2026-09-13 15:41:01 -07:00