7dee4cc068bbfbdb68b3769376808d84eaee4e0d
84 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb3011a202 |
fix(video_gen): resolve OpenRouter and DeepInfra credentials per profile, not from os.environ
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL straight from os.environ. That broke two setups: - A key added with `hermes auth add openrouter` (API key or OAuth) lives in the credential pool, not the environment. Chat and image_gen/openrouter find it through resolve_runtime_provider. video_gen reported OpenRouter unavailable, and generate() returned missing_credentials. - On a multiplexed gateway, os.environ holds the launch profile's .env. A routed profile's video jobs were submitted, polled and downloaded with the launch profile's key and billed to that account. A profile whose key lived only in its own .env could not use the backend at all. The backend now resolves (api_key, base_url) with resolve_runtime_provider(requested="openrouter"), the same call image_gen/openrouter makes. generate() resolves once and passes the pair to submit, poll and download. With a round-robin pool, resolving per request would poll with a different account's key than the one that created the job. OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through get_secret_str, as image_gen/deepinfra already does. |
||
|
|
550d74c62f |
fix: manage_catalog documented under its own setup toolset; post-hook ownership contract exercises it
tools-reference.md listed manage_catalog under the connections section, which the reference-docs contract resolves against the connections toolset. The agent-runtime post-hook ownership test enumerates every tool in AGENT_RUNTIME_POST_HOOK_TOOL_NAMES; manage_catalog now runs through both executor paths there. |
||
|
|
31e59b9441 |
feat: setup agent can search the catalog and install plugins and skills through the approval card
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog: - search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with whether each is already installed in the default profile. Read-only. - install: opens the same connection operation manage_connections opens, with rows of kind plugin / skill. Nothing installs until the user approves a row. An approved row installs into `default` (or the profile the Advanced modal named) through dashboard_install_plugin / the hub's headless install, so the catalog pin, kill list, security scan and live activation (#119644) are the host's. The row settles with the live MCP tool names and the plugin's skill. The model sends catalog ids and an action only; every other key is refused before anything runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the result is the `hermes plugins install` / `hermes skills install` pointer. - contract: plugin/skill targets follow the MCP transitions. - run.apply_answer / reissue route a card answer to the module that owns the operation. - tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never deferred behind tool_search. - docs: tools reference, toolsets reference, plugin catalog page. Linear NS-964. |
||
|
|
e636aedebf | fix: preserve known file baselines across partial rereads | ||
|
|
65b0ff38ea |
feat(terminal): heartbeat notifications for long background processes
`terminal(background=true, heartbeat=N)` emits a `heartbeat` event every N seconds (floor 60) carrying only the output produced since the previous one, plus the usual completion notice. The agent stays current on a long bounded job (merge train, full test suite, deploy) and reacts to a failure within N seconds instead of at exit. Why: `notify=true` fires once at exit, so multi-hour merge trains ran silently — the parent's prompt cache went cold and every "what's up" cost 5–10 rediscovery calls; 30 days of batch sessions show 4,930 hand status polls (14% of all tool calls). `notify=[pattern]` cannot serve this: its lifetime cap (8) exists precisely to stop periodic output from flooding the agent. Mechanism: one daemon timer thread for every heartbeat session (reader threads block on the pipe and cannot keep time); delta output via a total-ingested counter over the rolling buffer; heartbeats only where the completion notice can be delivered (same async-support/subagent gates), never after exit. `heartbeat_seconds` is checkpointed. Rendered on CLI, gateway and TUI through the existing process-notification paths. |
||
|
|
afc3b7c6f3 |
feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash merges ( |
||
|
|
6ae4cb88d5 |
fix(docs): reference pages name tools the registry never registered
The shipped tool-surface references still document the pre-consolidation surface: tools-reference.md lists cronjob/todo/process/project_create/ project_list/project_switch/open_preview/close_preview/read_preview/tour/tip (6 uncallable, 5 hidden dispatch-only aliases), and toolsets-reference.md still claims web_search is a member of the browser toolset — membership |
||
|
|
19b29df13b |
fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend tools a model override) so the LLM could route a single call to a different model family — a different endpoint and billing tier — than the one the user selected in `hermes tools`. image_generate never exposed this, and #83080 asked to extend it there; the ruling is the opposite: models do not choose models. The `model` property is gone from the static and dynamic video_generate schema and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is ignored and the configured `video_gen.model` (then the provider default) is what reaches the request. Config-side selection (`video_gen.model`, `video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the xAI plugin's explicit-model branch is no longer reachable from the tool layer. Refs #83080 |
||
|
|
2fbcd8b0ea |
docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored and generated pages) and the zh-Hans mirror: 1,868 route-style links (`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`. Every target was asserted to exist on disk; anchors and query strings are preserved; fenced code blocks and inline-code examples are untouched. Two dead targets found by the converter were fixed by hand first: memory-providers.md linked `/user-guide/plugins` (page is `user-guide/features/plugins`), and the zh-Hans learning-path still linked the removed `rl-training` page — ported the EN treatment (external Atropos link). Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links, 0 broken anchors. |
||
|
|
98b4efbf48 |
fix: clarify cards that cannot render are re-asked as plain text, never reported as user inactivity
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could fail without the user ever seeing a question, and the agent then waited out the full clarify_timeout and reported "[user did not respond within Nm]" (#112684): - the platform rejects the card at once -> the wait aborted with a delivery sentinel but the question was never re-asked; - the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts, a stale-thread retry) -> the runner never looked at the send future again and blocked for the whole timeout for a card that never posted; - no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported timed_out=True with an empty notice. gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers move out of the gateway/run.py facade) now retries every definitive card failure once through the adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a possibly-delivered send so a late failure releases the waiter with "[clarify prompt could not be delivered]". The no-surface case reports "[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already treats as a non-answer). Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s): card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt, after the numbered prompt is sent and "2" resolves to the choice. |
||
|
|
b3aef45987 |
fix(clarify): batch result carries the surface's no-answer notice; trim tests
Follow-up to the #112725 salvage. When a gateway clarify batch stops at an unanswered question, the batch payload only said `timed_out: true` with blank answers — so a card the platform REJECTED ("[clarify prompt could not be delivered]") read exactly like a user who walked away, which is the misreport #112684 describes ("a timeout must not be presented as user inactivity when the prompt was never delivered"). - gateway/run_turn_runner.py::_clarify_batch_sync: the unanswered question's own response text rides along as `notice`. - tools/clarify_tool.py::_run_batch/_batch_result: pass `notice` through into the result JSON beside `timed_out` (only when present); schema description mentions it. CLI/TUI batch callbacks send no notice -> byte-identical output. - tests: fold the single-question re-arm pin into the existing bracket-answer test (<=2 new tests), parametrize the end-to-end test over an undeliverable Telegram-shaped adapter so the delivery sentinel is pinned as `notice`. - docs: tools-reference.md describes `notice`. Part of #112684 |
||
|
|
ac63d0eea5 |
fix(agent): execution guidance and browser hints drop the web_search stripper; tests assert the invariant
The rebased guidance text no longer names web_search anywhere, so
execution_guidance_text()'s replace() calls (
|
||
|
|
e819846b10 |
Inspired by Amp: relative time bounds (7d/24h/2w) + wrapper forwarding for session_search after/before
Amp's thread feed supports relative time filters (`after:7d`, `updated_before:7d`) alongside ISO dates. Extend the salvaged after/before bounds (PR #86067 by @Moodtuner997) the same way: - `_parse_iso_bound()` now accepts relative durations `Nh`/`Nd`/`Nw` (case-insensitive) meaning "now minus N", alongside ISO dates/datetimes. Clearer error message names both accepted forms. - Forward after/before/exclude_session_ids through the public `session_search()` wrapper (the PR predates the wrapper/impl split; without this the SQL bounds were unreachable from the registry handler — same class as the earlier `detail` forwarding fix). Appended after `detail` to preserve positional compatibility. - Tool schema descriptions teach both forms. - Tests: relative after/before against the discovery shape, unit checks for h/d/w math, case-insensitivity, and bad-unit rejection. - Docs: tools-reference row mentions time bounds + exclude_session_ids. |
||
|
|
4613f895ab |
test: give the rewrite-hint fixtures a read baseline; align docs row with schema
The stale-write guard now refuses write_file on an existing file the task never read in full, so test_write_file_rewrite_hint's overwrite-without-read fixtures were refused before the hint could be computed. Reading first is the exact read->whole-file-rewrite pattern the hint exists for. tools-reference.md's write_file row now mirrors the WRITE_FILE_SCHEMA description (one-sentence contract + the recovery step) instead of a longer paraphrase. |
||
|
|
6569651b87 |
fix: harden salvage of #65605 — redaction-gated test, sibling test baseline, docs
- test_file_staleness redacted-read case now force-enables redaction (matches tests/agent/test_redact.py convention) so it exercises the sentinel path in hermetic CI where security.redact_secrets is unset. - test_write_verification CRLF case establishes a read baseline first (the new guard refuses unread existing-file overwrites by design). - tools-reference.md documents the read-before-overwrite contract. - contributors/emails mapping for DanSpicyTaco. |
||
|
|
cf35e7351e |
fix(connectors): manage_connections is absent for accounts the portal has not enabled (#111238)
A signed-in, paid Nous account that the portal had not enabled for connectors got `manage_connections` in its schema and a raw "tool gateway request failed with status 404" back from every call. The gateway answers 404 for any such account by design, and Hermes gated the tool on paid access or a free tool pool, which says nothing about that. The gate now reads the portal's own answer: a `managed_tools` token claim, plus the existing free-tier leg. The gate is also the tool's check_fn, so a session without the claim never sees the tool and the model has no 404 to narrate. A token without the claim reads as not enabled. |
||
|
|
ee2f5629b8 |
Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code
Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.
* feat(connectors): managed connect runs on the connection operation
Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).
Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.
What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
transition table. `operation.transition()` enforces it; a card cannot claim a managed
target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
settle -> result). The managed `observe` hook polls the gateway list once per tick for the
whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
`connect_card_available` instead of the link when the session platform is `desktop`.
Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.
* feat(desktop): connector card subscribes to the connection operation
The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.
- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
`applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
`Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
failed / expired reissues through `connectors.connect` on the open operation, Not now is a
per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
`HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
and compiles unchanged; PR3 moves it onto the operation.
anti-slop: no net-new findings (17 touched files vs 11d1a12472).
* fix(connectors): the card never parks the tool thread; every update carries the snapshot
Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).
- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
the tool thread on a private request-id Event until a `_respond` that no longer exists for
this event. `connection.respond` settled the operation but the tool waited its full deadline
before the watcher loop even started. The callback now only emits the card; the operation's
own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
(state, link, detail). The initial mint happened before the card existed, so the renderer
never saw the links and Connect stayed disabled; a Continue settlement stamped
`not_connected` on the backend while the card still showed `initiated`. The store now
overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
`_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
`interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
(the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.
tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.
* docs(connectors): prompts and docs describe the operation, not the deleted wait verb
The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.
* fix(connectors): the panel re-mints only a dead link
Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.
* test(connectors): the local-batch test answers the operation the way the card does
The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.
* ci: retrigger
* fix(connectors): the desktop card appears outside guided onboarding
Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.
The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").
The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.
ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.
message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.
* style(connectors): shorter comments, no module mock in the card router test
The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.
anti-slop: no net-new findings (25 touched files)
* fix(connectors): Connect on a waiting row opens the stored link
ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.
The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.
* fix(connectors): a settled card stays dead; the card binds to its tool call only
A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.
The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.
`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.
`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.
Sid's rule of record: a resolved card is fully dead; no path brings it back.
* fix(connectors): the watch loop settles once, on time, and never raises into the result
Three findings from the live review, one loop.
Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.
Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.
Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.
Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.
* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking
run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.
The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.
* fix(connectors): a failed Try again shows the failure, not the old dead link
The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.
One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.
* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected
`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.
A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.
* fix(connectors): the operation registers under the gateway session key
The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.
The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.
* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed
Three follow-ups from the verification of the fix pass.
The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.
Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.
`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.
`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
|
||
|
|
e0ef0eb9c3 |
manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema
One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.
MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.
Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.
`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.
Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.
`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.
The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.
* wip(desktop): connection.request store, resume restore, card routing for MCP targets
Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.
* fix(config): hermes update turns on the connections toolset for saved toolset lists
`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.
Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.
`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.
* refactor: anti-slop pass on the desktop slice; shorten added comments
Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.
* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed
The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.
session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.
lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.
vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.
* chore: drop __pycache__ files swept in by an over-broad git add
* fix(desktop): correlate the connection.request row with the model's tool call by reason
The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.
* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds
* fix(connections): settle reason derives from target state, never from the renderer
A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.
* fix(desktop): a pending connection card re-arms on resume and activate
The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.
* style: literal wording in added comments, docstrings and docs
* fix: shared gateway-event contract and config-schema category for the connection events
connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.
* style: import order (perfectionist) in the desktop and shared files this PR touches
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
|
||
|
|
3304d205be |
feat(video-gen): Kling 3.0 Standard + Pro families on the FAL backend
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro (fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key, aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and negative_prompt real, no seed/resolution keys per the published llms.txt schemas. Payload shapes pinned in tests; docs mention updated. |
||
|
|
c6f87deb2c |
feat(video): OpenRouter backend covers every model on the live video catalog
The salvaged #103267 plugin hardcoded a single model (minimax/hailuo-3-max) and rejected any other id. OpenRouter's public GET /api/v1/videos/models already publishes every generative model with its supported durations, resolutions, aspect ratios, frame-image support, audio and seed flags, and pricing SKUs, so the provider now reads that catalog (5-min TTL, offline snapshot fallback): - list_models(): all 25+ generative models (edit/upscale/avatar rows that take no duration are outside the unified video_generate surface and are dropped) with a per-second price label where the SKU is per-second - capabilities(): the CONFIGURED model's surface, so the dynamic schema only advertises audio/seed/resolutions the selected model honours - _build_payload(): clamps duration/resolution/aspect ratio to the model's live limits (nearest by value/height/ratio) and drops generate_audio/seed for models that lack them (the API 400s otherwise); reference images ride in input_references; local file inputs are refused (OpenRouter fetches URLs itself), data:image/ URLs from the sandbox chokepoint pass through - bearer key only ever goes to the configured origin (poll + /content), never to a provider-supplied unsigned_urls host (kept from #103267) Also drops the source-grep `_IGNORES_SEED` escape hatch #103267 added to the declaration⇄implementation sweep; the provider now implements seed for real. Docs list OpenRouter and DeepInfra as bundled video backends. Requested by Don Piedro Savastano (Discord): OpenRouter credit for video_generate. |
||
|
|
77e55b4d1f |
fix(desktop): show each in-app tip once, never lap the catalog again
The idle tip rotation walked the catalog as a ring: after the last tip it wrapped to the first, so a user who had already seen every tip kept getting "Start fresh", "Teach it once", ... again every six hours for as long as they used the app. Only the X stopped a tip, and letting a bubble time out (the normal way it leaves) counted for nothing. The walk is now one lap. nextTip also steps over every tip in the seen ledger ($tipShownAt, which already recorded every catalog tip that reached the screen), so a tip shows once however it left, and the rotation runs dry once every tip has had its moment. Settings > Reset clears the seen ledger and the cursor as well as the retired set, and its button counts what a Reset would actually bring back (shown or closed, counted once). Agent tips carry no catalog id and are untouched. Live repro (Playwright against the worktree's Vite renderer, all nine tips seeded as seen, clock fast-forwarded past settle + cooldown): origin/main re-showed "Start fresh"; fixed renderer shows nothing; a fresh user (nothing seen) still gets the first tip. |
||
|
|
f1ccf436a2 |
feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx: - search: POST https://api.perplexity.ai/search (documented Search API), search_context_size=low so `snippet` stays description-sized; results[].snippet -> description, max_results capped at the API's 20. - extract: POST /sdk/content/snippets — the query-relevant page-excerpt route behind `pplx content snippets` (the CLI's `content fetch` is deprecated upstream). web_extract has no query, so the URLs' path words serve as the relevance query; per-URL `error` entries survive a 200. - Wired into the same touchpoints as the other keyed vendors: legacy backend set + credential ladder + availability probe (web_tools), registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/ dump key lists, nous_subscription direct-credential detection, setup summary, test conftests, docs. Not a keyless-ring member (Perplexity has no anonymous tier). Related closed PRs #9192 / #23981 / #45225 predate the plugin ABC. |
||
|
|
428e084dcd |
feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199. |
||
|
|
b293157f7d |
feat(tools): withdraw tip and tour when the user switches them off
Off should mean the model is never told the tool exists. A switch that only made the call fail leaves Hermes offering walkthroughs it cannot give and promising to point at things it cannot point at, which reads as a broken agent rather than a respected preference. Both gate on the switch through a shared desktop_ui.user_enabled helper, which is the reactions check_fn generalized — same config read, same reason it has to be config rather than an env var: the switch belongs to the session's client, and the client may be on another machine. |
||
|
|
d6773cf26f |
refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and ring entry removed from keyless_mcp, legacy backend set / credential ladder / preference walks / rescue key map scrubbed. - TAVILY_API_KEY deregistered across config, setup, status, dump, and nous_subscription surfaces. The tvly- redaction pattern stays -- legacy keys in user envs still deserve masking. - Sibling test pins migrated (keenable/exa stand in where tavily was the fixture vendor); tavily test suite deleted. - Docs updated: web-search, configuration, integrations, environment-variables, tools-reference, web-dashboard, provider plugin dev guide. Live-verified from an isolated HERMES_HOME with all web creds blanked: zero-config resolution lands in the 4-vendor ring, live keyless ring search succeeds, no tavily anywhere in resolution order. |
||
|
|
b6bd681e89 |
feat(todo): nested subtasks via optional parent field
The todo tool now supports hierarchical task lists: an item's optional 'parent' field points at another item's id, making it a subtask. - tools/todo_tool.py: parent validated (self-ref dropped), dangling refs and cycles sanitized; merge mode can set/clear parent; post-compression injection renders the tree indented and keeps a finished parent visible while any descendant is still active; the in-progress reorder pass is skipped for nested lists (a flat move would tear subtasks from parents). - Schema cost: ~45 tokens added to the cached tool schema (one string property + one behavior sentence). - acp_adapter/tools.py: todo result markdown indents by parent depth. - Desktop: TodoItem carries parent; todoTree() DFS helper; composer status stack renders subtask rows indented (depth-capped), stabilizer compares depth. - Docs: tools-reference todo entry mentions nesting. Hydration/replay paths (gateway fresh-agent, API-server history) work unchanged: parent rides inside the same todos array. |
||
|
|
4bd2793367 |
feat(desktop): default the in-app tip rotation on (#96831)
Made opt-in when it fired a tip 45 seconds into a launch and another every six minutes, which is a cadence that owes you a choice. The pacing has since become a settling delay of five to ten minutes per launch and a six-hour cooldown persisted across them — roughly a tip a day, weeks to walk the catalog. At that weight the switch has nothing left to protect anyone from, and a discovery feature nobody meets is one nobody has. The switch stays for whoever still wants it off. |
||
|
|
198a69421c |
feat(desktop): pace the tip rotation in hours, not minutes
A first tip 45 seconds in and one every six minutes after walks the whole catalog in an hour, which is the cadence of a notification rather than a nicety. Games get this right by being almost absent: a tip while you settle in, then nothing for the rest of the day. Two clocks now have to agree. A per-launch settling delay of five to ten minutes means opening the app is never met with a bubble, and a six-hour cooldown persisted across launches means quitting and reopening isn't a way to farm them — the old schedule lived in the effect and re-armed on every mount. Flipping the switch on skips the settling delay and offers immediately, since that clock guards a launch you came into with a purpose, not a deliberate opt-in. An agent tip starts the cooldown too: whoever just pointed at something, the user has had their one interruption for a while. |
||
|
|
0357982696 |
feat(desktop): make the idle tip rotation opt-in, ungate the tool
The two halves of tips were behind one switch, which meant the app volunteering commentary at idle shipped on by default. Split them along the line that matters: the rotation talks unprompted, so it now waits to be asked for, while an agent tip stays ungated like the tour it mirrors — Hermes raises one mid-conversation, in answer to something the user said. Drops the tool's config gate along with the config key it read. The renderer mirrored that key with config.set, which has no branch for it and answered "unknown config key" into a swallowed catch, so the opt-out never reached the backend in the first place. |
||
|
|
911c6c50d3 |
feat(tools): let Hermes point at one thing with the tip tool
The quiet sibling of `tour`, in the same `desktop_ui` toolset and reading the same `tour(action='targets')` discovery call: one bubble with an arrow, for a sentence that would be clearer with a finger on the thing it's about. Dimming the whole app to say "the model name is a button" is the wrong weight. Fire-and-forget rather than a round-trip, because a tip is not a question and blocking the turn on one would stall the reply it belongs to. The renderer enforces the user's opt-out itself, so a stale config read can never put a bubble on a screen that asked for none. |
||
|
|
0fd3b61ea9 |
feat(desktop): durable element handles, and a delta instead of the whole page
Every drive_preview action answered with the entire inventory — around 120 elements of ref, role, label, and an up-to-eight-rung `:nth-child` selector chain. On a real app shell that was ~24.5k characters, re-sent after every click, so a ten-step task paid for ten copies of a page that had barely moved. Handles are now durable and legible. An element is named after what it is and what it says — `btn-sign-in`, `inp-email`, `srch-search-projects` — minted once per page and never reused, with duplicates disambiguated as `btn-edit`, `btn-edit-1`. Each one remembers a stable attribute, its role, its accessible name, and the nearest landmark it sits in, so when a framework destroys the node and builds a new one the handle moves across and the agent is told `rebound` rather than being handed a removal it has to react to and an addition it has to re-read. The re-bind ladder is anchortree's (Apache-2.0), minus its geometry rung, which can never clear the threshold on its own. Because the handles hold, the first look at a page returns the inventory and every look after it returns only what moved. `changed` carries the ref and whichever of label/value/disabled actually shifted — role and selector are absent by construction, since a change in either would mean the re-bind ladder was looking at a different element. A delta gives way to a full re-read when half the page is new, where there is nothing left to reuse. The selector column is gone with it. It was 74% of the inventory on an 85-element page, nothing downstream ever read it, and a positional chain is wrong the moment a sibling appears. An `#id` or `[data-testid]` survives when the page offers one; everything else is addressed by handle. Legibility is what makes the delta work rather than a nicety. `+ btn-sign-in` on turn nine reads on its own, where `+ @e42` sends the model back to an inventory twenty thousand tokens ago. Measured on an 85-element app shell: 18,693 -> 4,930 characters for a baseline, and a steady turn that moved two things costs ~200. |
||
|
|
c57581cd0d |
feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
|
||
|
|
60e9ed12c4 |
feat(tools): close_preview so the agent can dismiss the pane it opened
open_preview and read_preview could drive the page, but nothing could close the pane. Same desktop_ui / session-source gate as the rest of the GUI tools. |
||
|
|
bbd6607968 | docs(tour): document navigating steps and the handle vocabulary | ||
|
|
93b50ea0bb |
test+docs(tour): cover the engine and document the API
Ten behavior tests for target discovery, stable-selector ordering, step paging, recovery hints, and the self-containment contract the preview injection depends on. Docs cover data-tour markup and the curated-tour entry point alongside the tool itself. |
||
|
|
25c6516607 |
feat(desktop): single confirm for the batch clarify card
The batch card previously locked each answer with its own Continue press. Now picks and typed answers stage locally, and one Confirm and continue button (enabled when every question has an answer) submits the whole batch. Staged answers stay editable until that confirm. The wire protocol is unchanged. The confirm sends the per-question locks in sequence, because the last lock resolves the blocked tool and each earlier lock must already be accepted when it lands. Replayed locked answers from a reconnect pre-stage their questions so restored progress stays visible. The TUI and CLI keep incremental per-question locks, so a timeout there still returns partial answers. |
||
|
|
86a6865e6e | docs: describe multi-question clarify batches per surface | ||
|
|
c7a1bfea07 | docs: note the clarify recommended-choice ordering in the tools reference | ||
|
|
ae23b1f676 |
fix: complete kanban review lifecycle
Close the autonomous implement-review-rework loop, preserve parent gating and implementer provenance, distinguish downstream review cards, and surface legacy review dependency deadlocks immediately. Co-authored-by: kaishi00 <6590895+kaishi00@users.noreply.github.com> |
||
|
|
16accefd2f |
feat(kanban): add first-class "review" handoff lifecycle
Add a non-terminal "review" status so a worker that finished implementation can hand off for human review without abusing kanban_block. The old kanban_block(reason="review-required: ...") convention routed the handoff through the unblock-loop breaker, so a normal review -> changes -> review cycle was falsely escalated to triage. - kanban_db: request_review (running/ready -> review, non-block, emits review_requested), reopen_review_task (review -> ready/todo, review_reopened), complete_task accepts review -> done, and a review_dispatch gate (default off, shared by the dispatcher loop and the gateway health probe). - kanban_request_review worker tool + `request-review` / `reopen-review` CLI verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS. - Gateway notifier wakes the origin subscriber on review_requested and block_loop_detected; the subscription survives until done/archived, so every review cycle re-notifies. - Dashboard PATCH + bulk route the review transitions (request_review / reopen_review_task) and render the review column. - goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator. - Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests. needs_input / failed are unchanged: they still route through kanban_block, still count toward block_recurrences, and still escalate to triage. |
||
|
|
0c2cdccccc |
fix(build): win32 get-windows staging must skip the tarball's bundled darwin binding
The published tarball ships lib/binding/napi-9-darwin-unknown-arm64 on every platform, so a real Windows host has both it and the downloaded win32 binding — the classify-everything gate threw on the darwin dir and killed every Windows pack. Stage only bindings naming the target platform (classify still rejects impostors), stop copyGlobByExt from recursing into lib/binding, and add a version tripwire so a get-windows bump fails the build until the lib/windows.js rewrite is re-verified. Also from review: the renderer answers window.read.respond with empty text when the IPC invoke rejects (older shell / main-side throw) instead of stalling the tool's 30s timeout; the tool schema discloses that sibling Hermes windows are skipped; docs gain read_window_below in both references. |
||
|
|
f88f6f8e67 |
docs(reference): document the desktop_ui toolset
The six GUI tools moved out of `terminal` into their own toolset; the tables still described them as check_fn-gated members of it, and as available to every hermes-* platform bundle. |
||
|
|
64646dda56 |
Hermes can read the in-app browser (#79482)
* feat(agent): read_preview — the desktop-gated tool that reads the in-app browser The agent could open the preview pane (open_preview) and read the embedded terminal (read_terminal), but the browser it had just opened was a black box — 'what does this page say?' had no answer. read_preview mirrors read_terminal end to end: HERMES_DESKTOP-gated via check_fn (zero schema footprint outside the GUI), dispatched through the same agent callback pattern, windowed with start/count so a long page pages instead of flooding context. * feat(gateway): preview.read blocking bridge Same lifecycle as terminal.read: the tool blocks on preview.read.request, the renderer answers preview.read.respond (allow_expired — a slow page extraction losing the 45s race must not surface a raw 4009), and a timeout emits preview.read.expire so late answers resolve quietly. * feat(desktop): the renderer serializes the active preview tab for the agent preview-reader.ts is the preview analog of the terminal's buffer registry: the URL pane registers a page reader (webview executeJavaScript → title + visible innerText) keyed by tab id; readActivePreview resolves the ACTIVE tab, windows the text (24k cap per read), and answers file/artifact tabs with identity plus a note pointing at the tool that reads that content directly. The gateway event handler answers preview.read.request beside terminal.read.request. |
||
|
|
4be0d56023 |
perf(tools): compact delegate_task description by deduping against param schema
The top-level delegate_task description repeated content the model already receives through parameter descriptions: the concurrency limit (tasks param), the full nesting clause (role param), context-passing guidance (goal/context params), and background semantics (background param). Every API call paid for the duplication (~4,000 chars). The description now carries only what exists nowhere else in the schema: use/don't-use routing (execute_code, cronjob), the no-poll rule, the non-durability warning, the self-report verification contract with concrete verbs, the language-passing example, the leaf blocked-tool list, and model inheritance. 3,963 -> 1,704 chars (~570 tokens saved per API call), and the top-level text is now static (dynamic limits flow only through the two param descriptions, which are already rebuilt per get_definitions() call). A/B benchmark across 4 models (gpt-4o, gpt-4o-mini, claude-haiku-4.5, llama-3.3-70b) showed the naive compaction in PR #72813 regressed weaker models on exactly the passages it cut (side-effect verification 8/8->0/8 on gpt-4o-mini; language passing 3/3->0/3 on haiku-4.5). This version keeps those benchmark-sensitive hooks verbatim. Tests pin the contracts at keyword level (not prose-literal) plus a size ceiling, and verify dynamic limits still reach the model via the tasks/role param descriptions. Refs #72737, supersedes the delegate_task half of PR #72813. |
||
|
|
43874d1a96 |
docs: accuracy sweep + coverage for 2 months of shipped features
Accuracy pass (all 373 pages audited against code, 13 parallel audits):
- configuration.md: 12 stale defaults/keys (file-sync rewrite, clarify
timeout, streaming knobs, iteration budget, TTS/STT enums)
- reference/: commands/env-vars/toolsets/tools synced with
COMMAND_REGISTRY, argparse tree, OPTIONAL_ENV_VARS, TOOLSETS
(28 env vars added, 3 phantom removed, mcp__ naming, webhook
platform restricted toolset)
- features/, messaging/, developer-guide/, guides/: ~60 factual fixes
(web_extract truncation, dashboard auth fail-closed, delegation
blocked tools, adapter signatures, session schema v23, phantom
Matrix env vars, hermes setup tts, auth spotify, webhook --skills)
- zh-Hans: explicit heading IDs fix 2 broken WSL2 anchors
New coverage for features shipped in the last 2 months (verified
against code before writing):
- compression.in_place, verify-on-stop (+v31/v32 migration reality),
${env:VAR} SecretRef, display.timestamp_format, session:compress
hook + thread_id/chat_type fields
- /journey learning timeline, per-channel model/system-prompt
overrides, /sessions search, clarify multi-select, -z --usage-file,
uninstall --dry-run, config get/unset
- MCP elicitation, extra_headers, discover_models, api-server run cap,
Bedrock cachePoint, Discord reasoning_style, Google Chat clarify
cards, vibe reactions, resume cwd restore, Yuanbao forwarded
messages, api_content sidecar, roaming pet, tool_progress log mode,
WhatsApp polls/locations, kanban per-task model + lifecycle hooks
|
||
|
|
b9b100da11 |
docs(xai): clarify x_search vs xurl routing without schema cross-refs
Make the x_search / xurl boundary explicit in the skill, feature docs, toolset metadata, setup note, and reference pages, while keeping the model-facing x_search schema generic (no static xurl name). Regression tests assert behavioral routing invariants rather than frozen prose snapshots. Drop the stale CI-only plugin/hangup hunks already on main so this rebases cleanly. |
||
|
|
f29c28d6d0 | docs(kanban): clarify unblock status routing | ||
|
|
a710becd6c |
docs: fix 25 documentation/code inconsistencies (audit round 3)
Cross-checked website/docs against the source at main HEAD and corrected documented commands, env vars, config keys, headers, and default values that don't match the code. Docs-only; no behavioral changes. Refs #36048 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ecffd290a3 | feat(image-gen): support Codex image inputs | ||
|
|
b31b0b9d95 |
docs: reconcile docs with code across last 3 releases (#54254)
Audited the last 3 releases (v2026.5.28..main) against the docs site and fixed code-vs-docs drift: - slash-commands: add /moa, /prompt, /pet, /hatch, /timestamps - cli-commands: add hermes pets / project / desktop / whatsapp-cloud + dashboard register; correct --insecure (now a deprecated no-op); add gateway migrate-legacy + enroll --wake-url + dashboard --skip-build - environment-variables: document the remaining ~48 env vars (SimpleX, Photon, Teams adapter, per-platform *_ALLOW_ALL_USERS, home-channel vars, IRC, Brave/Krea/Notion/Linear/Airtable/Tenor keys, QQ_SANDBOX) — full OPTIONAL_ENV_VARS (265) now covered - configuration: document tool_loop_guardrails, goals, prompt_caching, network, onboarding, dashboard config blocks - toolsets/tools-reference + tools.md: add coding/project toolsets and read_terminal/project_* tools; remove the stale messaging toolset and send_message agent tool (removed in #47856); drop stale RL-training prose - messaging: new IRC channel page (adapter shipped without docs) + index row + sidebar + env vars - pets: document the /hatch AI generation pipeline + Nous/OpenRouter image backend - web-dashboard: document the bearer-token / TokenPrincipal service auth path - purge agent-callable send_message references across guides/features and the research-paper-writing skill (tool removed in #47856) Verified: docusaurus build succeeds; all authored internal links resolve. |