The gateway REST messages endpoint parsed limit/offset/order but never
read include_compacted, so clients always got the live post-compaction
transcript. Parse it like the dashboard router and pass it through to
SessionDB.get_messages.
Fixes#101939
The REST pages carry message content verbatim so the desktop's
extractEmbeddedImages can pull data URIs out of the text; a client reading
over a network had no way to ask for less (26.33 MiB per page read on the
measured conversation). inline_images=false (default true) routes content
through the same _coerce_message_text(image_urls=False) projection
session.resume uses, rendering [image] in place of the data URI so both
history surfaces agree by construction. Documented in the api-server
reference.
Fixes https://github.com/NousResearch/hermes-agent/issues/116511
Strict OpenAI clients reject the named hermes.tool.progress SSE frames. Setting
tool_progress_events: false under platforms.api_server (loaded into
PlatformConfig.extra by from_dict) now drops them; default stays on.
Reimplements the intent of #42640 against the adapter config actually read in
production. Overlaps #49069 (erikerosev).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).
ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
- request_restart() opens the same drain window as stop() (new turns refused,
in-flight work awaited), so pollers of GET /v1/runs/{id} need the
shutdown_requested_at marker from that moment as well; the marker is idempotent
so stop() re-marking after the restart wait is a no-op.
- Collapse the three adapter-level tests into one invariant that drives the real
GET /v1/runs/{id} route: live run keeps status=running but carries the marker
(durably), a status set after the boundary inherits it, a terminal run never does.
- Shutdown-path fakes gain _mark_api_runs_shutdown_requested (the real mixin method
where the fake already borrows _api_server_hook, a 0-stub where it has no adapter).
- Document the field on the runs API page.
Refs #115133.
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.
Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.
The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).
Fixes#89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
/v1/runs published a thinner {provider, model} twin under the same wire key that
/v1/chat/completions, /v1/responses and the session chat stream fill via
_sanitize_runtime_metadata (route_source, requested, cleaned ids). Route the served pair through
the same classmethod so one concept has one schema; requested comes from the run's
requested_model/requested_provider overrides, route_source from the model_routes/raw/global
vocabulary the other endpoints use. Doc example updated.
GET /v1/runs/{run_id} and the run.completed SSE event only echoed the
requested model and three token counters. After a fallback_providers
switch the run record still named the requested model and usage had no
cache figures, so a supervisor polling /v1/runs for cost attribution
booked the whole run to the wrong provider at the wrong price, with
cache reads counted as full-price input.
agent.provider / agent.model still hold the fallback pair when
run_conversation() returns (the primary is restored only at the start of
the next turn), and agent.session_cache_read_tokens has the cache reads.
Surface them on the completed run status and the terminal event:
- usage gains cache_read_tokens / cache_write_tokens (_USAGE_FIELDS)
- _run_agent_sync returns the served {provider, model} as a third value
and _finish stamps it as `runtime` on both the pollable status and the
run.completed event (the persisted idempotent record inherits it).
Ported onto the refactored _run_agent_sync / _finish shape from #102161;
the served pair is read from the agent (the reporter's diff in #102101)
rather than the turn record, so a run whose result dict lacks the keys is
still attributed. Tests trimmed to two invariants.
Fixes#102101
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
api_server.py keeps both interim_assistant_callback (#115903) and reasoning_callback (#115797) in every _create_agent/_run_agent slot.
api_server_openai_routes.py keeps both _ResponsesStream method sets and both __commentary__/__reasoning__ tag branches; emit_commentary closes any open reasoning item first.
POST /v1/runs + GET /v1/runs/{id}/events now carries mid-turn assistant
commentary as `message.interim` {text, already_streamed} — the same
contract the TUI gateway emits — so Runs clients can tell an active,
tool-heavy Codex turn from a stalled one instead of seeing only tool.*
events until run.completed (#67580).
APIServerAdapter._create_agent applies the same
`display.interim_assistant_messages` gate the messaging gateway and the
TUI apply (resolve_display_setting, per-platform override honoured):
when off, no callback is installed and nothing leaves the agent on any
of the three streaming surfaces. Dedup of repeated commentary stays in
agent/stream_delivery.py, where every surface already relies on it.
Tests: /v1/runs event contract; session SSE + /v1/responses item shape
plus the display gate (both red on the base commit). Docs: event
payloads on all three surfaces.
Co-authored-by: RoySRose <sungwook0115.kim@gmail.com>
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
The persisted conversation_history snapshot embeds the cumulative transcript
with every tool output verbatim, so a few large tool outputs made a single
response_store.db write ~677 KB (2.7x the configured rotation threshold)
while the response.completed payload for the same turn was already trimmed.
Add gateway.api_server.history_tool_output_max_chars (default 0 = store
verbatim, current behaviour). When set, tool rows and string tool-call
arguments longer than N chars are cut to the head plus the same
"...[K more chars]" marker _trim_tool_items uses, in a copied row, before the
snapshot is stored; user/assistant text and the agent's own in-memory
transcript are untouched. Opt-in rather than default because the stored
history is exactly what the model is replayed on the next chained turn.
Live (200 KB tool output, temp HERMES_HOME): stored row 200,659 B off ->
4,681 B with the cap at 4000.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.
Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.
Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
The doc claimed the cap covers 'every endpoint that starts' a run, but
POST /api/jobs/{id}/run and /api/cron/fire start agent runs through the
cron scheduler, outside _run_agent, and are neither counted nor
429-guarded. Name the covered routes and say cron-triggered runs are
governed by cron's own limits so operators don't rely on a cap that
isn't applied there.
POST /api/sessions/{id}/chat and its /stream variant run their turn through _run_agent, so each
one already counts toward gateway.api_server.max_concurrent_runs. Neither ever called
_concurrency_limited_response, so the two routes could never be refused while their turns pushed
every other caller — /v1/chat/completions, /v1/responses, /v1/runs — into HTTP 429. A gateway at
its cap therefore refused the endpoints that check it and admitted the ones that do not.
Cross-machine agent DMs land on exactly these routes (hermes peer dm posts the turn to
/api/sessions/{id}/chat), so on a fleet gateway N simultaneous DMs started N full agent turns
regardless of the configured cap, each one an agent in its own executor thread.
Both routes now take the same guard the other run-starting routes take, before the session is
prepared. One table asserts the contract for the whole class: every route that starts a turn
answers 429 with Retry-After at the cap and starts no turn, and is admitted under it.
Docs: the cap's scope in the API server guide named only the OpenAI-compatible and Runs
endpoints, which is what the code did.
Document that an untrusted-server write-capable MCP tool now parks a run in
waiting_for_approval and is resolved through POST /v1/runs/{id}/approval,
the same bridge dangerous-command approvals already use.
Part of #111526
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.
The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.
Salvages #111821 (@KoNit-K), part of #111815.
Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.
The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.
`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.
CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.
Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:
- `_request_owns_run` no longer admits run state that exists without an
owner stamp. Under gateway.multiplex_profiles every served profile holds
a valid key, so the "backward compatibility" branch made the boundary
allow-all whenever provenance was missing. Unstamped state now fails
closed; only an in-memory owner match or a durable idempotency record
under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
inside the request's profile scope, so its run is confined to the
creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
(`_release_run_owner_if_forgotten`) and runs at every retirement point
(task finally, SSE stream close, both sweep loops, chat-stream finally),
not only the terminal-status sweep — no stranded entries, no stateful id
ever left unowned.
Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).
Fixes#93689Fixes#90415
Supersedes #93747, #93704, #92822
Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.
- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
workers by exact delegation_id (heuristic shape/time grouping kept for
older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
returned by the dispatch and used for cache/delegation/live/<id>/.
Generic Hermes callers still use the existing browser backend when extension control is disabled or no server-bound controller identity exists.
Once the gateway binds a controller identity, missing scope, disconnect, or capability loss now fail closed instead of silently switching a control-this-tab request to another local or cloud browser. Covers the schema-build to dispatch disconnect race.
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.
Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.
Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
The non-streaming /v1/responses path built function_call and
function_call_output output items with no status field (and no item id),
while the SSE streaming path correctly emits status in_progress ->
completed. Spec-strict OpenAI clients reading the non-streaming output
array could interpret the status-less function_call items as pending
calls the CLIENT must execute — but these tools were already executed
server-side by the Hermes agent and are replayed for structured tool UI
only. Reported by a community user whose GPT-5.6 client concluded 'a
server should not tell an OpenAI client to execute a tool the server
already executed itself'.
- _extract_output_items now stamps status: completed and spec-shaped
item ids (fc_/fco_) on replayed items, matching the streaming path
- test updated to pin status + id shape
- docs example updated + explicit note that output tool calls are
replayed, never pending
The multiplexed listener now rejects the default API_SERVER_KEY on
/p/<profile>/ prefixes (fail-closed per-profile keys). Add the
multi-profile routing section with an explicit breaking-change callout
for the next release notes.
Add authenticated GET /api/model/options to the gateway API server,
sharing the dashboard/TUI picker payload builder so external clients
can sync to the user's configured Hermes provider catalog instead of
scraping the single OpenAI-compatible /v1/models alias.
- new shared hermes_cli.inventory.build_model_options_payload() wraps
build_models_payload with the stable picker shape and safe
custom-provider probe policy (probe current only on normal open,
probe all + cache bust on explicit refresh)
- dashboard web_server and TUI gateway model.options refactored onto
the shared builder; dashboard build moved off the event loop via
run_in_threadpool
- capabilities endpoint advertises model_options
- docs for both API server and programmatic integration
Salvaged from PR #54689 by @abundantbeing.
Follow-ups on the salvaged #54426 routing contract:
- Bare `model` without `provider` on the OpenAI-compatible endpoints
(/v1/chat/completions, /v1/responses) is now opt-in via
gateway.platforms.api_server.direct_model_requests (default off) —
generic OpenAI clients hardcode model names ('gpt-4o', ...) and
existing deployments rely on those falling back to the gateway
default. Explicit `provider` requests and the Hermes-native
session-chat + /v1/runs surfaces are always honored.
Idea credit: PR #22825 by @mssteuer.
- A model_routes alias with no `model` key can no longer leak the
alias string as the executing model name (defensive; parse-time
validation already drops such routes).
- Fix mis-indented _run_agent call args in _handle_session_chat_stream.
- Docs: document the opt-in flag.
Carry model, provider, and model_options through the API server's
execution surfaces (session chat, Chat Completions, Responses, /v1/runs)
without mutating global configuration. Precedence: session /model
override -> model_routes alias -> direct request selection -> global
defaults. Conflicting route/provider mixes fail closed with 400.
model_options stays request-scoped regardless of which selection wins.
Salvaged from PR #54426 by @abundantbeing.
The locale switcher appeared broken because hardcoded markdown links
(`](/docs/X)`) got double-prefixed by Docusaurus to `/docs/<locale>/docs/X`
(404) in non-English locales, and the MDX hero `<a href>` on the index
page escaped locale routing entirely.
Changes:
- Rewrite 922 `](/docs/X)` -> `](/X)` across 166 docs files (strip trailing
.md too). Docusaurus prepends locale + baseUrl itself.
- docs/index.md -> index.mdx; hero "Get Started" anchor -> Docusaurus
<Link> so it stays inside the active locale.
- Drop `ko` locale entirely from docusaurus.config.ts + delete i18n/ko/
(4 stale auto-translated kanban pages, <2% coverage, misleading).
Verified `npm run build` succeeds for both en and zh-Hans; `build/zh-Hans/
index.html` has no /docs/zh-Hans/docs/... double-prefixed paths.
PR2 will translate the 335 English docs into i18n/zh-Hans/.
Follow-up to #30869. Adds Portal mentions on user-facing pages that
naturally call for an LLM + tool credentials but didn't previously
acknowledge Portal as a one-stop option.
- getting-started/installation.md: tip after the 'after install' block
pointing at 'hermes setup --portal' for users who want everything wired
at once instead of piecewise via 'hermes model' + 'hermes tools'.
- user-guide/configuring-models.md: small tip near the top — the page is
literally about provider/model choice and previously had zero Portal
mention.
- user-guide/features/voice-mode.md: Prerequisites need both an LLM and
TTS — a Portal subscription is the single setup that covers both.
- user-guide/features/batch-processing.md: highlights Portal as a
predictable-cost option for parallel agent runs that hit many APIs.
- user-guide/features/api-server.md: backend needs models + tools; one
Portal sub gives a fully-equipped OpenAI-compatible endpoint.
- user-guide/windows-native.md: early-beta users on Windows benefit most
from skipping per-tool Windows-key-juggling.
- integrations/providers.md: updates the existing Tool Gateway tip and
the Nous Portal section to mention the new commands.
- user-guide/features/fallback-providers.md: Nous row in the provider
table now lists 'hermes setup --portal' as the fresh-install path.
Tone discipline: one Portal mention per page, concrete CLI commands
(no marketing copy), always solving a problem the page itself sets up.