65 Commits

Author SHA1 Message Date
Austin Pickett
7cadfaa668 fix(api_server): honour include_compacted on /api/sessions/{id}/messages
The gateway REST messages endpoint parsed limit/offset/order but never
read include_compacted, so clients always got the live post-compaction
transcript. Parse it like the dashboard router and pass it through to
SessionDB.get_messages.

Fixes #101939
2026-09-29 16:30:33 -04:00
Brooklyn Nicholson
9cd114b499 feat(web): inline_images=false on GET /api/sessions/{id}/messages
The REST pages carry message content verbatim so the desktop's
extractEmbeddedImages can pull data URIs out of the text; a client reading
over a network had no way to ask for less (26.33 MiB per page read on the
measured conversation). inline_images=false (default true) routes content
through the same _coerce_message_text(image_urls=False) projection
session.resume uses, rendering [image] in place of the data URI so both
history surfaces agree by construction. Documented in the api-server
reference.

Fixes https://github.com/NousResearch/hermes-agent/issues/116511
2026-09-27 19:06:26 -05:00
luyifan
5af9ac8ed1 fix(api): classify failed tool stream events
(cherry picked from commit 5ab0a5ee26e61d4e96d65e1dacdc2b2e1e524f1b)
2026-09-25 21:30:22 +05:30
kshitijk4poor
eea4419ddc fix(api_server): honour platforms.api_server.tool_progress_events for Chat Completions SSE (#12020)
Strict OpenAI clients reject the named hermes.tool.progress SSE frames. Setting
tool_progress_events: false under platforms.api_server (loaded into
PlatformConfig.extra by from_dict) now drops them; default stays on.
Reimplements the intent of #42640 against the adapter config actually read in
production. Overlaps #49069 (erikerosev).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-25 14:27:23 +05:30
teknium1
eb8960fead fix(api_server): keep one memory provider per session across requests (#120116)
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).

ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
2026-09-23 05:41:31 -07:00
teknium1
6d335529ea fix(api_server): stamp the drain boundary on restart too; one HTTP-level invariant; docs
- request_restart() opens the same drain window as stop() (new turns refused,
  in-flight work awaited), so pollers of GET /v1/runs/{id} need the
  shutdown_requested_at marker from that moment as well; the marker is idempotent
  so stop() re-marking after the restart wait is a no-op.
- Collapse the three adapter-level tests into one invariant that drives the real
  GET /v1/runs/{id} route: live run keeps status=running but carries the marker
  (durably), a status set after the boundary inherits it, a terminal run never does.
- Shutdown-path fakes gain _mark_api_runs_shutdown_requested (the real mixin method
  where the fake already borrows _api_server_hook, a 0-stub where it has no adapter).
- Document the field on the runs API page.

Refs #115133.
2026-09-20 12:13:44 -07:00
teknium1
d19963782b fix(api-server): anchor the Responses current turn on this turn's user row, not a history prefix match
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.

Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.

The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).

Fixes #89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
2026-09-19 12:11:23 -07:00
Teknium
3df8ad0d7f Merge pull request #115903 from NousResearch/fix/boa-res-R2-app-server-interim
fix(api-server): codex commentary reaches streaming clients on session SSE, /v1/runs and /v1/responses (#67580, salvage #67593 #67613)
2026-09-19 11:14:10 -07:00
teknium1
12cc17fd76 chore: merge origin/main (resolve gateway/platforms/api_server.py, gateway/platforms/api_server_openai_routes.py) 2026-09-19 10:54:05 -07:00
teknium1
6b415f8754 chore: merge origin/main (resolve gateway/platforms/api_server_openai_routes.py) 2026-09-19 10:51:28 -07:00
teknium1
6b5a8fad38 fix(api_server): emit the run record's runtime in the canonical _sanitize_runtime_metadata shape
/v1/runs published a thinner {provider, model} twin under the same wire key that
/v1/chat/completions, /v1/responses and the session chat stream fill via
_sanitize_runtime_metadata (route_source, requested, cleaned ids). Route the served pair through
the same classmethod so one concept has one schema; requested comes from the run's
requested_model/requested_provider overrides, route_source from the model_routes/raw/global
vocabulary the other endpoints use. Doc example updated.
2026-09-19 10:14:38 -07:00
Tranquil-Flow
1b02df86e3 fix(api_server): GET /v1/runs/{id} reports the served fallback runtime and cache-read tokens
GET /v1/runs/{run_id} and the run.completed SSE event only echoed the
requested model and three token counters. After a fallback_providers
switch the run record still named the requested model and usage had no
cache figures, so a supervisor polling /v1/runs for cost attribution
booked the whole run to the wrong provider at the wrong price, with
cache reads counted as full-price input.

agent.provider / agent.model still hold the fallback pair when
run_conversation() returns (the primary is restored only at the start of
the next turn), and agent.session_cache_read_tokens has the cache reads.
Surface them on the completed run status and the terminal event:

- usage gains cache_read_tokens / cache_write_tokens (_USAGE_FIELDS)
- _run_agent_sync returns the served {provider, model} as a third value
  and _finish stamps it as `runtime` on both the pollable status and the
  run.completed event (the persisted idempotent record inherits it).

Ported onto the refactored _run_agent_sync / _finish shape from #102161;
the served pair is read from the agent (the reporter's diff in #102101)
rather than the turn record, so a run whose result dict lacks the keys is
still attributed. Tests trimmed to two invariants.

Fixes #102101
2026-09-19 10:14:38 -07:00
teknium1
0a846a2e2d feat(api-server): reasoning on non-streaming replies; echoed reasoning input items are ignored
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:

- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
  `choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
  completed `reasoning` output item ahead of the message (and ahead of that step's
  `function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
  from the assistant messages the agent already persisted (`build_assistant_message`
  stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
  in non-stream mode the agent fires the callback per provider delta AND once more with
  the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
  `conversation_history`. Responses SDK clients replay a prior response's `output`
  list as the next `input`; before, the item became an empty `user` message (in
  `input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
  (`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
  `output_item.done` item and the non-streaming shape.

Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
2026-09-19 09:48:35 -07:00
teknium1
70cee9e092 feat(api-server): stream model reasoning on /v1/chat/completions and /v1/responses
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).

- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
  `_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
  keeping them distinct from answer text. The lossy 500-char `reasoning.available`
  progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
  (the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
  (`output_item.added`, `reasoning_summary_part.added`,
  `reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
  `output_item.done`), closed before the next message/function_call item opens
  and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.

Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
2026-09-19 09:48:35 -07:00
teknium1
01f3f90ad4 chore: stack on #115797 to resolve api_server.py / api_server_openai_routes.py conflict
api_server.py keeps both interim_assistant_callback (#115903) and reasoning_callback (#115797) in every _create_agent/_run_agent slot.
api_server_openai_routes.py keeps both _ResponsesStream method sets and both __commentary__/__reasoning__ tag branches; emit_commentary closes any open reasoning item first.
2026-09-19 03:35:41 -07:00
teknium1
67173a73cc fix(api-server): /v1/runs emits message.interim; display.interim_assistant_messages gates all surfaces
POST /v1/runs + GET /v1/runs/{id}/events now carries mid-turn assistant
commentary as `message.interim` {text, already_streamed} — the same
contract the TUI gateway emits — so Runs clients can tell an active,
tool-heavy Codex turn from a stalled one instead of seeing only tool.*
events until run.completed (#67580).

APIServerAdapter._create_agent applies the same
`display.interim_assistant_messages` gate the messaging gateway and the
TUI apply (resolve_display_setting, per-platform override honoured):
when off, no callback is installed and nothing leaves the agent on any
of the three streaming surfaces. Dedup of repeated commentary stays in
agent/stream_delivery.py, where every surface already relies on it.

Tests: /v1/runs event contract; session SSE + /v1/responses item shape
plus the display gate (both red on the base commit). Docs: event
payloads on all three surfaces.

Co-authored-by: RoySRose <sungwook0115.kim@gmail.com>
2026-09-19 00:59:21 -07:00
teknium1
004c8a51f0 feat(api-server): reasoning on non-streaming replies; echoed reasoning input items are ignored
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:

- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
  `choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
  completed `reasoning` output item ahead of the message (and ahead of that step's
  `function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
  from the assistant messages the agent already persisted (`build_assistant_message`
  stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
  in non-stream mode the agent fires the callback per provider delta AND once more with
  the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
  `conversation_history`. Responses SDK clients replay a prior response's `output`
  list as the next `input`; before, the item became an empty `user` message (in
  `input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
  (`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
  `output_item.done` item and the non-streaming shape.

Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
2026-09-19 00:46:31 -07:00
teknium1
064756ca45 feat(api-server): opt-in cap on tool outputs in the stored /v1/responses history
The persisted conversation_history snapshot embeds the cumulative transcript
with every tool output verbatim, so a few large tool outputs made a single
response_store.db write ~677 KB (2.7x the configured rotation threshold)
while the response.completed payload for the same turn was already trimmed.

Add gateway.api_server.history_tool_output_max_chars (default 0 = store
verbatim, current behaviour). When set, tool rows and string tool-call
arguments longer than N chars are cut to the head plus the same
"...[K more chars]" marker _trim_tool_items uses, in a copied row, before the
snapshot is stored; user/assistant text and the agent's own in-memory
transcript are untouched. Opt-in rather than default because the stored
history is exactly what the model is replayed on the next chained turn.

Live (200 KB tool output, temp HERMES_HOME): stored row 200,659 B off ->
4,681 B with the cap at 4000.
2026-09-19 00:44:22 -07:00
teknium1
b957782917 feat(api-server): stream model reasoning on /v1/chat/completions and /v1/responses
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).

- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
  `_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
  keeping them distinct from answer text. The lossy 500-char `reasoning.available`
  progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
  (the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
  (`output_item.added`, `reasoning_summary_part.added`,
  `reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
  `output_item.done`), closed before the next message/function_call item opens
  and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.

Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
2026-09-18 23:58:24 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
5811653914 fix(api): a run admitted but not yet started settles interrupted at shutdown
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.

Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.

Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
2026-09-18 09:24:21 -07:00
teknium1
8d11af8497 docs(api-server): scope the concurrent-run cap to the routes that enforce it
The doc claimed the cap covers 'every endpoint that starts' a run, but
POST /api/jobs/{id}/run and /api/cron/fire start agent runs through the
cron scheduler, outside _run_agent, and are neither counted nor
429-guarded. Name the covered routes and say cron-triggered runs are
governed by cron's own limits so operators don't rely on a cap that
isn't applied there.
2026-09-16 17:05:32 -07:00
John Paul Soliva
9934972e21 fix(api-server): the session-chat routes spend the concurrent-run budget without respecting it
POST /api/sessions/{id}/chat and its /stream variant run their turn through _run_agent, so each
one already counts toward gateway.api_server.max_concurrent_runs. Neither ever called
_concurrency_limited_response, so the two routes could never be refused while their turns pushed
every other caller — /v1/chat/completions, /v1/responses, /v1/runs — into HTTP 429. A gateway at
its cap therefore refused the endpoints that check it and admitted the ones that do not.

Cross-machine agent DMs land on exactly these routes (hermes peer dm posts the turn to
/api/sessions/{id}/chat), so on a fleet gateway N simultaneous DMs started N full agent turns
regardless of the configured cap, each one an agent in its own executor thread.

Both routes now take the same guard the other run-starting routes take, before the session is
prepared. One table asserts the contract for the whole class: every route that starts a turn
answers 429 with Retry-After at the cap and starts no turn, and is admitted under it.

Docs: the cap's scope in the API server guide named only the OpenAI-compatible and Runs
endpoints, which is what the code did.
2026-09-16 17:05:32 -07:00
teknium1
05fac10a75 docs(api-server): MCP trust-gate consent surfaces as approval.request on /v1/runs
Document that an untrusted-server write-capable MCP tool now parks a run in
waiting_for_approval and is resolved through POST /v1/runs/{id}/approval,
the same bridge dangerous-command approvals already use.

Part of #111526
2026-09-15 19:06:27 -07:00
teknium1
339fa6d918 fix(gateway): bounded redacted result preview on tool.completed run events
Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
2026-09-15 18:42:10 -07:00
teknium1
0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1
c5c71ea1ad fix(docs): document the 10 s SSE keepalive comment for custom parsers 2026-09-15 18:18:30 -07:00
Teknium
d5926b2494 fix: persist API delegation units once without waking the model 2026-09-09 10:55:32 -07:00
Teknium
33c78d189e docs: describe browser session continuation preflight 2026-09-07 06:00:05 -07:00
Teknium
7a86397a46 fix(api_server): fail closed on unstamped runs; claim session-chat-stream run owner (#93689)
Port the run-ownership invariants from PR #93747 onto main's `_run_owners`
model in gateway/platforms/api_server_runs.py:

- `_request_owns_run` no longer admits run state that exists without an
  owner stamp. Under gateway.multiplex_profiles every served profile holds
  a valid key, so the "backward compatibility" branch made the boundary
  allow-all whenever provenance was missing. Unstamped state now fails
  closed; only an in-memory owner match or a durable idempotency record
  under the caller's own scope admits a run.
- POST /api/sessions/{id}/chat/stream claims `_run_owners` at the run mint,
  inside the request's profile scope, so its run is confined to the
  creating profile like /v1/runs.
- Owner release is tied to "no run-keyed state survives"
  (`_release_run_owner_if_forgotten`) and runs at every retirement point
  (task finally, SSE stream close, both sweep loops, chat-stream finally),
  not only the terminal-status sweep — no stranded entries, no stateful id
  ever left unowned.

Docs: note that runs are per-profile scoped (replaces the now-false
visibility admonition proposed in PR #92822).

Fixes #93689
Fixes #90415
Supersedes #93747, #93704, #92822

Co-authored-by: RickyYii <237135932+RickyYii@users.noreply.github.com>
Co-authored-by: liuhao1024 <11816344+liuhao1024@users.noreply.github.com>
2026-09-02 06:17:47 -07:00
Teknium
a2600740e8 feat(delegate): tag every subagent progress line with its batch id
Concurrent or nested delegation batches (a parent's 9-way fan-out plus a
child's own 3-way fan-out) printed interleaved `✓ [3/3]` / `✓ [3/9]` lines
with nothing identifying which batch each belongs to.

- CLI: batch header `🔀 [6a66] delegating 9 tasks`; completion lines and
  child tree-view lines become `[6a66 3/9]`; spinner remaining-count tagged.
- Relay: `delegation_id` rides on every `subagent.*` event (TUI gateway
  payload, api_server SSE subagent.start/complete).
- TUI: `[6a66 3/9]` prefix on /agents rows; Desktop Agents pane groups
  workers by exact delegation_id (heuristic shape/time grouping kept for
  older backends) and shows the tag on the group header.
- Tag = last 4 hex of the deleg_xxxxxxxx id (format_batch_tag), same id
  returned by the dispatch and used for cache/delegation/live/<id>/.
2026-09-02 01:06:24 -07:00
David Dudok de Wit
e7433910e9 feat(bot-mode): add scoped cross-gateway Group Chat transport 2026-08-31 01:04:11 -07:00
abundantbeing
2039b572f5 fix(browser): keep bound controller routing authoritative
Generic Hermes callers still use the existing browser backend when extension control is disabled or no server-bound controller identity exists.

Once the gateway binds a controller identity, missing scope, disconnect, or capability loss now fail closed instead of silently switching a control-this-tab request to another local or cloud browser. Covers the schema-build to dispatch disconnect race.
2026-08-21 22:33:45 +05:30
abundantbeing
095a1d078c fix(browser): preserve controller work across reconnects
Treat unexpected controller transport loss as recoverable until each command's original deadline. Same-identity reconnects refresh transport and capability state, flush deferred cancels before new dispatch, and can complete already-started work.

Keep explicit detach and different controller/browser identity replacement terminal, owner-gate every inbound lifecycle frame, distinguish slow in-flight WebSocket writes from real send failures, and exclude browser-control session identity from shared shell snapshots.
2026-08-21 22:33:45 +05:30
abundantbeing
d524cc9a16 fix(browser): harden extension controller routing
Keep extension control opt-in and preserve existing browser backends unless an exact server-bound controller is available. Centralize protocol and capability admission across API and dashboard transports, make selected-controller results authoritative, bypass stale availability caches only inside bound requests, and serialize structured results for the existing tool contract.

Add a real browser_snapshot route-table/WebSocket E2E, strict admission and ownership regressions, public configuration and protocol documentation, and tests proving feature-off/no-controller compatibility.
2026-08-21 22:33:45 +05:30
Teknium
a51a4cb096 fix(api-server): mark replayed tool calls completed in Responses output items
The non-streaming /v1/responses path built function_call and
function_call_output output items with no status field (and no item id),
while the SSE streaming path correctly emits status in_progress ->
completed. Spec-strict OpenAI clients reading the non-streaming output
array could interpret the status-less function_call items as pending
calls the CLIENT must execute — but these tools were already executed
server-side by the Hermes agent and are replayed for structured tool UI
only. Reported by a community user whose GPT-5.6 client concluded 'a
server should not tell an OpenAI client to execute a tool the server
already executed itself'.

- _extract_output_items now stamps status: completed and spec-shaped
  item ids (fc_/fco_) on replayed items, matching the streaming path
- test updated to pin status + id shape
- docs example updated + explicit note that output tool calls are
  replayed, never pending
2026-08-08 05:22:02 -07:00
Teknium
43874d1a96 docs: accuracy sweep + coverage for 2 months of shipped features
Accuracy pass (all 373 pages audited against code, 13 parallel audits):
- configuration.md: 12 stale defaults/keys (file-sync rewrite, clarify
  timeout, streaming knobs, iteration budget, TTS/STT enums)
- reference/: commands/env-vars/toolsets/tools synced with
  COMMAND_REGISTRY, argparse tree, OPTIONAL_ENV_VARS, TOOLSETS
  (28 env vars added, 3 phantom removed, mcp__ naming, webhook
  platform restricted toolset)
- features/, messaging/, developer-guide/, guides/: ~60 factual fixes
  (web_extract truncation, dashboard auth fail-closed, delegation
  blocked tools, adapter signatures, session schema v23, phantom
  Matrix env vars, hermes setup tts, auth spotify, webhook --skills)
- zh-Hans: explicit heading IDs fix 2 broken WSL2 anchors

New coverage for features shipped in the last 2 months (verified
against code before writing):
- compression.in_place, verify-on-stop (+v31/v32 migration reality),
  ${env:VAR} SecretRef, display.timestamp_format, session:compress
  hook + thread_id/chat_type fields
- /journey learning timeline, per-channel model/system-prompt
  overrides, /sessions search, clarify multi-select, -z --usage-file,
  uninstall --dry-run, config get/unset
- MCP elicitation, extra_headers, discover_models, api-server run cap,
  Bedrock cachePoint, Discord reasoning_style, Google Chat clarify
  cards, vibe reactions, resume cwd restore, Yuanbao forwarded
  messages, api_content sidecar, roaming pet, tool_progress log mode,
  WhatsApp polls/locations, kanban per-task model + lifecycle hooks
2026-07-29 08:48:05 -07:00
Teknium
88ff722f94 docs(api-server): document profile-bound HTTP auth from #72285
The multiplexed listener now rejects the default API_SERVER_KEY on
/p/<profile>/ prefixes (fail-closed per-profile keys). Add the
multi-profile routing section with an explicit breaking-change callout
for the next release notes.
2026-07-28 21:45:44 -07:00
teknium1
4b4d2ae4cd docs(delegation,api): document stall detection, timeout metadata, /agents live status, runs-stream subagent lifecycle
Catch the docs up to this week's delegation work:

- delegation.md: structured timeout metadata fields (timeout_seconds/
  timed_out_after_seconds/timeout_phase, #72403); new 'Stall Detection
  for Background Subagents' section (progress-based monitor, thresholds,
  grace window, stalled event metadata, root-cause fix note — #72227/
  #72300/#72412); /agents per-child live activity on CLI + gateway;
  all-thread diagnostic dump note.
- api-server.md: /v1/runs events stream now documents subagent.start/
  subagent.complete forwarding, redaction, child_session_id, and the
  deliberate exclusion of per-tool child events (#72406).
- sidebars.ts: register the subagent-lifecycle-api developer guide page
  (shipped in #72501 but absent from the sidebar).

Docusaurus build verified.
2026-07-27 11:33:53 -07:00
abundantbeing
4a9447c726 fix(api-server): expose model options inventory
Add authenticated GET /api/model/options to the gateway API server,
sharing the dashboard/TUI picker payload builder so external clients
can sync to the user's configured Hermes provider catalog instead of
scraping the single OpenAI-compatible /v1/models alias.

- new shared hermes_cli.inventory.build_model_options_payload() wraps
  build_models_payload with the stable picker shape and safe
  custom-provider probe policy (probe current only on normal open,
  probe all + cache bust on explicit refresh)
- dashboard web_server and TUI gateway model.options refactored onto
  the shared builder; dashboard build moved off the event loop via
  run_in_threadpool
- capabilities endpoint advertises model_options
- docs for both API server and programmatic integration

Salvaged from PR #54689 by @abundantbeing.
2026-07-24 11:20:07 -07:00
teknium1
9f384783e7 fix(api): gate bare-model passthrough + route-alias model leak
Follow-ups on the salvaged #54426 routing contract:

- Bare `model` without `provider` on the OpenAI-compatible endpoints
  (/v1/chat/completions, /v1/responses) is now opt-in via
  gateway.platforms.api_server.direct_model_requests (default off) —
  generic OpenAI clients hardcode model names ('gpt-4o', ...) and
  existing deployments rely on those falling back to the gateway
  default. Explicit `provider` requests and the Hermes-native
  session-chat + /v1/runs surfaces are always honored.
  Idea credit: PR #22825 by @mssteuer.
- A model_routes alias with no `model` key can no longer leak the
  alias string as the executing model name (defensive; parse-time
  validation already drops such routes).
- Fix mis-indented _run_agent call args in _handle_session_chat_stream.
- Docs: document the opt-in flag.
2026-07-24 09:54:08 -07:00
abundantbeing
d66a82000c feat(api): honor provider-aware request routing
Carry model, provider, and model_options through the API server's
execution surfaces (session chat, Chat Completions, Responses, /v1/runs)
without mutating global configuration. Precedence: session /model
override -> model_routes alias -> direct request selection -> global
defaults. Conflicting route/provider mixes fail closed with 400.
model_options stays request-scoped regardless of which selection wins.

Salvaged from PR #54426 by @abundantbeing.
2026-07-24 09:54:08 -07:00
Teknium
6ced760a3d test(gateway): cover gateway.api_server YAML discovery + extra bridge; docs: document config.yaml support
Follow-up for salvaged PR #66633 (fixes #66630):
- 5 tests: nested gateway.api_server discovery, port/key/host/model_name
  bridged into extra, explicit extra wins over top-level key, non-platform
  gateway keys (streaming/timeout) not misparsed, gateway.platforms.api_server
  path regression-guarded.
- docs: api-server.md (en + zh-Hans) now documents the gateway.api_server
  YAML section instead of 'not yet supported'.
2026-07-23 11:54:32 -07:00
teknium1
837077dfae fix(api): stop producers after run transport expires 2026-07-12 04:15:47 -07:00
teknium1
6142203bd7 fix(gateway): ground readiness in live runtime state 2026-07-11 08:42:21 -07:00
Teknium
2d099fed1e docs: deep audit — registry drift, stale claims, 2-week PR coverage, dashboard screenshot (#40952)
Full-corpus correctness audit of the hand-written docs against the codebase,
plus a 2-week merged-PR coverage sweep and one live dashboard screenshot.

Correctness (verified against COMMAND_REGISTRY / PROVIDER_REGISTRY / TOOLSETS /
tools.registry / DEFAULT_CONFIG / source):
- reference: add /version slash command, context_engine toolset, openai-api +
  novita-ai to --provider; fix tool count 64->71; model_catalog ttl 24->1;
  add profile describe to summary table; add real provider env vars
  (LM_API_KEY/LM_BASE_URL, KIMI_CODING_API_KEY, ALIBABA_CODING_PLAN_*,
  ANTHROPIC_BASE_URL, COPILOT_API_BASE_URL); fix faq "Windows: not natively".
- user-guide: fix broken `hermes -w -q` (->-z) and `hermes logs --tail` (->-f);
  language list 8->16; aux slots 8->11; docker separate-dashboard claim;
  _SECURITY_ARGS -> _BASE_SECURITY_ARGS.
- features: curator prune_builtins truth + missing CLI verbs; codex-runtime aux
  keys (context_compression->compression, vision_detect->vision); kanban
  terminate endpoint + promote/reassign/schedule/diagnostics/edit + per-profile
  cap; mcp mTLS (client_cert/client_key); built-in-plugins nemo_relay +
  teams_pipeline; api-server run approval endpoint; computer-use frontmatter.
- features N-Z + integrations: StepFun step-3-mini->step-3.5-flash; web-search
  backends 4->8; tool-gateway image-model IDs; voice-mode STT/TTS enums; remove
  phantom `rl` toolset; nous-portal status subcommand.
- messaging: WeCom typing/streaming cols; telegram transport default edit->auto;
  sms host default; simplex/ntfy `gateway setup` + pairing approve; line
  smart-chunking; matrix MATRIX_DM_AUTO_THREAD.
- developer-guide: build-a-plugin code examples (register_command signature,
  ContextEngine/ImageGenProvider/MemoryProvider ABCs); model-provider-plugin
  entry-point group hermes.plugins->hermes_agent.plugins; PLUGIN.yaml->plugin.yaml;
  agent-loop stale LOC; web-search-provider phantom crawl().

PR coverage (2-week window, 149 feat PRs):
- desktop.md refreshed for ~15 shipped features (zh-Hans switcher, rebindable
  shortcuts + zoom + Cmd+K, status-bar model picker + YOLO toggle, session-by-id
  + archive, multi-profile concurrent + cross-profile @session, composer history,
  Providers pane, per-profile remote hosts, Grok OAuth, aux-pin warning).
- configuration.md gateway-streaming default corrected to per-platform.
- tool-gateway.md free tool pool entitlement note.

Media:
- New /img/dashboard/admin-config.png — live dashboard Config admin page
  (captured from a clean profile, no secrets/personalization).
2026-06-07 01:39:06 -07:00
Teknium
8b6beaab5f docs: 30-day overhaul — correctness audit, PR coverage, Nous Portal weave, sidebar reorg (#33782)
* docs(audit): correctness pass across getting-started, reference, features, messaging, developer-guide, guides, integrations, user-guide

* docs: add PR coverage for last 30d + Nous Portal weave + nav reorg + build fixes

- Add docs for top user-visible PRs that shipped without docs (api-server
  session control, kanban features, telegram pin/edit, provider client tag,
  xAI retired-model migration, cron name lookup, --branch update flag, etc.)
- Apply Nous Portal weave across 23 pages (tasteful one-liners on
  getting-started/learning-path, configuration, overview, vision, x-search,
  credential-pools, provider-routing, cron, codex-runtime, profiles, docker,
  messaging/index, multiple guides, plus FAQ + index promotion)
- Reorganize sidebar: split Messaging into Popular/M365/Chinese/Other,
  Reference into Command/Configuration/Tools-Skills sub-categories, add
  orphan developer-guide pages (web-search-provider-plugin,
  browser-supervisor), move features from Integrations back to Features,
  fold lone spotify into Media & Web.
- Regenerate skill stubs + catalogs (kanban-codex-lane, hermes-s6-container-
  supervision, web-pentest)
- Fix broken anchor links (security/cron, configuration/fallback, telegram
  large-files, adding-platform-adapters step-by-step)
2026-05-28 02:41:36 -07:00
Dusk1e
1a9ef83147 fix(security): require API_SERVER_KEY before dispatching API server work 2026-05-28 00:25:08 -07:00
Teknium
1d5deac346 fix(website): cross-locale doc links + drop empty ko locale (#31895)
The locale switcher appeared broken because hardcoded markdown links
(`](/docs/X)`) got double-prefixed by Docusaurus to `/docs/<locale>/docs/X`
(404) in non-English locales, and the MDX hero `<a href>` on the index
page escaped locale routing entirely.

Changes:
- Rewrite 922 `](/docs/X)` -> `](/X)` across 166 docs files (strip trailing
  .md too). Docusaurus prepends locale + baseUrl itself.
- docs/index.md -> index.mdx; hero "Get Started" anchor -> Docusaurus
  <Link> so it stays inside the active locale.
- Drop `ko` locale entirely from docusaurus.config.ts + delete i18n/ko/
  (4 stale auto-translated kanban pages, <2% coverage, misleading).

Verified `npm run build` succeeds for both en and zh-Hans; `build/zh-Hans/
index.html` has no /docs/zh-Hans/docs/... double-prefixed paths.

PR2 will translate the 335 English docs into i18n/zh-Hans/.
2026-05-24 23:16:20 -07:00
Teknium
64b3eb0dd7 docs: surface Nous Portal on pages where it solves a real problem the page describes (#30874)
Follow-up to #30869. Adds Portal mentions on user-facing pages that
naturally call for an LLM + tool credentials but didn't previously
acknowledge Portal as a one-stop option.

- getting-started/installation.md: tip after the 'after install' block
  pointing at 'hermes setup --portal' for users who want everything wired
  at once instead of piecewise via 'hermes model' + 'hermes tools'.
- user-guide/configuring-models.md: small tip near the top — the page is
  literally about provider/model choice and previously had zero Portal
  mention.
- user-guide/features/voice-mode.md: Prerequisites need both an LLM and
  TTS — a Portal subscription is the single setup that covers both.
- user-guide/features/batch-processing.md: highlights Portal as a
  predictable-cost option for parallel agent runs that hit many APIs.
- user-guide/features/api-server.md: backend needs models + tools; one
  Portal sub gives a fully-equipped OpenAI-compatible endpoint.
- user-guide/windows-native.md: early-beta users on Windows benefit most
  from skipping per-tool Windows-key-juggling.
- integrations/providers.md: updates the existing Tool Gateway tip and
  the Nous Portal section to mention the new commands.
- user-guide/features/fallback-providers.md: Nous row in the provider
  table now lists 'hermes setup --portal' as the fresh-install path.

Tone discipline: one Portal mention per page, concrete CLI commands
(no marketing copy), always solving a problem the page itself sets up.
2026-05-23 02:47:53 -07:00