Files
hermes-agent/website/docs/developer-guide/programmatic-integration.md

16 KiB

sidebar_position, title, description
sidebar_position title description
8 Programmatic Integration Three protocols for driving hermes-agent from external programs: ACP, the TUI gateway JSON-RPC, and the OpenAI-compatible HTTP API

Programmatic Integration

Hermes ships three protocols for driving the agent from external programs — IDE plugins, custom UIs, CI pipelines, embedded sub-agents. Pick the one that matches your transport and consumer.

Protocol Transport Best for Defined by
ACP JSON-RPC over stdio IDE clients (VS Code, Zed, JetBrains) that already speak the Agent Client Protocol acp_adapter/
TUI gateway JSON-RPC over stdio (or WebSocket) Custom hosts that want fine-grained control of sessions, slash commands, approvals, and streaming events tui_gateway/server.py
API server HTTP + Server-Sent Events OpenAI-compatible frontends (Open WebUI, LobeChat, LibreChat…) and language-agnostic web clients gateway/platforms/api_server.py

All three drive the same AIAgent core. They differ only in wire format and which set of features they expose.


ACP (Agent Client Protocol)

hermes acp starts a stdio JSON-RPC server speaking ACP. Used in production by VS Code (Zed Industries' ACP extension), Zed, and any JetBrains IDE with an ACP plugin.

Capabilities exposed: session creation, prompt submission, streaming agent message chunks, tool-call events, permission requests, session fork, cancel, and authentication. Tool output is rendered into ACP Diff/ToolCall content blocks the IDE understands.

Full lifecycle, event bridge, and approval flow: ACP Internals.

hermes acp                  # serve ACP on stdio
hermes acp --check          # verify ACP dependencies and adapter imports
hermes acp --setup          # interactive provider/model setup for ACP terminal auth

TUI Gateway JSON-RPC

tui_gateway/server.py is the protocol the Ink TUI (hermes --tui) and the embedded dashboard PTY bridge talk to. Any external host can speak the same protocol over stdio (or WebSocket via tui_gateway/ws.py).

Method catalog (selected)

prompt.submit           prompt.background       session.steer
session.create          session.list            session.active_list
session.activate        session.close           session.interrupt
session.history         session.compress        session.branch
session.title           session.usage           session.status
clarify.lock            config.set / config.get commands.catalog
command.resolve         command.dispatch        cli.exec
reload.mcp              reload.env              process.stop
delegation.status       subagent.interrupt      subagent.steer
spawn_tree.save / list / load
terminal.resize         clipboard.paste         image.attach

session.active_list, session.activate, and session.close are the process-local live-session controls used by the TUI session switcher. Use session.list / /resume for saved transcript discovery; use the active-session methods only for sessions that are currently open in the TUI gateway process.

Within one authenticated gateway, resuming or activating a live session attaches another event subscriber rather than replacing the previous connection. Streaming and terminal events go to all attached clients; disconnecting one client does not end a session another client is viewing. Existing submit exclusivity and configured busy-input policy remain in force. Attached clients can steer the session's subagents; browser-controller results still require the connection that registered that controller. This does not enable independent gateway processes to write the same session, nor does it imply durable prompt admission across an owner restart.

Rewinding history on prompt.submit

A rewind / edit / regenerate is a prompt.submit that drops part of the stored transcript before running the new turn. Because that write is a destructive rewrite of the session's durable rows, the gateway honors it only when the client states its intent:

Parameter Meaning
truncate_before_user_ordinal Zero-based index of the user turn to cut at. Everything from that turn onward is dropped. Display-only timeline rows (display_kind) are not counted. Must be a real integer — a JSON boolean is refused with code 4004.
truncate_before_row_id Integer SQLite row ID (messages.id / row_id) of the target user turn to cut at. Preferred durable address. When both ordinal and row ID are provided, gateway verifies they match (returning 4030 on mismatch). An unknown/stale row ID is refused with 4018 — it does not fall back to the ordinal.
confirm_truncate Required whenever an ordinal, message ID, or row ID is sent. Declares that this submit really is a rewind, not an ordinary send that happens to carry leftover parameters. Sending it without a target is refused with code 4004.
confirm_empty_truncate Additionally required when the cut would leave the transcript empty (ordinal 0).

A truncation parameter without confirm_truncate is refused with code 4004 or 4029 and nothing is written. Hosts that implement rewind must set the flag at the moment the user asks for it, and must never keep truncation parameters in state across ordinary submits. Prefer truncate_before_row_id (from resume row_id / _row_id) over ordinals; keep the ordinal as a back-compat / optimistic-row path only when no durable id is available yet.

On a successful truncating submit against a durable session, the prompt.submit result additionally carries survivor_user_row_ids — the fresh post-rewrite row IDs of the surviving user turns, in visible-user-ordinal order. The rewrite re-inserts the kept prefix as new rows, so every row ID the host cached before the rewind is stale afterward; rebind cached IDs from this list (a null entry means that turn has no durable ID — drop the cached one) or the next rewind targeting an older surviving turn will be refused with 4018.

Events streamed back

message.delta, message.complete, tool.start, tool.generating, tool.complete, gateway.ready, request.cancel, plus session lifecycle and error events.

Server→client requests (questions the agent asks you)

Approvals, clarify questions, sudo/secret prompts, vault unlock, MCP setup and the desktop read/act bridges are JSON-RPC requests from the gateway to the client, not events. The frame carries a string id, and the client answers with a normal JSON-RPC response bearing the same id:

← {"jsonrpc":"2.0","id":"srq-7","method":"approval","params":{"session_id":"…","request_id":"…","command":"rm -rf build","description":"…"}}
→ {"jsonrpc":"2.0","id":"srq-7","result":{"choice":"once"}}

Methods: approval → {choice}; clarify → {answer} (single) or {answers} / {} cancel (batch, with clarify.lock to lock one answer early); sudo, secret, vault.code, vault.unlock_prompt → {value}; connection → {settled_by, targets} (the manage_connections card: one outcome per target); terminal.read, window.read, preview.act, tour → {value} (JSON text). Respond with a JSON-RPC error (-32601) for a method your host does not implement so the agent fails fast instead of waiting out the timeout.

When the gateway withdraws a question (timeout, interrupt, answered from another surface) it emits request.cancel { id, method, reason }; clear only the matching prompt. session.resume / session.activate results and session.events.since carry open_requests — the still-open frames — so a reconnecting client re-renders (and can still answer) them.

Rebuilding the in-flight turn on reconnect

session.resume / session.activate results carry inflight — the turn still running (or the retained failed one) that history does not hold yet: user, assistant streamed so far, streaming, mid-turn corrections, and error fields. When the turn was started by the gateway rather than typed by a person (a background-process completion, an async delegation result, a hidden scaffolding prompt) inflight also carries the same display_kind / display_metadata the persisted messages row will get, so a client renders the live prompt exactly as it will render history after the turn lands — a process_complete timeline marker with display_metadata.display_text, nothing at all for hidden. Both fields are absent for genuine user input; never infer origin from the prompt text (a user quoting a marker string is still a user).

Pi-style RPC mapping

Every command in the Pi-mono RPC spec (issue #360) has a TUI-gateway equivalent:

Pi command Hermes equivalent
prompt prompt.submit (or ACP session/prompt)
steer session.steer
follow_up prompt.submit queued after current turn
abort session.interrupt
set_model command.dispatch for /model <provider:model> (mid-session, persistent)
compact session.compress
get_state session.status
get_messages session.history
switch_session session.resume
fork session.branch
ui_request / ui_response server→client requests clarify / sudo / secret / approval answered by JSON-RPC response frames

OpenAI-Compatible API Server

gateway/platforms/api_server.py exposes hermes over HTTP for any client that already speaks the OpenAI format. Useful when you want a web frontend, a curl-driven CI runner, or a non-Python consumer.

Endpoints:

POST /v1/chat/completions        OpenAI Chat Completions (streaming via SSE)
POST /v1/responses               OpenAI Responses API (stateful)
POST /v1/runs                    Start a run, returns run_id (202)
GET  /v1/runs/{id}               Run status
GET  /v1/runs/{id}/events        SSE stream of lifecycle events
POST /v1/runs/{id}/approval      Resolve a pending approval
POST /v1/runs/{id}/steer         Inject mid-run guidance at the next tool boundary
POST /v1/runs/{id}/stop          Interrupt the run
GET  /v1/capabilities            Machine-readable feature flags
POST /v1/browser-control/register Register a browser controller
GET  /v1/browser-control/ws       Browser-controller WebSocket
GET  /v1/models                  Lists hermes-agent
GET  /api/model/options          Provider-aware picker inventory
GET  /health, /health/detailed

Setup, headers (X-Hermes-Session-Id, X-Hermes-Session-Key), and frontend wiring: API Server.

Browser extensions can opt into the disabled-by-default controller protocol to drive the exact browser session that opened the Hermes conversation. The API and dashboard transports share one principal-bound broker and one explicit capability allowlist; see Browser-extension control.

Model catalog surfaces

The OpenAI-compatible API intentionally keeps GET /v1/models minimal: it is the compatibility endpoint frontends expect, not the full Hermes provider/model picker catalog.

If an external control plane needs Hermes' curated provider rows, per-model pricing, or capability hints, use one of the authenticated picker surfaces:

  • API server REST: GET /api/model/options with the API-server bearer key
  • Dashboard backend REST: GET /api/model/options with X-Hermes-Session-Token
  • TUI gateway RPC: model.options

Those surfaces share the same payload builder and the same custom-provider probe policy:

  • Normal open: probe only the current custom provider so offline saved endpoints do not stall the picker.
  • Explicit refresh (refresh=1 or refresh: true): bust the provider-model cache and probe all saved custom providers so live catalogs repopulate fully.

Use /v1/models for OpenAI-client compatibility. Use /api/model/options or model.options when you are building a Hermes-aware model picker.

POST /v1/runs/{id}/steer is the HTTP equivalent of Hermes /steer: it does not create a new user turn or immediately rewrite the assistant output already in flight. Instead, the text is appended to the live run and becomes visible to the agent after the next tool boundary, so it can course-correct without discarding the current tool-calling loop.

/v1/runs/{id}/steer is only accepted while the run status is running. Queued, approval-paused, stopping, cancelled, failed, and completed runs return 409 run_not_accepting_steer, even if the server still retains internal agent references during cooperative shutdown.

A 200 (and the run.steered event) means the text was queued, not that the agent consumed it. If a steer lands after the agent's final response — with no later tool boundary to deliver it at — the undelivered text is returned as pending_steer on the terminal event (run.completed, run.failed, or run.cancelled) and run status, so the client can replay it as the next user turn instead of losing it.

Terminal run status

The terminal status of a run is derived from how the agent's turn actually ended, and the terminal event name always matches it (run.<status>):

Turn outcome Status Terminal event Flags on the event / status
Final answer produced completed run.completed completed: true
Interrupted (/stop, or an interrupt inside the agent) cancelled run.cancelled completed: false, interrupted: true, turn_exit_reason naming the issuer — interrupted_by_user for a human stop, interrupted_by_system(<issuer>) / interrupted_during_api_call(<issuer>) when a watchdog (e.g. cron_inactivity_watchdog, turn_liveness_watchdog, gateway_inactivity_watchdog, session_turn_lease_lost) ended the turn
Provider/agent failure failed run.failed completed: false, error
Ended without finishing (iteration budget, truncated or partial reply) failed run.failed completed: false, partial when applicable, turn_exit_reason (e.g. max_iterations_reached(60/60)), output with any fallback text

A run is never reported as completed with completed: false or partial: true in the same payload. The same rule applies to /api/sessions/{id}/chat/stream, whose assistant.completed payload carries the real completed / partial / interrupted flags and whose terminal event is run.completed, run.failed, or run.cancelled.


Which one should I use?

  • You're writing an IDE plugin and the IDE already speaks ACP → ACP. Zero protocol work on the IDE side.
  • You're writing a custom desktop / web / TUI host and want every Hermes feature (slash commands, approvals, clarify, multi-agent, session branching) → TUI gateway JSON-RPC.
  • You want any OpenAI-compatible frontend, a language-agnostic HTTP client, or curl-driven automation → API server.
  • You want a Python in-process embed without a subprocess → import run_agent.AIAgent directly. See Agent Loop.

Model hot-swapping

Mid-session model switching works on every surface — it's the /model slash command under the hood.

  • CLI / TUI: /model claude-sonnet-4 or /model openrouter:anthropic/claude-sonnet-4.6
  • TUI gateway RPC: command.dispatch with {"command": "/model claude-sonnet-4"}
  • ACP: the IDE sends the slash command as a prompt; the agent dispatches it
  • API server: include a model field in the request body

Provider-aware resolution (the same model name picks the right format for whatever provider you're on) is built in. See hermes_cli/model_switch.py.


A note on --mode rpc

Hermes does not have a --mode rpc flag. The three protocols above already cover the use cases — ACP for IDE-protocol clients, the TUI gateway for stdio JSON-RPC hosts, and the API server for HTTP. If you find a real gap that none of them fill, open an issue with the concrete consumer you're building.