`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and has no tools of its own. Every real request it hears becomes a normal Hermes turn on the open chat — any model/provider the session selected, full toolset, memory, approvals — and the voice paraphrases the reply aloud. Backend - tools/voice_live.py: mode/credential/persona resolution and the one server-side step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer, pinned to client delegation; the OpenAI key never reaches the renderer. Voice persona follows the vendor prompting guide (role, style, labelled delegation policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn model-input note (transcript in, speakable prose out). - REST: GET /api/audio/voice-live/status (mode + readiness, non-secret), POST /api/audio/voice-live/session (SDP exchange). The offer is passed byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF". - prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the note rides the model input via the existing _prepend_note seam, the persisted user row stays the user's words, the system prompt stays byte-stable. - config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*. Desktop - lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript accumulation, session.commentary/thinking/instructions appends (500-token chunking), mute, graceful close waiting for session.closed. - hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation; delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking appends; reply streamed back per sentence; spoken stop phrase ends the chat; a newer delegation interrupts an in-flight turn. - use-composer-voice mounts both engines and latches one at conversation start from the backend-resolved status; gpt-live without a key falls back to chained with a notice. Settings → Voice gets the mode dropdown, voice picker, persona. Live-verified on the worktree desktop build (headless Electron, CDP, synthetic mic): "what is 17 times 23 and which model are you on" → delegation → Hermes (Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through OpenRouter"; follow-up "double that" resolved from the spoken context → 782; "run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as quiet context → spoken kernel version; "stop" closed the session (reason=close_requested). Chained mode creates no RTCPeerConnection.
Website
This website is built using Docusaurus, a modern static website generator.
Installation
yarn
Local Development
yarn start
This command starts a local development server and opens up a browser window. Most changes are reflected live without having to restart the server.
Build
yarn build
This command generates static content into the build directory and can be served using any static contents hosting service.
Deployment
Using SSH:
USE_SSH=true yarn deploy
Not using SSH:
GIT_USER=<Your GitHub username> yarn deploy
If you are using GitHub pages for hosting, this command is a convenient way to build the website and push to the gh-pages branch.
Diagram Linting
CI runs ascii-guard to lint docs for ASCII box diagrams. Use Mermaid (````mermaid`) or plain lists/tables instead of ASCII boxes to avoid CI failures.