`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS
loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and
has no tools of its own. Every real request it hears becomes a normal Hermes
turn on the open chat — any model/provider the session selected, full toolset,
memory, approvals — and the voice paraphrases the reply aloud.
Backend
- tools/voice_live.py: mode/credential/persona resolution and the one server-side
step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer,
pinned to client delegation; the OpenAI key never reaches the renderer. Voice
persona follows the vendor prompting guide (role, style, labelled delegation
policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn
model-input note (transcript in, speakable prose out).
- REST: GET /api/audio/voice-live/status (mode + readiness, non-secret),
POST /api/audio/voice-live/session (SDP exchange). The offer is passed
byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF".
- prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the
note rides the model input via the existing _prepend_note seam, the persisted
user row stays the user's words, the system prompt stays byte-stable.
- config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*.
Desktop
- lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript
accumulation, session.commentary/thinking/instructions appends (500-token
chunking), mute, graceful close waiting for session.closed.
- hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation;
delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking
appends; reply streamed back per sentence; spoken stop phrase ends the chat;
a newer delegation interrupts an in-flight turn.
- use-composer-voice mounts both engines and latches one at conversation start
from the backend-resolved status; gpt-live without a key falls back to chained
with a notice. Settings → Voice gets the mode dropdown, voice picker, persona.
Live-verified on the worktree desktop build (headless Electron, CDP, synthetic
mic): "what is 17 times 23 and which model are you on" → delegation → Hermes
(Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through
OpenRouter"; follow-up "double that" resolved from the spoken context → 782;
"run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as
quiet context → spoken kernel version; "stop" closed the session
(reason=close_requested). Chained mode creates no RTCPeerConnection.