Files
hermes-agent/tools
Teknium f923faa0b8 feat(voice): GPT-Live voice chat mode — a full-duplex voice frontend that delegates to Hermes (Desktop)
`voice.voice_chat_mode: gpt-live` swaps the desktop's chained STT → turn → TTS
loop for OpenAI's gpt-live-1: one voice model that listens while it speaks and
has no tools of its own. Every real request it hears becomes a normal Hermes
turn on the open chat — any model/provider the session selected, full toolset,
memory, approvals — and the voice paraphrases the reply aloud.

Backend
- tools/voice_live.py: mode/credential/persona resolution and the one server-side
  step the API needs — POST /v1/live/sessions exchanging the renderer's SDP offer,
  pinned to client delegation; the OpenAI key never reaches the renderer. Voice
  persona follows the vendor prompting guide (role, style, labelled delegation
  policy describing Hermes as the backend). VOICE_LIVE_TURN_NOTE is the per-turn
  model-input note (transcript in, speakable prose out).
- REST: GET /api/audio/voice-live/status (mode + readiness, non-secret),
  POST /api/audio/voice-live/session (SDP exchange). The offer is passed
  byte-exact: a stripped trailing CRLF is a vendor 400 "unmarshal SDP: EOF".
- prompt.submit accepts surface=voice-live (+ voice_context) beside hud; the
  note rides the model input via the existing _prepend_note seam, the persisted
  user row stays the user's words, the system prompt stays byte-stable.
- config_defaults: voice.voice_chat_mode (chained|gpt-live), voice.gpt_live.*.

Desktop
- lib/voice-live.ts: RTCPeerConnection + oai-events data channel owner, transcript
  accumulation, session.commentary/thinking/instructions appends (500-token
  chunking), mute, graceful close waiting for session.closed.
- hooks/use-voice-live-conversation.ts: same public shape as useVoiceConversation;
  delegation → prompt.submit(surface=voice-live); tool activity → quiet thinking
  appends; reply streamed back per sentence; spoken stop phrase ends the chat;
  a newer delegation interrupts an in-flight turn.
- use-composer-voice mounts both engines and latches one at conversation start
  from the backend-resolved status; gpt-live without a key falls back to chained
  with a notice. Settings → Voice gets the mode dropdown, voice picker, persona.

Live-verified on the worktree desktop build (headless Electron, CDP, synthetic
mic): "what is 17 times 23 and which model are you on" → delegation → Hermes
(Claude Sonnet 4.5 via OpenRouter) → spoken "391 … Claude Sonnet 4.5 through
OpenRouter"; follow-up "double that" resolved from the spoken context → 782;
"run uname -r" ran the terminal tool with "Hermes is working: terminal" fed as
quiet context → spoken kernel version; "stop" closed the session
(reason=close_requested). Chained mode creates no RTCPeerConnection.
2026-09-11 19:14:24 -07:00
..