Lowest-hop voice path in both directions for desktop + remote gateway:
mic audio goes straight to the profile's STT provider and reply text is
synthesized on the desktop with the profile's TTS provider. The
desktop-gateway link carries only text (which the chat stream carries
anyway). No second key store: GET /api/audio/voice-config returns the
profile's resolved provider/model/language/key using the exact resolution
chains transcription_tools/tts_tool use, over the authenticated REST
channel. Keys live in renderer memory only.
Backend:
- tools/voice_client_config.py: single resolver; per-provider client
wire shapes (openai-multipart, xai-stt, elevenlabs-stt, openai-speech,
elevenlabs-tts). Server-host-only providers (local whisper, edge,
command/plugin) and missing credentials resolve to {mode: relay}.
xAI OAuth stays relay (bearer refreshes server-side).
- web_server.py: GET /api/audio/voice-config, profile-scoped via the
same _config_profile_scope seam as /api/audio/transcribe.
- config_defaults.py: voice.client_direct gate (default true).
Desktop:
- lib/voice-client-direct.ts: config fetch keyed by (connection,
profile) with 60s TTL, provider-direct STT + TTS calls, sentence
cutter mirroring the server pipeline's contract.
- Dictation (use-prompt-actions + session-tile) tries client-direct
first; null -> existing relay unchanged; provider rejections surface.
- voice-playback.ts: client-direct speech session as the top rung of
startSpeechStream/playSpeechText; WS relay + POST fallback unchanged
below it. Barge-in via the same stopVoicePlayback sequence bump.
Validation: 13/13 backend E2E (real temp HERMES_HOME + real resolution),
live FastAPI TestClient E2E (direct + gate-flip), 15/15 client tests
(wire shapes, scope-keyed caching, rejection surfacing, sentence cutter),
sibling suites 72/72 + 36/36, tsc + eslint + ruff clean.
Docs: voice-mode.md client-direct section ships in this PR.