Files
ethernet 82a5affdd3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/backup.py
#	tests/hermes_cli/test_gateway_restart_loop.py
#	website/docs/developer-guide/web-search-provider-plugin.md
#	website/docs/getting-started/installation.md
#	website/docs/getting-started/updating.md
#	website/docs/index.mdx
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/docker.md
#	website/docs/user-guide/windows-wsl-quickstart.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/plugins/index.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/web-search-provider-plugin.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/index.mdx
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/cli-commands.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/environment-variables.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/docker.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/features/plugins.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/security.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/windows-wsl-quickstart.md
2026-09-18 18:41:16 -04:00

16 KiB
Raw Permalink Blame History

sidebar_position, title, description
sidebar_position title description
11 Wake Word Hands-free 'Hey Hermes' wake word — start a voice session by speaking, the 'Hey Siri' way

Wake Word ("Hey Hermes")

The wake word turns Hermes into a hands-free assistant across the CLI, TUI, and desktop app: with one setting on, Hermes listens in the background for a spoken trigger phrase. Say it, and Hermes starts a fresh session, opens the microphone, captures your command via the normal voice pipeline, and answers — exactly like "Hey Siri" or "Alexa". Use surface to pick which one listens.

Detection runs entirely on-device. The always-on listener only watches for the wake phrase; no audio leaves your machine until you actually speak a command to the agent.

How it works

  1. With wake_word.enabled: true (or after /wake on), a lightweight hotword detector listens on your configured input device, or the process default microphone when wake_word.input_device is unset.
  2. When it hears the wake phrase it pauses itself (freeing the mic), starts a new session, and records one utterance with voice mode's silence detection.
  3. Your speech is transcribed and sent to the agent. After it replies, the listener resumes automatically and waits for the next wake word.

It is off by default — nothing listens until you turn it on.

On the desktop app, a hands-free voice conversation can be ended by simply saying "stop" (or "never mind", "goodbye", "cancel", "that's all") — the spoken command ends the conversation instead of being sent to the agent. Only a whole-utterance stop command matches, so a real request like "stop the docker container" still goes through normally.

Remote desktop (client capture)

When the desktop app connects to a remote Hermes backend (for example a headless Docker host or a machine in another room), the backend often has no microphone. Server-side PortAudio then fails with “Failed to open the wake-word microphone.”

Hermes supports client capture for that case:

  1. The desktop arms wake with capture: client (automatic for the GUI when the backend has no local input device, or set explicitly below).
  2. The selected wake engine still runs on the backend (same engines, same models).
  3. The desktop opens the local Mac/PC microphone, resamples to 16 kHz mono int16, and streams short frames via the wake.feed RPC.
  4. On detection the backend emits wake.detected as usual; the desktop starts the normal voice pipeline on the client mic.
wake_word:
  enabled: true
  capture: auto    # auto | local | client
  # auto   — local PortAudio unless the desktop arms with client_capture
  # local  — always open the backend mic (CLI/TUI default)
  # client — always expect wake.feed PCM from the desktop (remote-friendly)

The desktop GUI always passes client_capture: true on wake.start, so remote backends without a mic arm in client mode automatically. CLI and TUI keep local capture unless you set capture: client explicitly.

Privacy note: with client capture, wake PCM travels over the authenticated desktop↔backend WebSocket (same channel as the rest of the session). Detection still does not send audio to third-party wake APIs; the engine is local to the backend process.

Engines

Engine Cost API key Notes
openWakeWord Free None TFLite through pyopen-wakeword. Includes the "hey hermes" model. Custom models require a .tflite file. Not available on Intel macOS or native Windows ARM64.
sherpa Free None Open-vocabulary detection for typed phrases. Downloads an English model on first use. Supports native Windows ARM64.
Porcupine Free tier / paid PORCUPINE_ACCESS_KEY Picovoice engine; built-in keywords + custom .ppn files

The default provider is auto. It selects the first platform-supported engine in this order: openWakeWord → sherpa → Porcupine. The platform is that of the Python backend, not a remote desktop client:

  • Native Windows ARM64 and Intel macOS: sherpa (free, no key).
  • Windows x64, Apple Silicon, and supported Linux targets: openWakeWord (free, no key).

An explicit provider stays selected, even if this platform does not support it. Hermes reports the requirement error instead of silently switching engines. Existing explicit settings are not migrated. To opt into automatic selection, run hermes config set wake_word.provider auto. Wake detection stays off until you enable it.

The default phrase label is "hey hermes". For openWakeWord, Hermes includes its trained TFLite model. The pyopen-wakeword package includes the shared feature-extraction models, so this engine does not download models when it starts.

If the selected engine is missing, Hermes requests its PM extra when you enable wake-word detection. security.allow_lazy_installs controls this installation. A new dependency environment can require a Hermes restart before the engine loads. Packaged builds include the engine dependencies supported by their target.

The pyopen-wakeword macOS wheel contains an ARM64-only library despite its universal2 label. Hermes excludes that engine on Intel Macs and native Windows ARM64. Sherpa provides keyless detection on both targets.

Porcupine's default keyword is "jarvis", not "hey hermes". Its phrase setting is only a display label; choose a built-in keyword or supply a custom .ppn model to change what it detects. Get an access key at console.picovoice.ai and store PORCUPINE_ACCESS_KEY in your profile's .env, not config.yaml.

The supported pyopen-wakeword wheels target Apple Silicon with macOS 15 or later, glibc Linux 2.35 or later, and Windows x64. These requirements apply to that engine, not every Hermes feature. Termux's core/ACP package does not include this wake stack.

Quick start

# In an interactive `hermes` session:
/wake on        # start listening (installs the engine on first use)
/wake status    # show phrase, provider, and state
/wake off       # stop listening

In the desktop app, hover the microphone in the composer and click the ear that fans out of it. The ear is solid when the wake word is listening.

The toggle IS the setting: turning the wake word on or off — via /wake or the desktop ear button — also writes wake_word.enabled to ~/.hermes/config.yaml, so your choice persists across sessions. You can also flip it by hand:

wake_word:
  enabled: true

Configuration

wake_word:
  enabled: false
  surface: auto               # eligible surface: "auto" | "cli" | "tui" | "gui"
  input_device: null           # PortAudio input index or device-name substring; null = process default
  capture: auto               # auto | local | client — where PCM is captured (see Remote desktop)
  provider: auto              # auto | openwakeword | sherpa | porcupine (requires an access key)
  phrase: "hey hermes"        # cosmetic label only — detection is keyed by the model/keyword below
  sensitivity: 0.6            # 0.0-1.0 — higher = stricter (fewer false triggers), consistent across all engines
  confirmation_frames: 3      # openWakeWord only — consecutive over-threshold frames required to fire
  start_new_session: true     # start a fresh session on wake vs. continue the current one
  openwakeword:
    model: hey_hermes         # bundled default, or an absolute path to a custom .tflite
  porcupine:
    keyword: jarvis           # built-in keyword OR path to a custom .ppn

sensitivity and start_new_session apply to all three engines. For sherpa, phrase selects the detection phrase. For openWakeWord and Porcupine, phrase is a display label; their model or keyword selects the detection phrase.

input_device is passed directly to the wake listener's PortAudio (sounddevice) stream. Use either a numeric device index or an unambiguous device-name substring. This setting only changes wake-word capture; desktop push-to-talk still uses the desktop application's microphone path.

Reducing false triggers on ambient speech

openWakeWord scores one short (~80ms) audio frame at a time, so a stray phoneme in background conversation can occasionally spike a single frame over the threshold and fire the wake word unintentionally. Two knobs control this:

  • confirmation_frames (default 3, openWakeWord only) — how many consecutive over-threshold frames are required before the wake fires. A real "hey hermes" holds a high score across several frames; an ambient blip spikes just one. Raise it (e.g. 4–5) if you still get false triggers in a noisy room; the cost is a few tens of milliseconds of extra latency. 1 restores the old fire-on-first-frame behavior.
  • sensitivity (default 0.6) — the detection threshold, 0.0–1.0. Higher is stricter (fewer false triggers). This direction is consistent across all engines — for openWakeWord it's the raw per-frame score threshold, for sherpa it maps onto the keyword threshold, and for Porcupine it's inverted internally so "higher = stricter" holds there too. The 0.6 default sits above openWakeWord's permissive 0.5 baseline, which let near-misses like "hey hor" through; raise toward 0.8 if you still get false fires, lower it if real "hey hermes" utterances are missed.

The sherpa and porcupine engines decode the whole phrase internally, so they don't have the single-frame-spike problem and ignore confirmation_frames (but they still honor sensitivity).

The openwakeword provider name now selects pyopen-wakeword. Its wheel includes the TFLite library and shared feature models. Hermes uses the bundled hey_hermes.tflite model by default. ONNX wake models and the inference_framework setting are no longer supported.

Surfaces (CLI, TUI, GUI)

The wake word works in all three Hermes surfaces, and surface picks which one owns the listener and opens the new session when it fires:

surface Behavior
auto (default) All local surfaces are eligible; the first one to arm owns the listener.
cli Only the classic hermes CLI.
tui Only hermes --tui.
gui Only the desktop app.

The detector is on-device and single-mic, so only one surface listens at a time, including when Hermes surfaces run in separate processes. Ownership is sticky: the first eligible claimant keeps the listener until it stops, disconnects, or its process exits. Hermes does not silently fail over to another open surface. Set surface when you want to pin ownership instead of using first-claim wins. The TUI and desktop GUI share the same Python backend (tui_gateway), which runs the detector server-side and yields the mic to voice capture while a command records.

Using a different phrase

"Hey Hermes" is the default detection phrase with openWakeWord and sherpa. Porcupine uses its configured keyword instead ("jarvis" by default). To wake on something else, the easiest path on supported platforms is the open-vocabulary engine:

Option A — sherpa (any phrase, zero training)

Type the phrase you want; it's tokenized at runtime — "hey coder", "computer", "wake up neo", anything:

wake_word:
  enabled: true
  provider: sherpa
  phrase: "hey coder"        # detection key — just type your phrase

The small English KWS model (~13 MB) downloads once on first use. Each profile can set its own phrase — "hey <profile>" for every profile you run.

Waking a specific profile (desktop)

With the sherpa engine, ONE listener can wake ANY profile. Every profile whose config has wake_word.enabled: true is enrolled automatically; its phrase defaults to hey <profile name> when unset. Say a profile's phrase and the desktop app live-switches to that profile, opens a fresh session there, and starts hands-free voice:

  • "hey hermes" → default profile
  • "hey coder" → the coder profile
  • "hey trader" → the trader profile

Set wake_word.profile_routing: false on the listener's profile to opt out and listen only for its own phrase. The CLI and TUI are single-profile processes: a wake phrase belonging to another profile prints the switch command (hermes -p <profile>) instead of routing.

Names are matched acoustically by their English subword sounds: two-word phrases with distinct, 2+ syllable names work best. Very short names, heavy non-English phonology, or two profiles with similar-sounding names will degrade accuracy — tune per-profile sensitivity if needed.

Option B — openWakeWord (free, trained model)

For a different phrase, obtain or train a compatible openWakeWord TFLite model. Set its absolute path in the configuration. Hermes does not resolve built-in names such as hey_jarvis or download their models for you.

wake_word:
  enabled: true
  provider: openwakeword
  phrase: "computer"
  openwakeword:
    model: /absolute/path/to/computer.tflite

Training references:

:::tip Pick a distinctive phrase Wake phrases that don't collide with everyday speech generalize best. Two syllables with an uncommon word ("hermes" qualifies) beat common words like "hello" or "stop". :::

Option C — Porcupine (custom keyword in seconds)

Create a "Hey Hermes" keyword in the Picovoice Console, download the .ppn, and:

wake_word:
  enabled: true
  provider: porcupine
  phrase: "hey hermes"
  porcupine:
    keyword: ~/.hermes/wakewords/hey_hermes.ppn

Set your access key in ~/.hermes/.env:

PORCUPINE_ACCESS_KEY=your-key-here

Requirements

  • A working microphone and the sounddevice + numpy audio stack (shared with voice mode).
  • An STT provider for transcribing the spoken command — local faster-whisper works out of the box; see Voice Mode for the full provider list.
  • A TTS provider for speaking the reply (the default edge-tts works with no key). The wake flow is fully hands-free, so the toggle refuses to arm until both STT and TTS are ready — hermes tools (Voice section) sets them up.
  • The wake engine deps (auto-installed, or hermes-agent[wake]).

/wake status reports exactly what's missing if the listener won't start.

"Listening" but never wakes (macOS)

macOS grants microphone access per process. STT working in the desktop app proves the renderer has mic access — the wake listener runs in the Python backend, which needs its own grant. Without it, CoreAudio hands the backend a "working" stream that only ever delivers silence, so the ear shows listening but the phrase never fires. Hermes detects this (/wake status shows "mic delivers only silence"; the desktop's folded voice menu carries the same hint on its trigger). Fix: System Settings → Privacy & Security → Microphone → enable the Hermes backend (it may appear as your terminal, python, or Hermes), then toggle the wake word off and on.

"Listening" but receives silence (Windows)

Desktop push-to-talk and wake-word capture use different microphone paths. Push-to-talk uses the desktop application's browser capture, while the wake-word listener opens a PortAudio stream in the Python backend. One can work while the other selects a silent or unusable Windows input.

/wake status reports the selected input device and Windows audio host API. When it reports silence, set wake_word.input_device to the numeric index or an unambiguous name of the working PortAudio input, then toggle the wake word:

hermes config set wake_word.input_device "Microphone Array"

Use null to return to the process default:

hermes config set wake_word.input_device null

Notes & limits

  • Local surfaces only. The wake word runs in the CLI, TUI, and desktop GUI — wherever a local microphone is available. It does not run in the messaging gateway (Telegram, Discord, …), which has no mic.
  • One mic at a time. The detector releases the microphone while a command is recording and reclaims it once the turn ends, so it won't fight voice capture.
  • Privacy. Hotword detection is local. Set sensitivity higher if you get false triggers, lower if it misses you.