Files
hermes-agent/tests/tools/test_voice_cli_integration.py
Teknium 9bcbe7b5df feat(i18n): pluggable, layered language packs across core, Desktop and TUI (#126296)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales

* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter

ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.

* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs

* chore(tui): split en catalog siblings by lane (slash sibling)

* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports

* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts

* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list

* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes

* feat(plugins): report language-pack layers in the mid-run activation summary

* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n

StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.

* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter

- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
  the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
  layers partial packs (nested or flat dotted) over bundled/en via
  mergeTranslations; a string over a function-valued en entry becomes a
  positional {0}/{1} formatter; $appLocaleVersion bumps so translators
  re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
  registry as source 'backend' (method-not-found is silent); re-synced on
  socket open, display.language change and profile switch. A saved pack-only
  language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
  tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
  from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).

* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)

Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.

* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()

- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
  ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
  accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
  properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces

* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)

* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()

Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.

locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).

* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()

- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions

* i18n(platforms): route Google Chat and Teams user-facing text through t()

Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.

Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.

* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()

LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.

* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)

Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.

* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)

_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.

* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)

RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.

* i18n(cli): live-work dock, subagent monitor and render copy through t()

cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.

* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()

* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)

get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.

* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output

* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time

Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).

* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)

cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.

* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()

- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
  at call time; category labels via slash.category.*; help/alias/usage suffixes
  via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
  bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
  lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
  status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
  approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
  bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
  heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
  sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
  keeps its English constant; callers use history_unreadable() ->
  gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.

* i18n(telegram): route adapter chat copy through t()

Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.

Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.

* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()

- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor

* i18n(discord): route adapter chat copy through t()

Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().

* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)

* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()

- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
  existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
  via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
  _DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
  Updating/Generating) are one full template per variant; plurals use
  <key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
  translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).

* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English

* test(i18n): pin Telegram/Discord adapter catalog wiring

Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.

* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge

* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint

* i18n(tr): translate bundled catalog + tui pack

* i18n(ja): translate bundled catalog + tui pack

* i18n(ko): translate bundled catalog + tui pack

* i18n(zh): translate bundled catalog + tui pack

* i18n(fr): translate bundled catalog + tui pack

* i18n(af): translate bundled catalog + tui pack

* i18n(uk): translate bundled catalog + tui pack

* i18n(ar): translate bundled catalog + tui pack

* i18n(pt): translate bundled catalog + tui pack

* i18n(it): translate bundled catalog + tui pack

* i18n(es): translate bundled catalog + tui pack

* i18n(zh-hant): translate bundled catalog + tui pack

* i18n(ru): translate bundled catalog + tui pack

* i18n(hu): translate bundled catalog + tui pack

* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)

* i18n(de): translate bundled catalog + tui pack

* i18n(ga): translate bundled catalog + tui pack

* test(i18n): fixture matches _normalize_lang(lang, home) signature

* i18n(tui): scaffold userMessages/slashCmd en siblings

* i18n(tui): wire secure prompts + content tables

* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays

Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.

* i18n(tui): wire slash ops/wake replies

* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)

* i18n(tui): wire slash core/debug/setup replies

* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)

* i18n(tui): wire slash session/topup/subscription replies

* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)

* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys

* i18n(tui): wire userMessages copy through the userMessages namespace

* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json

* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)

Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).

* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)

* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)

* tui: i18n-export-en script (English templates for pack translators)

* docs(i18n): bundled TUI packs are the bottom layer of the tui surface

* i18n(ru): translate TUI pack

* i18n(ar): translate TUI pack

* i18n(es): translate TUI pack

* i18n(pt): translate TUI pack

* i18n(ko): translate TUI pack

* i18n(de): translate TUI pack

* i18n(ja): translate TUI pack

* i18n(fr): translate TUI pack

* i18n(tr): translate TUI pack

* i18n(it): translate TUI pack

* i18n(zh): translate TUI pack

* i18n(zh-hant): translate TUI pack

* i18n(hu): translate TUI pack

* i18n(uk): translate TUI pack

1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.

Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).

* i18n(ga): translate TUI pack

* i18n(af): translate TUI pack

* plugin_guard: locale catalogs in language packs step down the agent-config family

A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.

* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)

* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header

Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.

* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel

- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
  and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
  agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)

* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})

* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property

* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal

* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export

The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.

* test: unbreak two main-red timing tests the PR merge-ref inherits

- test_local_runtime racing fake publishes the modern state record (legacy pid-only
  records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
  a fixed 0.35s, which a loaded CI runner does not always meet

* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)

* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak

SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.

* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings

* chore(i18n): regenerate desktop key catalog after main sync

---------

Co-authored-by: Teknium <teknium@nousresearch.com>
2026-09-28 14:16:18 -07:00

686 lines
26 KiB
Python

"""Tests for CLI voice mode integration -- markdown stripping, voice state
management, TTS/STT wiring, barge-in and the full-duplex listener."""
import json
import queue
import threading
from types import SimpleNamespace
from unittest.mock import MagicMock, patch
import pytest
from agent.i18n import t
def _make_voice_cli(**overrides):
"""Create a minimal HermesCLI with only voice-related attrs initialized.
Uses ``__new__()`` to bypass ``__init__`` so no config/env/API setup is
needed. Only the voice state attributes (from __init__ lines 3749-3758)
are populated.
"""
from cli import HermesCLI
cli = HermesCLI.__new__(HermesCLI)
cli._voice_lock = threading.Lock()
cli._voice_mode = False
cli._voice_tts = False
cli._voice_recorder = None
cli._voice_recording = False
cli._voice_processing = False
cli._voice_continuous = False
cli._voice_tts_done = threading.Event()
cli._voice_tts_done.set()
cli._voice_tts_stop = None
cli._voice_barge_capture = threading.Event()
cli._voice_last_tts_text = ""
cli._voice_barge_phase = None
cli._pending_input = queue.Queue()
cli._app = None
cli._attached_images = []
cli.console = SimpleNamespace(width=80)
for k, v in overrides.items():
setattr(cli, k, v)
return cli
# ============================================================================
# Markdown stripping — import real function from tts_tool
# ============================================================================
from tools.tts_text_normalize import _strip_markdown_for_tts
class TestMarkdownStripping:
def test_empty_after_stripping_returns_empty(self):
text = "```python\nprint('hello')\n```"
result = _strip_markdown_for_tts(text)
assert result == ""
def test_complex_response(self):
text = (
"## Answer\n\n"
"Here's how to do it:\n\n"
"```python\ndef hello():\n print('hi')\n```\n\n"
"Run it with `python main.py`. "
"See [docs](https://example.com) for more.\n\n"
"- Step one\n- Step two\n\n"
"---\n\n"
"**Good luck!**"
)
result = _strip_markdown_for_tts(text)
assert "```" not in result
assert "https://" not in result
assert "**" not in result
assert "---" not in result
assert "Answer" in result
assert "Good luck!" in result
assert "docs" in result
# ============================================================================
# Real behavior tests — CLI voice methods via _make_voice_cli()
# ============================================================================
class TestHandleVoiceCommandReal:
"""Tests _handle_voice_command routing with real CLI instance."""
def _cli(self):
cli = _make_voice_cli()
cli._enable_voice_mode = MagicMock()
cli._disable_voice_mode = MagicMock()
cli._toggle_voice_tts = MagicMock()
cli._show_voice_status = MagicMock()
return cli
@patch("cli._cprint")
def test_on_calls_enable(self, _cp):
cli = self._cli()
cli._handle_voice_command("/voice on")
cli._enable_voice_mode.assert_called_once()
@patch("cli._cprint")
def test_unknown_subcommand(self, mock_cp):
cli = self._cli()
cli._handle_voice_command("/voice foobar")
cli._enable_voice_mode.assert_not_called()
cli._disable_voice_mode.assert_not_called()
# Should print usage via _cprint
assert any("Unknown" in str(c) or "unknown" in str(c)
for c in mock_cp.call_args_list)
class TestEnableVoiceModeReal:
"""Tests _enable_voice_mode with real CLI instance."""
@patch("cli._cprint")
@patch("hermes_cli.config.load_config", return_value={"voice": {}})
@patch("tools.voice_mode.check_voice_requirements",
return_value={"available": True, "details": "OK"})
@patch("tools.voice_mode.detect_audio_environment",
return_value={"available": True, "warnings": []})
def test_success_sets_voice_mode(self, _env, _req, _cfg, _cp):
cli = _make_voice_cli()
cli._enable_voice_mode()
assert cli._voice_mode is True
@patch("cli._cprint")
@patch("hermes_cli.config.load_config", side_effect=Exception("broken config"))
@patch("tools.voice_mode.check_voice_requirements",
return_value={"available": True, "details": "OK"})
@patch("tools.voice_mode.detect_audio_environment",
return_value={"available": True, "warnings": []})
def test_config_exception_still_enables(self, _env, _req, _cfg, _cp):
cli = _make_voice_cli()
cli._enable_voice_mode()
assert cli._voice_mode is True
class TestVoiceBeepConfigReal:
"""Tests the CLI voice beep toggle."""
@patch("hermes_cli.config.load_config", return_value={"voice": {"beep_enabled": False}})
def test_beeps_can_be_disabled(self, _cfg):
cli = _make_voice_cli()
assert cli._voice_beeps_enabled() is False
@patch("cli._cprint")
@patch("cli.threading.Thread")
@patch("tools.voice_mode.play_beep")
@patch("tools.voice_mode.create_audio_recorder")
@patch(
"tools.voice_mode.check_voice_requirements",
return_value={
"available": True,
"audio_available": True,
"stt_available": True,
"details": "OK",
"missing_packages": [],
},
)
@patch(
"hermes_cli.config.load_config",
return_value={
"voice": {
"beep_enabled": False,
"silence_threshold": 200,
"silence_duration": 3.0,
}
},
)
def test_start_recording_skips_beep_when_disabled(
self, _cfg, _req, mock_create, mock_beep, mock_thread, _cp
):
recorder = MagicMock()
recorder.supports_silence_autostop = True
mock_create.return_value = recorder
mock_thread.return_value = MagicMock(start=MagicMock())
cli = _make_voice_cli()
cli._voice_start_recording()
recorder.start.assert_called_once()
mock_beep.assert_not_called()
class TestMaxRecordingSecondsConfigReal:
"""voice.max_recording_seconds must reach the recorder from config.
Regression for the dead-config fix: the predicate alone can stay green
while the CLI wiring regresses, so pin the actual assignment made by
``_voice_start_recording`` for the valid / disabled / corrupted cases.
"""
def _start_with_voice_cfg(self, voice_cfg):
with patch("cli._cprint"), \
patch("cli.threading.Thread", return_value=MagicMock(start=MagicMock())), \
patch("tools.voice_mode.play_beep"), \
patch("tools.voice_mode.create_audio_recorder") as mock_create, \
patch(
"tools.voice_mode.check_voice_requirements",
return_value={
"available": True,
"audio_available": True,
"stt_available": True,
"details": "OK",
"missing_packages": [],
},
), \
patch("hermes_cli.config.load_config", return_value={"voice": voice_cfg}):
recorder = MagicMock()
recorder.supports_silence_autostop = True
mock_create.return_value = recorder
cli = _make_voice_cli()
cli._voice_start_recording()
return recorder
def test_configured_cap_reaches_recorder(self):
recorder = self._start_with_voice_cfg({"max_recording_seconds": 45})
assert recorder._max_recording_seconds == 45
def test_bool_falls_back_to_documented_default(self):
# bool is a subclass of int — ``max_recording_seconds: true`` must not
# become a 1-second cap; it falls back to the documented 120 default,
# mirroring the silence-param corruption handling.
recorder = self._start_with_voice_cfg({"max_recording_seconds": True})
from hermes_cli.config import DEFAULT_CONFIG
assert recorder._max_recording_seconds == DEFAULT_CONFIG["voice"]["max_recording_seconds"]
class TestDisableVoiceModeReal:
"""Tests _disable_voice_mode with real CLI instance."""
@patch("cli._cprint")
@patch("tools.voice_mode.stop_playback")
def test_all_flags_reset(self, _sp, _cp):
cli = _make_voice_cli(_voice_mode=True, _voice_tts=True,
_voice_continuous=True)
cli._disable_voice_mode()
assert cli._voice_mode is False
assert cli._voice_tts is False
assert cli._voice_continuous is False
@patch("cli._cprint")
@patch("tools.voice_mode.stop_playback", side_effect=RuntimeError("boom"))
def test_stop_playback_exception_swallowed(self, _sp, _cp):
cli = _make_voice_cli(_voice_mode=True)
cli._disable_voice_mode()
assert cli._voice_mode is False
class TestVoiceSpeakResponseReal:
"""Tests _voice_speak_response with real CLI instance."""
def test_async_scheduling_clears_done_before_thread_start(self):
cli = _make_voice_cli(_voice_tts=True)
starts = []
class FakeThread:
def __init__(self, target=None, args=(), daemon=None):
self.target = target
self.args = args
self.daemon = daemon
def start(self):
starts.append(cli._voice_tts_done.is_set())
with patch("cli.threading.Thread", FakeThread):
cli._voice_speak_response_async("Hello")
assert starts == [False]
assert not cli._voice_tts_done.is_set()
@patch("cli._cprint")
def test_early_return_when_tts_off(self, _cp):
cli = _make_voice_cli(_voice_tts=False)
with patch("tools.tts_tool.text_to_speech_tool") as mock_tts:
cli._voice_speak_response("Hello")
mock_tts.assert_not_called()
@patch("cli._cprint")
@patch("cli.os.unlink")
@patch("cli.os.path.getsize", return_value=1000)
@patch("cli.os.path.isfile", return_value=True)
@patch("cli.os.makedirs")
@patch("tools.voice_mode.play_audio_file")
@patch("tools.tts_tool.text_to_speech_tool")
def test_play_audio_uses_returned_file_paths(
self, mock_tts, mock_play, _mkd, _isf, _gsz, _unl, _cp
):
def fake_tts(**kwargs):
mp3_path = kwargs["output_path"]
ogg_path = mp3_path.rsplit(".", 1)[0] + ".ogg"
# The tool result is authoritative — file_paths drives playback
return json.dumps({
"success": True,
"file_path": ogg_path,
"file_paths": [ogg_path],
})
mock_tts.side_effect = fake_tts
cli = _make_voice_cli(_voice_tts=True)
cli._voice_speak_response("Hello world")
# Should play the returned OGG path, not the requested MP3 path
mock_play.assert_called_once_with(
mock_tts.call_args.kwargs["output_path"].rsplit(".", 1)[0] + ".ogg"
)
class TestVoiceStopAndTranscribeReal:
"""Tests _voice_stop_and_transcribe with real CLI instance."""
@patch("cli._cprint")
def test_guard_not_recording(self, _cp):
cli = _make_voice_cli(_voice_recording=False)
with patch("tools.voice_mode.transcribe_recording") as mock_tr:
cli._voice_stop_and_transcribe()
mock_tr.assert_not_called()
@patch("cli._cprint")
@patch("tools.voice_mode.play_beep")
def test_no_speech_detected(self, _beep, _cp):
recorder = MagicMock()
recorder.stop.return_value = None
cli = _make_voice_cli(_voice_recording=True, _voice_recorder=recorder)
cli._voice_stop_and_transcribe()
assert cli._pending_input.empty()
@patch("cli._cprint")
@patch("cli.os.unlink")
@patch("cli.os.path.isfile", return_value=True)
@patch("hermes_cli.config.load_config", return_value={"stt": {}})
@patch("tools.voice_mode.transcribe_recording",
return_value={"success": True, "transcript": "hello world"})
@patch("tools.voice_mode.play_beep")
def test_successful_transcription_queues_input(
self, _beep, _tr, _cfg, _isf, _unl, _cp
):
recorder = MagicMock()
recorder.stop.return_value = "/tmp/test.wav"
cli = _make_voice_cli(_voice_recording=True, _voice_recorder=recorder)
cli._voice_stop_and_transcribe()
queued = cli._pending_input.get_nowait()
# Voice transcripts are wrapped in the _VoiceInputMessage sentinel so
# only genuine STT output gets the voice prefix (#65827).
from cli import _VoiceInputMessage
assert isinstance(queued, _VoiceInputMessage)
assert str(queued) == "hello world"
def test_non_local_stt_keeps_generic_transcribing_status(self):
recorder = MagicMock()
recorder.stop.return_value = "/tmp/test.wav"
cli = _make_voice_cli(_voice_recording=True, _voice_recorder=recorder)
with patch("cli._cprint") as mock_print, \
patch("cli.os.path.isfile", return_value=False), \
patch(
"hermes_cli.config.load_config",
return_value={"stt": {"provider": "openai", "model": "whisper-1"}},
), \
patch("tools.voice_mode.transcribe_recording",
return_value={"success": True, "transcript": "hello"}) as mock_transcribe, \
patch("tools.voice_mode.play_beep"):
cli._voice_stop_and_transcribe()
messages = [call.args[0] for call in mock_print.call_args_list]
assert any(t("cli.voice.transcribing") in message for message in messages)
assert all("Hugging Face" not in message for message in messages)
mock_transcribe.assert_called_once_with("/tmp/test.wav", model="whisper-1")
# ---------------------------------------------------------------------------
# Barge-in capture — the interruption is transcribed and queued directly
# ---------------------------------------------------------------------------
class TestVoiceBargeCaptureSubmit:
"""_voice_submit_barge_utterance: the barge monitor's captured WAV becomes
the next turn without a re-record round trip."""
def test_transcript_is_queued_and_wav_removed(self, tmp_path, monkeypatch):
cli = _make_voice_cli()
cli._voice_barge_capture.set()
wav = tmp_path / "barge.wav"
wav.write_bytes(b"RIFF")
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {"success": True, "transcript": "stop, do it differently"},
)
cli._voice_submit_barge_utterance(str(wav))
queued = cli._pending_input.get_nowait()
from cli import _VoiceInputMessage
assert isinstance(queued, _VoiceInputMessage)
assert str(queued) == "stop, do it differently"
assert not cli._voice_barge_capture.is_set()
assert not wav.exists()
def test_no_speech_hands_mic_back_without_queueing(self, tmp_path, monkeypatch):
cli = _make_voice_cli(_voice_mode=True, _voice_continuous=True)
cli._voice_barge_capture.set()
wav = tmp_path / "barge.wav"
wav.write_bytes(b"RIFF")
restarted = threading.Event()
cli._voice_start_recording = lambda: restarted.set()
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {"success": True, "transcript": "", "no_speech": True},
)
cli._voice_submit_barge_utterance(str(wav))
assert cli._pending_input.empty()
assert not cli._voice_barge_capture.is_set()
assert restarted.wait(2.0) # continuous mode resumes listening
def test_playback_phase_echo_of_own_tts_is_dropped(self, tmp_path, monkeypatch):
"""#75780: a playback-phase capture that closely matches the TTS
text Hermes just spoke is speaker bleed, not real user speech --
it must be dropped instead of queued as the next turn, and the mic
handed back so continuous mode keeps listening."""
cli = _make_voice_cli(_voice_mode=True, _voice_continuous=True)
cli._voice_barge_capture.set()
cli._voice_barge_phase = "playback"
cli._voice_last_tts_text = "네, 방금도 제 답변이 그대로 다시 입력됐어요."
wav = tmp_path / "barge.wav"
wav.write_bytes(b"RIFF")
restarted = threading.Event()
cli._voice_start_recording = lambda: restarted.set()
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {
"success": True,
"transcript": "네 방금 네 방금도 제 답변이 그대로 다시 입력됐어요.",
},
)
cli._voice_submit_barge_utterance(str(wav))
assert cli._pending_input.empty() # not queued as a user turn
assert not cli._voice_barge_capture.is_set()
assert restarted.wait(2.0) # mic handed back instead of self-triggering another turn
def test_playback_phase_genuine_interjection_is_still_queued(self, tmp_path, monkeypatch):
"""A real user interjection during playback -- unrelated to the TTS
text -- must still reach the agent."""
cli = _make_voice_cli()
cli._voice_barge_capture.set()
cli._voice_barge_phase = "playback"
cli._voice_last_tts_text = "The weather today is sunny with a light breeze."
wav = tmp_path / "barge.wav"
wav.write_bytes(b"RIFF")
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {
"success": True,
"transcript": "actually can you check my calendar for tomorrow",
},
)
cli._voice_submit_barge_utterance(str(wav))
queued = cli._pending_input.get_nowait()
assert str(queued) == "actually can you check my calendar for tomorrow"
def test_generation_phase_transcript_not_echo_checked(self, tmp_path, monkeypatch):
"""Generation-phase barges (no TTS playing) are never treated as
echo, even if the transcript happens to match old TTS text."""
cli = _make_voice_cli()
cli._voice_barge_capture.set()
cli._voice_barge_phase = "generation"
cli._voice_last_tts_text = "stop, do it differently"
wav = tmp_path / "barge.wav"
wav.write_bytes(b"RIFF")
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {"success": True, "transcript": "stop, do it differently"},
)
cli._voice_submit_barge_utterance(str(wav))
queued = cli._pending_input.get_nowait()
assert str(queued) == "stop, do it differently"
# ============================================================================
# Full-duplex agent-turn listener — CLI phase behaviour
# ============================================================================
class TestVoiceFullDuplexListener:
"""_voice_full_duplex_listener: one mic for the whole turn. Generation-
phase speech interrupts the in-flight agent turn; playback-phase speech
cuts TTS; the capture is submitted either way."""
def _cli(self, monkeypatch, *, listen, voice_cfg=None, **overrides):
cli = _make_voice_cli(
_voice_mode=True, _voice_continuous=True, **overrides
)
cli.agent = None
monkeypatch.setattr(
"hermes_cli.config.load_config",
lambda: {"voice": dict(voice_cfg or {"barge_in": True})},
)
monkeypatch.setattr("tools.voice_mode.full_duplex_listen", listen)
monkeypatch.setattr("tools.voice_mode.is_audio_output_active", lambda: False)
monkeypatch.setattr("tools.voice_mode.stop_playback", lambda: None)
return cli
def test_generation_trip_interrupts_agent_and_submits(self, monkeypatch, tmp_path):
"""Speech during generation → agent.interrupt() (the same seam the
typed interrupt uses) + pending TTS pipeline cut + capture queued."""
wav = tmp_path / "fd.wav"
wav.write_bytes(b"RIFF")
def fake_listen(should_stop, is_playing=None, on_trigger=None, **_kw):
on_trigger("generation")
return str(wav)
cli = self._cli(monkeypatch, listen=fake_listen, _agent_running=True)
interrupted = threading.Event()
cli.agent = SimpleNamespace(interrupt=lambda: interrupted.set())
pipe_stop = threading.Event()
cli._voice_tts_stop = pipe_stop
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {"success": True, "transcript": "actually wait"},
)
cli._voice_full_duplex_listener()
assert interrupted.is_set()
assert pipe_stop.is_set() # stale reply's TTS can never play
from cli import _VoiceInputMessage
queued = cli._pending_input.get_nowait()
assert isinstance(queued, _VoiceInputMessage)
assert str(queued) == "actually wait"
assert not cli._voice_barge_capture.is_set()
def test_listener_arms_at_submit_and_survives_into_playback(self, monkeypatch):
"""Lifecycle: should_stop is False during generation AND during
pending TTS (survives the phase transition — no re-arm race), and
True once the turn is fully done."""
probes = {}
def fake_listen(should_stop, is_playing=None, on_trigger=None, **_kw):
# generation: agent running, TTS not started
probes["generation"] = should_stop()
# transition: agent done, TTS still pending
cli._agent_running = False
cli._voice_tts_done.clear()
probes["playback_pending"] = should_stop()
# turn fully done
cli._voice_tts_done.set()
probes["done"] = should_stop()
return None
cli = self._cli(monkeypatch, listen=fake_listen, _agent_running=True)
cli._voice_tts_done.set()
cli._voice_full_duplex_listener()
assert probes["generation"] is False
assert probes["playback_pending"] is False # same listener spans phases
assert probes["done"] is True
def test_stop_phrase_mid_generation_interrupts_and_ends_chat(self, monkeypatch, tmp_path):
"""Bare 'stop' during generation = stop everything: the turn is
interrupted at trip time AND the voice chat is disabled."""
wav = tmp_path / "fd.wav"
wav.write_bytes(b"RIFF")
def fake_listen(should_stop, is_playing=None, on_trigger=None, **_kw):
on_trigger("generation")
return str(wav)
cli = self._cli(monkeypatch, listen=fake_listen, _agent_running=True)
interrupted = threading.Event()
cli.agent = SimpleNamespace(interrupt=lambda: interrupted.set())
disabled = []
cli._disable_voice_mode = lambda: disabled.append(True)
monkeypatch.setattr(
"tools.voice_mode.transcribe_recording",
lambda path, model=None: {"success": True, "transcript": "stop"},
)
monkeypatch.setattr(
"tools.voice_mode_transcript.is_voice_stop_phrase",
lambda text: text.strip().lower() == "stop",
)
cli._voice_full_duplex_listener()
assert interrupted.is_set() # turn interrupted at trip
assert disabled == [True] # chat ended by the stop phrase
assert cli._pending_input.empty() # stop phrase never reaches the agent
# ============================================================================
# Typed stop phrase — typing "stop" during a voice chat ends it
# ============================================================================
class TestTypedVoiceStop:
"""_typed_voice_stop: a TYPED bare stop phrase during an active voice chat
ends the chat (same as saying "stop"); outside voice mode it passes
through to the agent untouched."""
def _cli(self, **overrides):
cli = _make_voice_cli(**overrides)
cli._disable_calls = []
cli._disable_voice_mode = lambda: cli._disable_calls.append(True)
return cli
@pytest.fixture(autouse=True)
def _pin_stop_phrases(self, monkeypatch):
# Hermetic: don't let a dev machine's voice.stop_phrases config
# change which utterances count as a stop phrase.
monkeypatch.setattr(
"tools.voice_mode_transcript._load_voice_stop_phrases", lambda: ("stop",)
)
def test_typed_stop_ends_voice_chat_when_voice_on(self):
cli = self._cli(_voice_mode=True)
assert cli._typed_voice_stop("stop") is True
assert cli._disable_calls == [True]
def test_longer_typed_message_passes_through_in_voice_mode(self):
cli = self._cli(_voice_mode=True)
assert cli._typed_voice_stop("stop the docker container") is False
assert cli._disable_calls == []
# ============================================================================
# Fallback (whole-file) TTS path arms the full-duplex listener
# ============================================================================
class TestFallbackSpeakArmsBargeMonitor:
"""_voice_speak_response_async must arm _voice_full_duplex_listener in
continuous voice mode. This is the safety net for speak calls outside a
chat turn — the primary arm happens at utterance-submit in chat()."""
def _cli(self, **overrides):
cli = _make_voice_cli(**overrides)
cli._monitor_calls = []
cli._monitor_armed = threading.Event()
def _armed():
cli._monitor_calls.append(True)
cli._monitor_armed.set()
cli._voice_full_duplex_listener = _armed
cli._voice_speak_response = lambda text: None
return cli
def test_monitor_armed_in_continuous_voice_mode(self):
cli = self._cli(_voice_mode=True, _voice_tts=True, _voice_continuous=True)
cli._voice_speak_response_async("a reply")
assert cli._monitor_armed.wait(5.0), "listener was never armed"
assert len(cli._monitor_calls) == 1
def test_no_monitor_outside_continuous_mode(self):
cli = self._cli(_voice_mode=True, _voice_tts=True, _voice_continuous=False)
cli._voice_speak_response_async("a reply")
# Nothing to wait for — a short negative window is enough to prove the
# speak thread came and went without arming the mic.
assert not cli._monitor_armed.wait(0.05)
assert cli._monitor_calls == []