Wyoming speech-to-text server for Home Assistant using NVIDIA Parakeet

Loads parakeet-mlx in-process (no HTTP hop) and serves it over the Wyoming
protocol. On an M4 Mac mini this transcribes typical voice commands in ~110ms
versus ~1150ms for a whisper.cpp large-v3 setup, with identical accuracy on a
ten-command benchmark.

Two behaviours matter beyond speed: silence returns an empty string rather
than whisper's "Thank you." hallucination, and there is no decoder context
carried between requests.

Notable implementation details, all covered by mutation-checked regression
tests:

- MLX streams are thread-local, so the model is loaded and evaluated on a
  single dedicated worker thread. Splitting those raises
  "There is no Stream(cpu, 1) in current thread".
- parakeet_mlx.load_audio() shells out to ffmpeg, which is unnecessary here
  since Wyoming delivers 16kHz mono PCM. The mel is built directly via
  get_logmel(), whose input must be float32 -- it views the complex STFT
  output as the input dtype, so anything narrower doubles the mel bin count.
- Wyoming's run loop has no except clause, so an exception escaping
  handle_event closes the connection without sending a Transcript and Home
  Assistant waits indefinitely. Failures are caught and returned as an empty
  transcript instead.

Defaults to parakeet-tdt-0.6b-v2 rather than the newer multilingual v3
because v2 emits digits ("21 degrees") where v3 spells numbers out, and
Home Assistant's local intent matching expects digits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-29 03:04:48 +01:00
co-authored by Claude Opus 5
commit 4286e88344
18 changed files with 1168 additions and 0 deletions
+45
View File
@@ -0,0 +1,45 @@
"""Tests for CLI wiring and the Info advertised to Home Assistant."""
from wyoming.info import Info
from wyoming_parakeet.__main__ import DEFAULT_MODEL, build_info, build_parser
def test_default_model_is_v2_not_v3():
"""Deliberate: v3 is newer and multilingual but spells numbers out
("twenty-one degrees" vs "21 degrees"). Home Assistant's local intent
matching wants digits, and both pipelines run prefer_local_intents, so v3
would still look accurate while pushing commands onto the LLM fallback.
If you are changing this, re-run test/wy-test.py and check cmd2/cmd4/cmd7."""
assert DEFAULT_MODEL == "mlx-community/parakeet-tdt-0.6b-v2"
def test_model_defaults_are_applied():
args = build_parser().parse_args(["--uri", "tcp://0.0.0.0:7892"])
assert args.model == DEFAULT_MODEL
assert args.language == "en"
assert args.debug is False
def test_model_can_be_overridden():
args = build_parser().parse_args(
["--uri", "tcp://0.0.0.0:7892", "--model", "mlx-community/other"]
)
assert args.model == "mlx-community/other"
def test_info_advertises_the_running_model():
"""Home Assistant's Wyoming config flow reads this; if it is malformed the
integration cannot be added at all."""
info = build_info("mlx-community/other", "en")
assert Info.is_type(info.event().type)
program = info.asr[0]
assert program.installed is True
assert program.models[0].name == "mlx-community/other"
assert program.models[0].languages == ["en"]
def test_info_round_trips_through_an_event():
info = build_info(DEFAULT_MODEL, "en")
restored = Info.from_event(info.event())
assert restored.asr[0].models[0].name == DEFAULT_MODEL