Wyoming speech-to-text server for Home Assistant using NVIDIA Parakeet
Loads parakeet-mlx in-process (no HTTP hop) and serves it over the Wyoming
protocol. On an M4 Mac mini this transcribes typical voice commands in ~110ms
versus ~1150ms for a whisper.cpp large-v3 setup, with identical accuracy on a
ten-command benchmark.
Two behaviours matter beyond speed: silence returns an empty string rather
than whisper's "Thank you." hallucination, and there is no decoder context
carried between requests.
Notable implementation details, all covered by mutation-checked regression
tests:
- MLX streams are thread-local, so the model is loaded and evaluated on a
single dedicated worker thread. Splitting those raises
"There is no Stream(cpu, 1) in current thread".
- parakeet_mlx.load_audio() shells out to ffmpeg, which is unnecessary here
since Wyoming delivers 16kHz mono PCM. The mel is built directly via
get_logmel(), whose input must be float32 -- it views the complex STFT
output as the input dtype, so anything narrower doubles the mel bin count.
- Wyoming's run loop has no except clause, so an exception escaping
handle_event closes the connection without sending a Transcript and Home
Assistant waits indefinitely. Failures are caught and returned as an empty
transcript instead.
Defaults to parakeet-tdt-0.6b-v2 rather than the newer multilingual v3
because v2 emits digits ("21 degrees") where v3 spells numbers out, and
Home Assistant's local intent matching expects digits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,45 @@
|
||||
"""Tests for CLI wiring and the Info advertised to Home Assistant."""
|
||||
from wyoming.info import Info
|
||||
|
||||
from wyoming_parakeet.__main__ import DEFAULT_MODEL, build_info, build_parser
|
||||
|
||||
|
||||
def test_default_model_is_v2_not_v3():
|
||||
"""Deliberate: v3 is newer and multilingual but spells numbers out
|
||||
("twenty-one degrees" vs "21 degrees"). Home Assistant's local intent
|
||||
matching wants digits, and both pipelines run prefer_local_intents, so v3
|
||||
would still look accurate while pushing commands onto the LLM fallback.
|
||||
If you are changing this, re-run test/wy-test.py and check cmd2/cmd4/cmd7."""
|
||||
assert DEFAULT_MODEL == "mlx-community/parakeet-tdt-0.6b-v2"
|
||||
|
||||
|
||||
def test_model_defaults_are_applied():
|
||||
args = build_parser().parse_args(["--uri", "tcp://0.0.0.0:7892"])
|
||||
assert args.model == DEFAULT_MODEL
|
||||
assert args.language == "en"
|
||||
assert args.debug is False
|
||||
|
||||
|
||||
def test_model_can_be_overridden():
|
||||
args = build_parser().parse_args(
|
||||
["--uri", "tcp://0.0.0.0:7892", "--model", "mlx-community/other"]
|
||||
)
|
||||
assert args.model == "mlx-community/other"
|
||||
|
||||
|
||||
def test_info_advertises_the_running_model():
|
||||
"""Home Assistant's Wyoming config flow reads this; if it is malformed the
|
||||
integration cannot be added at all."""
|
||||
info = build_info("mlx-community/other", "en")
|
||||
assert Info.is_type(info.event().type)
|
||||
|
||||
program = info.asr[0]
|
||||
assert program.installed is True
|
||||
assert program.models[0].name == "mlx-community/other"
|
||||
assert program.models[0].languages == ["en"]
|
||||
|
||||
|
||||
def test_info_round_trips_through_an_event():
|
||||
info = build_info(DEFAULT_MODEL, "en")
|
||||
restored = Info.from_event(info.event())
|
||||
assert restored.asr[0].models[0].name == DEFAULT_MODEL
|
||||
Reference in New Issue
Block a user