Wyoming speech-to-text server for Home Assistant using NVIDIA Parakeet
Loads parakeet-mlx in-process (no HTTP hop) and serves it over the Wyoming
protocol. On an M4 Mac mini this transcribes typical voice commands in ~110ms
versus ~1150ms for a whisper.cpp large-v3 setup, with identical accuracy on a
ten-command benchmark.
Two behaviours matter beyond speed: silence returns an empty string rather
than whisper's "Thank you." hallucination, and there is no decoder context
carried between requests.
Notable implementation details, all covered by mutation-checked regression
tests:
- MLX streams are thread-local, so the model is loaded and evaluated on a
single dedicated worker thread. Splitting those raises
"There is no Stream(cpu, 1) in current thread".
- parakeet_mlx.load_audio() shells out to ffmpeg, which is unnecessary here
since Wyoming delivers 16kHz mono PCM. The mel is built directly via
get_logmel(), whose input must be float32 -- it views the complex STFT
output as the input dtype, so anything narrower doubles the mel bin count.
- Wyoming's run loop has no except clause, so an exception escaping
handle_event closes the connection without sending a Transcript and Home
Assistant waits indefinitely. Failures are caught and returned as an empty
transcript instead.
Defaults to parakeet-tdt-0.6b-v2 rather than the newer multilingual v3
because v2 emits digits ("21 degrees") where v3 spells numbers out, and
Home Assistant's local intent matching expects digits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
"""Wyoming event handler backed by an in-process parakeet-mlx model."""
|
||||
import logging
|
||||
import time
|
||||
|
||||
from wyoming.asr import Transcribe, Transcript
|
||||
from wyoming.audio import AudioChunk, AudioChunkConverter, AudioStop
|
||||
from wyoming.event import Event
|
||||
from wyoming.info import Describe, Info
|
||||
from wyoming.server import AsyncEventHandler
|
||||
|
||||
from .engine import SAMPLE_RATE
|
||||
|
||||
_LOGGER = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class ParakeetEventHandler(AsyncEventHandler):
|
||||
def __init__(self, wyoming_info: Info, cli_args, engine, *args, **kwargs):
|
||||
super().__init__(*args, **kwargs)
|
||||
self.cli_args = cli_args
|
||||
self.wyoming_info_event = wyoming_info.event()
|
||||
self.engine = engine
|
||||
self.audio = bytes()
|
||||
self.converter = AudioChunkConverter(rate=SAMPLE_RATE, width=2, channels=1)
|
||||
|
||||
async def handle_event(self, event: Event) -> bool:
|
||||
if AudioChunk.is_type(event.type):
|
||||
if not self.audio:
|
||||
_LOGGER.debug("Receiving audio")
|
||||
self.audio += self.converter.convert(AudioChunk.from_event(event)).audio
|
||||
return True
|
||||
|
||||
if AudioStop.is_type(event.type):
|
||||
duration = len(self.audio) / (SAMPLE_RATE * 2)
|
||||
started = time.monotonic()
|
||||
try:
|
||||
text = await self.engine.transcribe(self.audio)
|
||||
_LOGGER.info(
|
||||
"%.2fs audio -> %.0fms :: %r",
|
||||
duration,
|
||||
(time.monotonic() - started) * 1000,
|
||||
text,
|
||||
)
|
||||
except Exception:
|
||||
# wyoming's run loop has no except clause, so letting this
|
||||
# propagate closes the connection without ever sending a
|
||||
# Transcript -- Home Assistant then waits for a response that
|
||||
# will never arrive. Fail fast with an empty result instead.
|
||||
_LOGGER.exception("Transcription failed after %.2fs audio", duration)
|
||||
text = ""
|
||||
finally:
|
||||
self.audio = bytes()
|
||||
|
||||
await self.write_event(Transcript(text=text).event())
|
||||
return False
|
||||
|
||||
if Transcribe.is_type(event.type):
|
||||
return True
|
||||
|
||||
if Describe.is_type(event.type):
|
||||
await self.write_event(self.wyoming_info_event)
|
||||
_LOGGER.debug("Sent info")
|
||||
return True
|
||||
|
||||
return True
|
||||
Reference in New Issue
Block a user