Loads parakeet-mlx in-process (no HTTP hop) and serves it over the Wyoming
protocol. On an M4 Mac mini this transcribes typical voice commands in ~110ms
versus ~1150ms for a whisper.cpp large-v3 setup, with identical accuracy on a
ten-command benchmark.
Two behaviours matter beyond speed: silence returns an empty string rather
than whisper's "Thank you." hallucination, and there is no decoder context
carried between requests.
Notable implementation details, all covered by mutation-checked regression
tests:
- MLX streams are thread-local, so the model is loaded and evaluated on a
single dedicated worker thread. Splitting those raises
"There is no Stream(cpu, 1) in current thread".
- parakeet_mlx.load_audio() shells out to ffmpeg, which is unnecessary here
since Wyoming delivers 16kHz mono PCM. The mel is built directly via
get_logmel(), whose input must be float32 -- it views the complex STFT
output as the input dtype, so anything narrower doubles the mel bin count.
- Wyoming's run loop has no except clause, so an exception escaping
handle_event closes the connection without sending a Transcript and Home
Assistant waits indefinitely. Failures are caught and returned as an empty
transcript instead.
Defaults to parakeet-tdt-0.6b-v2 rather than the newer multilingual v3
because v2 emits digits ("21 degrees") where v3 spells numbers out, and
Home Assistant's local intent matching expects digits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
65 lines
2.3 KiB
Python
65 lines
2.3 KiB
Python
"""Wyoming event handler backed by an in-process parakeet-mlx model."""
|
|
import logging
|
|
import time
|
|
|
|
from wyoming.asr import Transcribe, Transcript
|
|
from wyoming.audio import AudioChunk, AudioChunkConverter, AudioStop
|
|
from wyoming.event import Event
|
|
from wyoming.info import Describe, Info
|
|
from wyoming.server import AsyncEventHandler
|
|
|
|
from .engine import SAMPLE_RATE
|
|
|
|
_LOGGER = logging.getLogger(__name__)
|
|
|
|
|
|
class ParakeetEventHandler(AsyncEventHandler):
|
|
def __init__(self, wyoming_info: Info, cli_args, engine, *args, **kwargs):
|
|
super().__init__(*args, **kwargs)
|
|
self.cli_args = cli_args
|
|
self.wyoming_info_event = wyoming_info.event()
|
|
self.engine = engine
|
|
self.audio = bytes()
|
|
self.converter = AudioChunkConverter(rate=SAMPLE_RATE, width=2, channels=1)
|
|
|
|
async def handle_event(self, event: Event) -> bool:
|
|
if AudioChunk.is_type(event.type):
|
|
if not self.audio:
|
|
_LOGGER.debug("Receiving audio")
|
|
self.audio += self.converter.convert(AudioChunk.from_event(event)).audio
|
|
return True
|
|
|
|
if AudioStop.is_type(event.type):
|
|
duration = len(self.audio) / (SAMPLE_RATE * 2)
|
|
started = time.monotonic()
|
|
try:
|
|
text = await self.engine.transcribe(self.audio)
|
|
_LOGGER.info(
|
|
"%.2fs audio -> %.0fms :: %r",
|
|
duration,
|
|
(time.monotonic() - started) * 1000,
|
|
text,
|
|
)
|
|
except Exception:
|
|
# wyoming's run loop has no except clause, so letting this
|
|
# propagate closes the connection without ever sending a
|
|
# Transcript -- Home Assistant then waits for a response that
|
|
# will never arrive. Fail fast with an empty result instead.
|
|
_LOGGER.exception("Transcription failed after %.2fs audio", duration)
|
|
text = ""
|
|
finally:
|
|
self.audio = bytes()
|
|
|
|
await self.write_event(Transcript(text=text).event())
|
|
return False
|
|
|
|
if Transcribe.is_type(event.type):
|
|
return True
|
|
|
|
if Describe.is_type(event.type):
|
|
await self.write_event(self.wyoming_info_event)
|
|
_LOGGER.debug("Sent info")
|
|
return True
|
|
|
|
return True
|