Add benchmark harness and compare against 11 other STT backends
Measures median latency, strict exact-match accuracy, digit formatting and silence behaviour across 48 clips (16 Home Assistant commands x 3 macOS TTS voices), on an M4 Mac mini. Headline: Parakeet v2 at 102ms median is 5x faster than the best whisper.cpp configuration and lands within one clip of it on accuracy. mlx-whisper large-v3 is the only backend to score 48/48, at 11x the latency. Moonshine is 2x faster again but gives up real accuracy (35/48). Also quantifies the reason this defaults to v2 over v3: v3 returned digits for only 10 of 21 number-bearing commands, against 21/21 for v2, which is most of the gap between their exact-match scores. The harness feeds audio to every backend as an array rather than a path -- mlx-whisper and moonshine otherwise shell out to ffmpeg, which this project deliberately does not require. Clips are gitignored; bench/make_clips.sh regenerates them. Caveats are documented in the README: this is clean synthetic TTS, so it measures latency rigorously and accuracy only as a domain smoke test, and faster-whisper is CPU-only on Apple Silicon because CTranslate2 has no Metal backend. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -10,23 +10,84 @@ bridge and inference.
|
||||
|
||||
## Why
|
||||
|
||||
Measured on an M4 Mac mini against ten typical Home Assistant voice commands,
|
||||
replacing a whisper.cpp setup:
|
||||
|
||||
| Backend | Mean latency | Correct | Silent input |
|
||||
|---|---|---|---|
|
||||
| whisper.cpp `large-v3` | ~1150 ms | 10/10 | `"Thank you."` |
|
||||
| whisper.cpp `large-v3-turbo` | ~570 ms | 10/10 | `"Thank you."` |
|
||||
| **parakeet-tdt-0.6b-v2** | **~110 ms** | **10/10** | `""` |
|
||||
On an M4 Mac mini, Parakeet transcribes a typical Home Assistant command in
|
||||
**~100 ms** — roughly 5× faster than the best whisper.cpp configuration and
|
||||
9× faster than whisper `large-v3`, while matching them on accuracy. See
|
||||
[Benchmarks](#benchmarks) for the full comparison against eleven other
|
||||
backends.
|
||||
|
||||
Two things matter beyond raw speed:
|
||||
|
||||
- **Silence returns an empty string.** Whisper hallucinates `"Thank you."` on
|
||||
digital silence, which reaches your conversation agent as a real utterance.
|
||||
- **Silence returns an empty string.** Every whisper variant tested
|
||||
hallucinates on digital silence (`"Thank you."`, or `"you"` for
|
||||
faster-whisper), which reaches your conversation agent as a real utterance.
|
||||
Parakeet and Moonshine return nothing.
|
||||
- **No cross-request contamination.** whisper.cpp's server carries decoder
|
||||
context between requests unless you pass `-nc`, and will return the
|
||||
*previous* utterance — in testing, roughly one time in five.
|
||||
|
||||
## Benchmarks
|
||||
|
||||
All figures measured on one machine: **M4 Mac mini (10-core, 32 GB), macOS 26.5**.
|
||||
48 clips — 16 Home Assistant commands rendered through three macOS TTS voices
|
||||
(Daniel, Samantha, Karen). Every backend loads once, transcribes all 48 clips
|
||||
to warm up, then runs a timed pass. Latency is the **median** of that pass;
|
||||
means are skewed by first-request kernel compilation.
|
||||
|
||||
Reproduce with `./bench/make_clips.sh && ./bench/benchmark.py --backend ...`.
|
||||
|
||||
| Backend | Runtime | Median | p90 | Exact | Digits | Silence |
|
||||
|---|---|--:|--:|--:|--:|---|
|
||||
| moonshine tiny | ONNX CPU | 28 ms | 44 ms | 31/48 | 15/21 | `""` |
|
||||
| moonshine base | ONNX CPU | 53 ms | 70 ms | 35/48 | 20/21 | `""` |
|
||||
| **parakeet-tdt-0.6b-v2** | **MLX** | **102 ms** | **112 ms** | **46/48** | **21/21** | `""` |
|
||||
| parakeet-tdt-0.6b-v3 | MLX | 128 ms | 149 ms | 36/48 | 10/21 | `""` |
|
||||
| faster-whisper tiny.en | CPU int8 | 182 ms | 204 ms | 44/48 | 21/21 | `"you"` |
|
||||
| faster-whisper base.en | CPU int8 | 326 ms | 347 ms | 44/48 | 19/21 | `"you"` |
|
||||
| whisper.cpp large-v3-turbo | Metal + CoreML | 518 ms | 537 ms | 47/48 | 21/21 | `"thank you"` |
|
||||
| mlx-whisper large-v3-turbo | MLX | 828 ms | 848 ms | 47/48 | 21/21 | `"thank you"` |
|
||||
| whisper.cpp large-v3 | Metal + CoreML | 941 ms | 1042 ms | 47/48 | 21/21 | `"thank you"` |
|
||||
| faster-whisper small.en | CPU int8 | 974 ms | 1032 ms | 47/48 | 21/21 | `"you"` |
|
||||
| mlx-whisper large-v3 | MLX | 1170 ms | 1251 ms | 48/48 | 21/21 | `"thank you"` |
|
||||
| faster-whisper distil-large-v3 | CPU int8 | 4269 ms | 4304 ms | 46/48 | 21/21 | `"thank you"` |
|
||||
|
||||
**Exact** is a strict string match after normalising case, punctuation and
|
||||
whitespace. **Digits** counts how many of the 21 number-bearing clips came
|
||||
back with digits rather than spelled-out words — see
|
||||
[Model choice](#model-choice-v2-not-v3) for why that matters more than it looks.
|
||||
**Silence** is the output for three seconds of digital silence.
|
||||
|
||||
What the numbers say:
|
||||
|
||||
- **Parakeet v2 has the best latency/accuracy trade-off here.** It is 5×
|
||||
faster than the best whisper.cpp configuration and lands within one clip of
|
||||
it on accuracy.
|
||||
- **`mlx-whisper large-v3` is the accuracy ceiling** — the only backend to
|
||||
score 48/48 — but costs 11× the latency to get there.
|
||||
- **Moonshine is genuinely faster**, at 2× Parakeet's speed, and it also
|
||||
handles silence cleanly. It gives up real accuracy for it (35/48), so it is
|
||||
the right pick only if latency dominates everything else.
|
||||
- **Parakeet v3's 10/21 on digits** is the ITN problem quantified. Its exact
|
||||
match (36/48) is dragged down almost entirely by that one behaviour.
|
||||
- Both clips Parakeet v2 misses are the same word — "aircon", which the TTS
|
||||
voices render as "air con" / "aircan". Accuracy differences at the top of
|
||||
this table are concentrated in a couple of awkward tokens, not spread out.
|
||||
|
||||
### Caveats — read these before trusting the table
|
||||
|
||||
- **This is not a WER benchmark.** The audio is clean synthetic TTS from three
|
||||
similar English voices, with no noise, accents, crosstalk or far-field
|
||||
effects. It measures latency rigorously and accuracy only as a domain smoke
|
||||
test. For real word error rates see the
|
||||
[Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard).
|
||||
- **faster-whisper is CPU-only on Apple Silicon.** CTranslate2 has no Metal
|
||||
backend, so those rows show CPU int8 performance. On an NVIDIA GPU they
|
||||
would look completely different — do not read this as a verdict on
|
||||
faster-whisper generally, only on what it does on this hardware.
|
||||
- **Latency is raw inference**, excluding Wyoming protocol overhead. End to
|
||||
end through this server, expect roughly 15–35 ms on top.
|
||||
- One machine, one run each. Treat differences of a few percent as noise.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Apple Silicon Mac (MLX is Metal/ANE-backed)
|
||||
@@ -73,8 +134,12 @@ multilingual, because of inverse text normalisation:
|
||||
Home Assistant's local intent matching (hassil) expects digits. If your
|
||||
pipeline has `prefer_local_intents` enabled, a model that spells numbers out
|
||||
still *looks* accurate while quietly pushing commands off the fast local path
|
||||
onto your LLM fallback. v3 is the better choice if you need languages other
|
||||
than English — just be aware of the trade.
|
||||
onto your LLM fallback.
|
||||
|
||||
Measured on the benchmark corpus, v3 returned digits for only **10 of 21**
|
||||
number-bearing commands, against **21/21** for v2. That single behaviour is
|
||||
most of the gap between their exact-match scores. v3 remains the right choice
|
||||
if you need languages other than English — just size the trade-off first.
|
||||
|
||||
## Updating the model
|
||||
|
||||
|
||||
Reference in New Issue
Block a user