Add benchmark harness and compare against 11 other STT backends

Measures median latency, strict exact-match accuracy, digit formatting and
silence behaviour across 48 clips (16 Home Assistant commands x 3 macOS TTS
voices), on an M4 Mac mini.

Headline: Parakeet v2 at 102ms median is 5x faster than the best whisper.cpp
configuration and lands within one clip of it on accuracy. mlx-whisper
large-v3 is the only backend to score 48/48, at 11x the latency. Moonshine is
2x faster again but gives up real accuracy (35/48).

Also quantifies the reason this defaults to v2 over v3: v3 returned digits
for only 10 of 21 number-bearing commands, against 21/21 for v2, which is
most of the gap between their exact-match scores.

The harness feeds audio to every backend as an array rather than a path --
mlx-whisper and moonshine otherwise shell out to ffmpeg, which this project
deliberately does not require.

Clips are gitignored; bench/make_clips.sh regenerates them. Caveats are
documented in the README: this is clean synthetic TTS, so it measures latency
rigorously and accuracy only as a domain smoke test, and faster-whisper is
CPU-only on Apple Silicon because CTranslate2 has no Metal backend.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-29 03:32:57 +01:00
co-authored by Claude Opus 5
parent c846bc245c
commit 4475f6d39c
6 changed files with 373 additions and 12 deletions
+77 -12
View File
@@ -10,23 +10,84 @@ bridge and inference.
## Why
Measured on an M4 Mac mini against ten typical Home Assistant voice commands,
replacing a whisper.cpp setup:
| Backend | Mean latency | Correct | Silent input |
|---|---|---|---|
| whisper.cpp `large-v3` | ~1150 ms | 10/10 | `"Thank you."` |
| whisper.cpp `large-v3-turbo` | ~570 ms | 10/10 | `"Thank you."` |
| **parakeet-tdt-0.6b-v2** | **~110 ms** | **10/10** | `""` |
On an M4 Mac mini, Parakeet transcribes a typical Home Assistant command in
**~100 ms** — roughly 5× faster than the best whisper.cpp configuration and
9× faster than whisper `large-v3`, while matching them on accuracy. See
[Benchmarks](#benchmarks) for the full comparison against eleven other
backends.
Two things matter beyond raw speed:
- **Silence returns an empty string.** Whisper hallucinates `"Thank you."` on
digital silence, which reaches your conversation agent as a real utterance.
- **Silence returns an empty string.** Every whisper variant tested
hallucinates on digital silence (`"Thank you."`, or `"you"` for
faster-whisper), which reaches your conversation agent as a real utterance.
Parakeet and Moonshine return nothing.
- **No cross-request contamination.** whisper.cpp's server carries decoder
context between requests unless you pass `-nc`, and will return the
*previous* utterance — in testing, roughly one time in five.
## Benchmarks
All figures measured on one machine: **M4 Mac mini (10-core, 32 GB), macOS 26.5**.
48 clips — 16 Home Assistant commands rendered through three macOS TTS voices
(Daniel, Samantha, Karen). Every backend loads once, transcribes all 48 clips
to warm up, then runs a timed pass. Latency is the **median** of that pass;
means are skewed by first-request kernel compilation.
Reproduce with `./bench/make_clips.sh && ./bench/benchmark.py --backend ...`.
| Backend | Runtime | Median | p90 | Exact | Digits | Silence |
|---|---|--:|--:|--:|--:|---|
| moonshine tiny | ONNX CPU | 28 ms | 44 ms | 31/48 | 15/21 | `""` |
| moonshine base | ONNX CPU | 53 ms | 70 ms | 35/48 | 20/21 | `""` |
| **parakeet-tdt-0.6b-v2** | **MLX** | **102 ms** | **112 ms** | **46/48** | **21/21** | `""` |
| parakeet-tdt-0.6b-v3 | MLX | 128 ms | 149 ms | 36/48 | 10/21 | `""` |
| faster-whisper tiny.en | CPU int8 | 182 ms | 204 ms | 44/48 | 21/21 | `"you"` |
| faster-whisper base.en | CPU int8 | 326 ms | 347 ms | 44/48 | 19/21 | `"you"` |
| whisper.cpp large-v3-turbo | Metal + CoreML | 518 ms | 537 ms | 47/48 | 21/21 | `"thank you"` |
| mlx-whisper large-v3-turbo | MLX | 828 ms | 848 ms | 47/48 | 21/21 | `"thank you"` |
| whisper.cpp large-v3 | Metal + CoreML | 941 ms | 1042 ms | 47/48 | 21/21 | `"thank you"` |
| faster-whisper small.en | CPU int8 | 974 ms | 1032 ms | 47/48 | 21/21 | `"you"` |
| mlx-whisper large-v3 | MLX | 1170 ms | 1251 ms | 48/48 | 21/21 | `"thank you"` |
| faster-whisper distil-large-v3 | CPU int8 | 4269 ms | 4304 ms | 46/48 | 21/21 | `"thank you"` |
**Exact** is a strict string match after normalising case, punctuation and
whitespace. **Digits** counts how many of the 21 number-bearing clips came
back with digits rather than spelled-out words — see
[Model choice](#model-choice-v2-not-v3) for why that matters more than it looks.
**Silence** is the output for three seconds of digital silence.
What the numbers say:
- **Parakeet v2 has the best latency/accuracy trade-off here.** It is 5×
faster than the best whisper.cpp configuration and lands within one clip of
it on accuracy.
- **`mlx-whisper large-v3` is the accuracy ceiling** — the only backend to
score 48/48 — but costs 11× the latency to get there.
- **Moonshine is genuinely faster**, at 2× Parakeet's speed, and it also
handles silence cleanly. It gives up real accuracy for it (35/48), so it is
the right pick only if latency dominates everything else.
- **Parakeet v3's 10/21 on digits** is the ITN problem quantified. Its exact
match (36/48) is dragged down almost entirely by that one behaviour.
- Both clips Parakeet v2 misses are the same word — "aircon", which the TTS
voices render as "air con" / "aircan". Accuracy differences at the top of
this table are concentrated in a couple of awkward tokens, not spread out.
### Caveats — read these before trusting the table
- **This is not a WER benchmark.** The audio is clean synthetic TTS from three
similar English voices, with no noise, accents, crosstalk or far-field
effects. It measures latency rigorously and accuracy only as a domain smoke
test. For real word error rates see the
[Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard).
- **faster-whisper is CPU-only on Apple Silicon.** CTranslate2 has no Metal
backend, so those rows show CPU int8 performance. On an NVIDIA GPU they
would look completely different — do not read this as a verdict on
faster-whisper generally, only on what it does on this hardware.
- **Latency is raw inference**, excluding Wyoming protocol overhead. End to
end through this server, expect roughly 1535 ms on top.
- One machine, one run each. Treat differences of a few percent as noise.
## Requirements
- Apple Silicon Mac (MLX is Metal/ANE-backed)
@@ -73,8 +134,12 @@ multilingual, because of inverse text normalisation:
Home Assistant's local intent matching (hassil) expects digits. If your
pipeline has `prefer_local_intents` enabled, a model that spells numbers out
still *looks* accurate while quietly pushing commands off the fast local path
onto your LLM fallback. v3 is the better choice if you need languages other
than English — just be aware of the trade.
onto your LLM fallback.
Measured on the benchmark corpus, v3 returned digits for only **10 of 21**
number-bearing commands, against **21/21** for v2. That single behaviour is
most of the gap between their exact-match scores. v3 remains the right choice
if you need languages other than English — just size the trade-off first.
## Updating the model