The same thousand clips, the same pipeline, the same VAD signal — and two numbers that pull against each other.
Every model transcribes the same 1,000 clips, played at real-time pace, through one Pipecat pipeline behind one Silero VAD — only the STT service changes. Accuracy is semantic word error rate: a difference from the reference counts only if it would change what an LLM understood, so filler words and number formatting are free while wrong names and negations are not. Latency is TTFS, time to final segment — streaming STT has no request to time from, so the clock runs from the VAD’s stop signal, less its own hangover, to the final transcript.
Seven models sit on the frontier — no one else beats them on both axes, and they read as one staircase. NVIDIA Nemotron 3.0 ASR anchors the fast end at a 221 ms median and the lowest P99 of any model, 252 ms, for 1.95% pooled WER. Deepgram nova-3-general takes 26 ms more, at 247 ms, and brings the error rate down to 1.62%. Then a knot of three inside 33 ms and 0.07 points of each other — Soniox stt-rt-v4 at 249 ms, Soniox stt-rt-v5 at 260 ms, and AssemblyAI universal-3-5-pro at 282 ms, running 1.29% down to 1.22% — where the choice is close to a coin flip. Past the knot, AssemblyAI universal-3-6-pro, the successor to u3.5-pro, spends 25 ms more for a quarter point less error: the second-lowest rate, 0.96%, at a 307 ms median, with its tail in hand — 401 ms at P95, 498 ms at P99. At the far end, Meta muse-voice-transcribe-1.0 reaches 0.83%, the lowest of any model, for a 392 ms median, and transcribes 89.2% of clips perfectly — 2.2 points clear of the next model. It pays for that in the tail: a 1,292 ms P95 and a 1,922 ms P99 against that median, the widest spread in the table, so the accuracy arrives with a pause long enough to break a turn about once every twenty calls — where universal-3-6-pro, 85 ms quicker at the median, gives up 0.13 points of WER for a P99 under half a second. Beating Speechmatics linden-1 on both plotted axes is what unseats it — 1.05% for 369 ms no longer holds a frontier spot, though it still betters Speechmatics’ legacy real-time service on every error and latency measure. Under a fifth of a second separates the frontier end to end, and walking it cuts the error rate by 57%.
- Models benchmarked
- 2415 vendors · 1,000 samples each
- Lowest pooled WER
- 0.83%Meta muse-voice · at a 392 ms median
- Fastest transcription
- 247 msDeepgram nova-3 · 1.62% WER
- Most consistent
- 53 msSoniox stt-rt-v5 · 260 ms P50 → 313 ms P99
24 models · 15 vendors · 1,000 samples each · one Pipecat pipeline, only the STT service swapped
24 models · 1,000 samples each · linear scale from zero · hover a row for its full numbers
| Model | Transcripts | Perfect | WER mean | Pooled WER | TTFS P50 | TTFS P95 | TTFS P99 |
|---|---|---|---|---|---|---|---|
| 01AssemblyAI universal-3-6-pro ◆ | 99.9% | 87.0% | 1.06% | 0.96% | 307 | 401 | 498 |
| 02AssemblyAI universal-3-5-pro ◆ | 99.9% | 84.7% | 1.44% | 1.22% | 282 | 354 | 393 |
| 03AssemblyAI u3-rt-pro | 99.8% | 83.9% | 1.74% | 1.34% | 335 | 534 | 613 |
| 04AssemblyAI universal-streaming-english | 99.8% | 66.8% | 3.49% | 3.02% | 256 | 362 | 417 |
| 05AWS | 100.0% | 77.4% | 1.68% | 1.75% | 1,136 | 1,527 | 1,897 |
| 06Azure | 100.0% | 82.9% | 1.21% | 1.18% | 1,016 | 1,345 | 1,791 |
| 07Cartesia ink-2 | 100.0% | 84.2% | 1.47% | 1.25% | 299 | 328 | 1,584 |
| 08Cartesia ink-whisper | 99.9% | 60.5% | 3.92% | 4.36% | 266 | 364 | 898 |
| 09Deepgram nova-3-general ◆ | 99.8% | 76.5% | 1.71% | 1.62% | 247 | 298 | 326 |
| 10ElevenLabs scribe_v2_realtime | 99.7% | 81.3% | 3.16% | 3.12% | 281 | 348 | 407 |
| 11Google gemini-3.5-transcribe-live | 99.9% | 78.0% | 2.24% | 2.24% | 458 | 532 | 599 |
| 12Google latest-long | 100.0% | 69.0% | 2.84% | 2.85% | 878 | 1,155 | 1,570 |
| 13Gradium default | 99.8% | 65.1% | 3.56% | 3.71% | 570 | 596 | 622 |
| 14Meta muse-voice-transcribe-1.0 ◆ | 99.9% | 89.2% | 0.97% | 0.83% | 392 | 1,292 | 1,922 |
| 15Mistral voxtral-mini-transcribe-realtime-2602 | 99.3% | 68.8% | 4.44% | 4.97% | 525 | 973 | 1,913 |
| 16NVIDIA Nemotron 3.0 ASR (en) ◆ | 100.0% | 76.1% | 1.90% | 1.95% | 221 | 238 | 252 |
| 17NVIDIA Nemotron 3.5 ASR (multilingual) | 99.6% | 62.0% | 4.54% | 4.58% | 236 | 253 | 266 |
| 18OpenAI gpt-4o-transcribe | 99.3% | 75.9% | 3.24% | 3.06% | 637 | 965 | 1,655 |
| 19OpenAI gpt-realtime-whisper | 100.0% | 72.5% | 2.92% | 2.73% | 740 | 878 | 1,080 |
| 20Smallest AI pulse | 100.0% | 72.4% | 2.30% | 2.37% | 398 | 533 | 1,593 |
| 21Soniox stt-rt-v5 ◆ | 99.8% | 83.3% | 1.34% | 1.27% | 260 | 305 | 313 |
| 22Soniox stt-rt-v4 ◆ | 99.8% | 84.1% | 1.25% | 1.29% | 249 | 281 | 310 |
| 23Speechmatics linden-1 | 99.5% | 84.5% | 1.21% | 1.05% | 369 | 438 | 690 |
| 24Speechmatics | 99.7% | 83.2% | 1.40% | 1.07% | 495 | 676 | 736 |
Numbers transcribed from the results summary the repo treats as its single source of truth; vendors contribute rows one at a time, so the board moves. Rows keep the repo’s published order — vendor first, no ranking implied. Both axes matter and neither dominates, so the table stays raw and the two figures above carry the argument; sorting on either metric alone would seat a model at the top that the other metric disqualifies. ◆ marks the seven models on the frontier for pooled WER against median TTFS, the pair the repo’s own trade-off chart plots, derived here by comparing the published numbers. Semantic WER is judged by Claude, which ignores punctuation, capitalization, contractions, plurals, filler words, and number formatting, and counts substitutions that change meaning, hallucinated words, and wrong names, numbers, or negations. Pooled WER divides total errors by total reference words rather than averaging per-sample rates, so it weights long utterances more heavily and reorders the field: legacy Speechmatics climbs from the seventh-lowest mean WER, 1.40%, to the fourth-lowest pooled, 1.07%, while Azure, tied with Speechmatics linden-1 for third on mean at 1.21%, falls to fifth on pooled. Ranking on mean instead would unseat Soniox stt-rt-v5 and AssemblyAI universal-3-5-pro from the frontier and seat no one new; Meta leads both measures either way. The Speechmatics row with no model name is the legacy real-time endpoint, kept for comparison; linden-1 runs on its Agent STT endpoint. Perfect is the share of runs scored at 0% semantic WER; Transcripts is the share of samples that returned a transcription at all. TTFS is the final transcription frame’s arrival minus the speaker’s speech-end time, taken as the VAD stopped-speaking frame less its configured stop_secs hangover. The two NVIDIA Nemotron rows were run self-hosted, so their TTFS carries no public-network round trip and does not compare directly with the hosted services; the tiles above name hosted APIs for that reason, and the prose carries NVIDIA’s numbers. Audio is drawn from smart-turn-data-v3.1-train and the reference transcriptions are generated with Gemini and human-reviewed; both are published as stt-benchmark-data.