Skip to content

Benchmark 02 · August 2026

Speech-to-text: accuracy and latency for realtime voice agents

Speech-to-text services from across the industry transcribe the same 1,000 audio samples through an identical Pipecat pipeline, scored on semantic word error rate — only the errors that would change an LLM’s understanding — and on TTFS, the time from the moment the speaker falls silent to the final transcript.

The same thousand clips, the same pipeline, the same VAD signal — and two numbers that pull against each other.

Every model transcribes the same 1,000 clips, played at real-time pace, through one Pipecat pipeline behind one Silero VAD — only the STT service changes. Accuracy is semantic word error rate: a difference from the reference counts only if it would change what an LLM understood, so filler words and number formatting are free while wrong names and negations are not. Latency is TTFS, time to final segment — streaming STT has no request to time from, so the clock runs from the VAD’s stop signal, less its own hangover, to the final transcript.

Seven models sit on the frontier — no one else beats them on both axes, and they read as one staircase. NVIDIA Nemotron 3.0 ASR anchors the fast end at a 221 ms median and the lowest P99 of any model, 252 ms, for 1.95% pooled WER. Deepgram nova-3-general takes 26 ms more, at 247 ms, and brings the error rate down to 1.62%. Then a knot of three inside 33 ms and 0.07 points of each other — Soniox stt-rt-v4 at 249 ms, Soniox stt-rt-v5 at 260 ms, and AssemblyAI universal-3-5-pro at 282 ms, running 1.29% down to 1.22% — where the choice is close to a coin flip. Past the knot, AssemblyAI universal-3-6-pro, the successor to u3.5-pro, spends 25 ms more for a quarter point less error: the second-lowest rate, 0.96%, at a 307 ms median, with its tail in hand — 401 ms at P95, 498 ms at P99. At the far end, Meta muse-voice-transcribe-1.0 reaches 0.83%, the lowest of any model, for a 392 ms median, and transcribes 89.2% of clips perfectly — 2.2 points clear of the next model. It pays for that in the tail: a 1,292 ms P95 and a 1,922 ms P99 against that median, the widest spread in the table, so the accuracy arrives with a pause long enough to break a turn about once every twenty calls — where universal-3-6-pro, 85 ms quicker at the median, gives up 0.13 points of WER for a P99 under half a second. Beating Speechmatics linden-1 on both plotted axes is what unseats it — 1.05% for 369 ms no longer holds a frontier spot, though it still betters Speechmatics’ legacy real-time service on every error and latency measure. Under a fifth of a second separates the frontier end to end, and walking it cuts the error rate by 57%.

Models benchmarked
2415 vendors · 1,000 samples each
Lowest pooled WER
0.83%Meta muse-voice · at a 392 ms median
Fastest transcription
247 msDeepgram nova-3 · 1.62% WER
Most consistent
53 msSoniox stt-rt-v5 · 260 ms P50 → 313 ms P99
The trade-offAccuracy against latencyOne dot per model — median TTFS on a log axis against pooled semantic WER. Both are better lower, so the strong corner is bottom-left. Filled dots on the dashed staircase are the Pareto frontier: the 7 models no other model beats on both metrics. Three of them land on top of each other and share a braced label. Hover any dot for its full numbers.
250ms500ms1s1%2%3%4%5%Pooled semantic WERMedian TTFS (log scale) →Soniox v4 · Soniox v5 · AssemblyAI u3.5-proAssemblyAI universal-3-6-pro — 0.96% pooled semantic WER (1.06% mean) · TTFS median 307ms, P95 401ms, P99 498ms · on the frontierAssemblyAI u3.6-proAssemblyAI universal-3-5-pro — 1.22% pooled semantic WER (1.44% mean) · TTFS median 282ms, P95 354ms, P99 393ms · on the frontierAssemblyAI u3.5-proAssemblyAI u3-rt-pro — 1.34% pooled semantic WER (1.74% mean) · TTFS median 335ms, P95 534ms, P99 613msAssemblyAI u3-rt-proAssemblyAI universal-streaming-english — 3.02% pooled semantic WER (3.49% mean) · TTFS median 256ms, P95 362ms, P99 417msAssemblyAI streaming-enAWS — 1.75% pooled semantic WER (1.68% mean) · TTFS median 1136ms, P95 1527ms, P99 1897msAWSAzure — 1.18% pooled semantic WER (1.21% mean) · TTFS median 1016ms, P95 1345ms, P99 1791msAzureCartesia ink-2 — 1.25% pooled semantic WER (1.47% mean) · TTFS median 299ms, P95 328ms, P99 1584msCartesia ink-2Cartesia ink-whisper — 4.36% pooled semantic WER (3.92% mean) · TTFS median 266ms, P95 364ms, P99 898msCartesia ink-whisperDeepgram nova-3-general — 1.62% pooled semantic WER (1.71% mean) · TTFS median 247ms, P95 298ms, P99 326ms · on the frontierDeepgram nova-3ElevenLabs scribe_v2_realtime — 3.12% pooled semantic WER (3.16% mean) · TTFS median 281ms, P95 348ms, P99 407msElevenLabs scribe v2Google gemini-3.5-transcribe-live — 2.24% pooled semantic WER (2.24% mean) · TTFS median 458ms, P95 532ms, P99 599msGemini 3.5 liveGoogle latest-long — 2.85% pooled semantic WER (2.84% mean) · TTFS median 878ms, P95 1155ms, P99 1570msGoogle latest-longGradium default — 3.71% pooled semantic WER (3.56% mean) · TTFS median 570ms, P95 596ms, P99 622msGradiumMeta muse-voice-transcribe-1.0 — 0.83% pooled semantic WER (0.97% mean) · TTFS median 392ms, P95 1292ms, P99 1922ms · on the frontierMeta muse-voiceMistral voxtral-mini-transcribe-realtime-2602 — 4.97% pooled semantic WER (4.44% mean) · TTFS median 525ms, P95 973ms, P99 1913msMistral voxtral-miniNVIDIA Nemotron 3.0 ASR (en) — 1.95% pooled semantic WER (1.90% mean) · TTFS median 221ms, P95 238ms, P99 252ms · on the frontierNVIDIA 3.0NVIDIA Nemotron 3.5 ASR (multilingual) — 4.58% pooled semantic WER (4.54% mean) · TTFS median 236ms, P95 253ms, P99 266msNVIDIA 3.5 multiOpenAI gpt-4o-transcribe — 3.06% pooled semantic WER (3.24% mean) · TTFS median 637ms, P95 965ms, P99 1655msgpt-4o-transcribeOpenAI gpt-realtime-whisper — 2.73% pooled semantic WER (2.92% mean) · TTFS median 740ms, P95 878ms, P99 1080msgpt-realtime-whisperSmallest AI pulse — 2.37% pooled semantic WER (2.30% mean) · TTFS median 398ms, P95 533ms, P99 1593msSmallest pulseSoniox stt-rt-v5 — 1.27% pooled semantic WER (1.34% mean) · TTFS median 260ms, P95 305ms, P99 313ms · on the frontierSoniox v5Soniox stt-rt-v4 — 1.29% pooled semantic WER (1.25% mean) · TTFS median 249ms, P95 281ms, P99 310ms · on the frontierSoniox v4Speechmatics linden-1 — 1.05% pooled semantic WER (1.21% mean) · TTFS median 369ms, P95 438ms, P99 690msSpeechmatics linden-1Speechmatics — 1.07% pooled semantic WER (1.40% mean) · TTFS median 495ms, P95 676ms, P99 736msSpeechmatics legacy

24 models · 15 vendors · 1,000 samples each · one Pipecat pipeline, only the STT service swapped

Tail latencyMedian, P95, and P99 by modelEach bar runs from a model's median TTFS to its P99, notched at P95, so a bar's length is the size of its tail in milliseconds — a model with no tail draws almost no bar. Full ink marks a P99 past 3× the median: fast until, once in a hundred calls, it isn't. Ordered by median.
NVIDIA 3.0
221P99 252ms
NVIDIA 3.5 multi
236P99 266ms
Deepgram nova-3
247P99 326ms
Soniox v4
249P99 310ms
AssemblyAI streaming-en
256P99 417ms
Soniox v5
260P99 313ms
Cartesia ink-whisper
266P99 898ms
ElevenLabs scribe v2
281P99 407ms
AssemblyAI u3.5-pro
282P99 393ms
Cartesia ink-2
299P99 1.58s
AssemblyAI u3.6-pro
307P99 498ms
AssemblyAI u3-rt-pro
335P99 613ms
Speechmatics linden-1
369P99 690ms
Meta muse-voice
392P99 1.92s
Smallest pulse
398P99 1.59s
Gemini 3.5 live
458P99 599ms
Speechmatics legacy
495P99 736ms
Mistral voxtral-mini
525P99 1.91s
Gradium
570P99 622ms
gpt-4o-transcribe
637P99 1.66s
gpt-realtime-whisper
740P99 1.08s
Google latest-long
878P99 1.57s
Azure
1016P99 1.79s
AWS
1136P99 1.90s
500ms1s1.5s2s

24 models · 1,000 samples each · linear scale from zero · hover a row for its full numbers

ModelTranscriptsPerfectWER meanPooled WERTTFS P50TTFS P95TTFS P99
01AssemblyAI universal-3-6-pro ◆99.9%87.0%1.06%0.96%307401498
02AssemblyAI universal-3-5-pro ◆99.9%84.7%1.44%1.22%282354393
03AssemblyAI u3-rt-pro99.8%83.9%1.74%1.34%335534613
04AssemblyAI universal-streaming-english99.8%66.8%3.49%3.02%256362417
05AWS100.0%77.4%1.68%1.75%1,1361,5271,897
06Azure100.0%82.9%1.21%1.18%1,0161,3451,791
07Cartesia ink-2100.0%84.2%1.47%1.25%2993281,584
08Cartesia ink-whisper99.9%60.5%3.92%4.36%266364898
09Deepgram nova-3-general ◆99.8%76.5%1.71%1.62%247298326
10ElevenLabs scribe_v2_realtime99.7%81.3%3.16%3.12%281348407
11Google gemini-3.5-transcribe-live99.9%78.0%2.24%2.24%458532599
12Google latest-long100.0%69.0%2.84%2.85%8781,1551,570
13Gradium default99.8%65.1%3.56%3.71%570596622
14Meta muse-voice-transcribe-1.0 ◆99.9%89.2%0.97%0.83%3921,2921,922
15Mistral voxtral-mini-transcribe-realtime-260299.3%68.8%4.44%4.97%5259731,913
16NVIDIA Nemotron 3.0 ASR (en) ◆100.0%76.1%1.90%1.95%221238252
17NVIDIA Nemotron 3.5 ASR (multilingual)99.6%62.0%4.54%4.58%236253266
18OpenAI gpt-4o-transcribe99.3%75.9%3.24%3.06%6379651,655
19OpenAI gpt-realtime-whisper100.0%72.5%2.92%2.73%7408781,080
20Smallest AI pulse100.0%72.4%2.30%2.37%3985331,593
21Soniox stt-rt-v5 ◆99.8%83.3%1.34%1.27%260305313
22Soniox stt-rt-v4 ◆99.8%84.1%1.25%1.29%249281310
23Speechmatics linden-199.5%84.5%1.21%1.05%369438690
24Speechmatics99.7%83.2%1.40%1.07%495676736

Numbers transcribed from the results summary the repo treats as its single source of truth; vendors contribute rows one at a time, so the board moves. Rows keep the repo’s published order — vendor first, no ranking implied. Both axes matter and neither dominates, so the table stays raw and the two figures above carry the argument; sorting on either metric alone would seat a model at the top that the other metric disqualifies. ◆ marks the seven models on the frontier for pooled WER against median TTFS, the pair the repo’s own trade-off chart plots, derived here by comparing the published numbers. Semantic WER is judged by Claude, which ignores punctuation, capitalization, contractions, plurals, filler words, and number formatting, and counts substitutions that change meaning, hallucinated words, and wrong names, numbers, or negations. Pooled WER divides total errors by total reference words rather than averaging per-sample rates, so it weights long utterances more heavily and reorders the field: legacy Speechmatics climbs from the seventh-lowest mean WER, 1.40%, to the fourth-lowest pooled, 1.07%, while Azure, tied with Speechmatics linden-1 for third on mean at 1.21%, falls to fifth on pooled. Ranking on mean instead would unseat Soniox stt-rt-v5 and AssemblyAI universal-3-5-pro from the frontier and seat no one new; Meta leads both measures either way. The Speechmatics row with no model name is the legacy real-time endpoint, kept for comparison; linden-1 runs on its Agent STT endpoint. Perfect is the share of runs scored at 0% semantic WER; Transcripts is the share of samples that returned a transcription at all. TTFS is the final transcription frame’s arrival minus the speaker’s speech-end time, taken as the VAD stopped-speaking frame less its configured stop_secs hangover. The two NVIDIA Nemotron rows were run self-hosted, so their TTFS carries no public-network round trip and does not compare directly with the hosted services; the tiles above name hosted APIs for that reason, and the prose carries NVIDIA’s numbers. Audio is drawn from smart-turn-data-v3.1-train and the reference transcriptions are generated with Gemini and human-reviewed; both are published as stt-benchmark-data.