Pipecat research
Voice AI, measured in the open
Controlled, repeatable studies of the parts that make up a voice agent — the models, the voices, the turn-taking. Every benchmark pins the full Pipecat pipeline and changes one variable, so when a number moves it’s the component, not the harness.
All benchmarks
03 published
Benchmark 03
PhoneBench Alpha 1: phone-assistant tool calling and dialogue
15 models scored on multi-turn phone-assistant tool calling and dialogue across realistic customer-service scenarios, judged against reference responses by a panel of LLM judges — with time-to-first-answer-token and estimated cost per minute measured alongside accuracy.
15 models · LLM judge panelBenchmark 02
Speech-to-text: accuracy and latency for realtime voice agents
Speech-to-text services from across the industry transcribe the same 1,000 audio samples through an identical Pipecat pipeline, scored on semantic word error rate — only the errors that would change an LLM’s understanding — and on TTFS, the time from the moment the speaker falls silent to the final transcript.
Semantic WER · TTFS latencyBenchmark 01
Voice readiness: text LLMs on 30-turn conversations
46 model configs from 7 providers run the same scripted 30-turn conversations — tool use, instruction following, and knowledge-base grounding judged on every turn — then get re-ranked inside the ~700 ms latency budget voice actually allows.
46 configs · 7 providers