Pipecat research
Voice AI, measured in the open
Controlled, repeatable studies of the parts that make up a voice agent — the models, the voices, the turn-taking. Every benchmark pins the full Pipecat pipeline and changes one variable, so when a number moves it’s the component, not the harness.
All benchmarks
02 published
Benchmark 02
PhoneBench Alpha 1: phone-assistant tool calling and dialogue
15 models scored on multi-turn phone-assistant tool calling and dialogue across realistic customer-service scenarios, judged against reference responses by a panel of LLM judges — with time-to-first-answer-token and estimated cost per minute measured alongside accuracy.
15 models · LLM judge panelBenchmark 01
Voice readiness: text LLMs on 30-turn conversations
46 model configs from 7 providers run the same scripted 30-turn conversations — tool use, instruction following, and knowledge-base grounding judged on every turn — then get re-ranked inside the ~700 ms latency budget voice actually allows.
46 configs · 7 providers