Skip to main content
Live source snapshot

AudioAgentBench

Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
6 service workflows; synthetic and human audio
Primary metric
Checks passed
Owner
Arcada Labs
Available evidence
Machine-readable public run records

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

Protocol details
Aggregated from benchmark-owner run records
RankModelChecks passedEvidence
1Grok Voice Think Fast 2.080.2%10,617 / 13,232
2Grok Voice Think Fast 1.079.9%5,370 / 6,722
3Grok Realtime78.3%10,526 / 13,440
4Ultravox v0.778.1%9,367 / 11,991
5GPT Realtime77.5%10,415 / 13,440
6Gemini 2.5 Flash Native Audio74.7%10,045 / 13,440
7GPT Realtime 273.1%4,909 / 6,720
8Gemini 3.1 Flash Live71.3%9,583 / 13,440
9Amazon Nova 2 Sonic53.7%5,999 / 11,163
10GLM Realtime Flash29.6%2,722 / 9,195
11GLM Realtime Air18.3%510 / 2,790

Interpretation limit

Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.