Live source snapshot
AudioAgentBench
Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 6 service workflows; synthetic and human audio
- Primary metric
- Checks passed
- Owner
- Arcada Labs
- Available evidence
- Machine-readable public run records
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
Aggregated from benchmark-owner run records
| Rank | Model | Checks passed | Evidence |
|---|---|---|---|
| 1 | Grok Voice Think Fast 2.0 | 80.2% | 10,617 / 13,232 |
| 2 | Grok Voice Think Fast 1.0 | 79.9% | 5,370 / 6,722 |
| 3 | Grok Realtime | 78.3% | 10,526 / 13,440 |
| 4 | Ultravox v0.7 | 78.1% | 9,367 / 11,991 |
| 5 | GPT Realtime | 77.5% | 10,415 / 13,440 |
| 6 | Gemini 2.5 Flash Native Audio | 74.7% | 10,045 / 13,440 |
| 7 | GPT Realtime 2 | 73.1% | 4,909 / 6,720 |
| 8 | Gemini 3.1 Flash Live | 71.3% | 9,583 / 13,440 |
| 9 | Amazon Nova 2 Sonic | 53.7% | 5,999 / 11,163 |
| 10 | GLM Realtime Flash | 29.6% | 2,722 / 9,195 |
| 11 | GLM Realtime Air | 18.3% | 510 / 2,790 |
Interpretation limit
Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.