Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

AudioAgentBench

Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
6 service workflows; synthetic and human audio
Primary metric
Checks passed
Owner
Arcada Labs
Available evidence
Machine-readable public run records

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

Protocol details
Aggregated from benchmark-owner run records
Voice task-completion results: checks passed per model with evidence counts.
RankModelChecks passedEvidence
1Grok Voice Think Fast 2.080.2%10,617 / 13,232
2Grok Voice Think Fast 1.079.9%5,370 / 6,722
3Grok Realtime78.3%10,526 / 13,440
4Ultravox v0.778.1%9,367 / 11,991
5GPT Realtime77.5%10,415 / 13,440
6Gemini 2.5 Flash Native Audio74.7%10,045 / 13,440
7GPT Realtime 273.1%4,909 / 6,720
8Gemini 3.1 Flash Live71.3%9,583 / 13,440
9Amazon Nova 2 Sonic53.7%5,999 / 11,163
10GLM Realtime Flash29.6%2,722 / 9,195
11GLM Realtime Air18.3%510 / 2,790

Interpretation limit

Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.