# AudioAgentBench: Voice Benchmark Profile

> Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

- Scope: 6 service workflows; synthetic and human audio
- Measurement lanes: Task completion, Conversation dynamics
- Primary metric: Checks passed
- Available evidence: Machine-readable public run records
- Owner: Arcada Labs
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.

## Primary sources

- [Code](https://github.com/Design-Arena/audio-agent-bench)
- [Owner page](https://audioarena.ai/methodology)

Canonical page: https://benchlm.ai/voice-benchmarks/audioagentbench
