# Voice LLM Benchmarks and Voice Agent Leaderboards

> Source-backed results for spoken reasoning, voice-agent task completion, full-duplex interaction, and latency. Incompatible protocols remain separate.

- Source snapshots refreshed: 2026-09-18
- Tracked benchmarks: 41
- Display policy: protocol-specific results only; no combined voice score

## VoiceBench leaders

| Rank | Model | Architecture | Overall |
|---|---|---|---|
| 1 | BR-Voice-Reasoner | vision_audio_llm | 90.46 |
| 2 | LFG-3 | audiollm | 89.88 |
| 3 | NVIDIA Nemotron 3 Nano Omni 30B A3B | vision_audio_llm | 89.39 |
| 4 | Ultravox-GLM-4P7 | audiollm | 88.86 |
| 5 | Qwen3-Omni-30B-A3B-Thinking | vision_audio_llm | 88.80 |
| 6 | Ultravox-GLM-4P7 (thinking) | audiollm | 88.79 |
| 7 | Whisper-v3-large + GPT-4o | cascaded | 87.80 |
| 8 | Ultravox-GLM-4P6 | audiollm | 87.05 |
| 9 | LFG-2 | audiollm | 86.98 |
| 10 | GPT-4o-Audio | omni_hd | 86.75 |

## AudioAgentBench human-audio leaders

| Model | Checks passed | Evidence |
|---|---|---|
| grok-voice-think-fast-2.0 | 80.2% | 10617/13232 |
| grok-voice-think-fast-1.0 | 79.9% | 5370/6722 |
| grok-realtime | 78.3% | 10526/13440 |
| ultravox-v0.7 | 78.1% | 9367/11991 |
| gpt-realtime | 77.5% | 10415/13440 |
| gemini-2.5-flash-native-audio-preview-12-2025 | 74.7% | 10045/13440 |
| gpt-realtime-2 | 73.1% | 4909/6720 |
| gemini-3.1-flash-live-preview | 71.3% | 9583/13440 |
| amazon.nova-2-sonic-v1:0 | 53.7% | 5999/11163 |
| glm-realtime-flash | 29.6% | 2722/9195 |

## Browse

- [Benchmark directory](/voice-benchmarks/benchmarks)
- [Methodology](/voice-benchmarks/methodology)

Canonical page: https://benchlm.ai/voice-benchmarks
