Reference coverage
VoiceAgentBench
A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 5,500+ synthetic queries; 7 languages
- Primary metric
- Capability-specific task scores
- Owner
- Krutrim AI Labs
- Available evidence
- Paper and public dataset
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
No refreshable result table yet
The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.
Interpretation limit
Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.