Skip to main content
Reference coverage

VoiceAgentBench

A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
5,500+ synthetic queries; 7 languages
Primary metric
Capability-specific task scores
Owner
Krutrim AI Labs
Available evidence
Paper and public dataset

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

No refreshable result table yet

The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.

Interpretation limit

Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.