Reference coverage
EVA-Bench
Separates enterprise task accuracy from conversational experience across realistic service scenarios.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 213 scenarios; 3 enterprise domains; 12 systems
- Primary metric
- EVA-A accuracy and EVA-X experience
- Owner
- ServiceNow Research
- Available evidence
- Paper, code, and public dataset
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?
Available results
No refreshable result table yet
The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.
Interpretation limit
Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.