# EVA-Bench: Voice Benchmark Profile

> Separates enterprise task accuracy from conversational experience across realistic service scenarios.

- Scope: 213 scenarios; 3 enterprise domains; 12 systems
- Measurement lanes: Task completion, Conversation dynamics, Voice experience
- Primary metric: EVA-A accuracy and EVA-X experience
- Available evidence: Machine-readable benchmark-owner leaderboard, code, and public dataset
- Owner: ServiceNow Research
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.

## Primary sources

- [Paper](https://arxiv.org/abs/2605.13841)
- [Code](https://github.com/ServiceNow/eva)
- [Dataset](https://huggingface.co/datasets/ServiceNow-AI/eva)
- [Owner page](https://servicenow.github.io/eva)

Canonical page: https://benchlm.ai/voice-benchmarks/eva-bench
