Skip to main content
Reference coverage

EVA-Bench

Separates enterprise task accuracy from conversational experience across realistic service scenarios.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
213 scenarios; 3 enterprise domains; 12 systems
Primary metric
EVA-A accuracy and EVA-X experience
Owner
ServiceNow Research
Available evidence
Paper, code, and public dataset

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?

Available results

No refreshable result table yet

The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.

Interpretation limit

Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.