VITA-Bench
We show this table for reference; we do not rank on it.
An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.
Benchmark score on VITA-Bench — September 27, 2026
We compile the VITA-Bench rows from benchmark-owner or independent runs and provider self-reports. Qwen3.7 Max leads the table at 47.9%, followed by Qwen3.7 Plus (45.6%) and Qwen3.6 Plus (44.3%). We do not use these results to rank models overall.
Qwen3.7 Max
Alibaba
Qwen3.7 Plus
Alibaba
Qwen3.6 Plus
Alibaba
12 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (12 models)
ScoreAmong the reported VITA-Bench rows, Qwen3.7 Max is first at 47.9%. The third row is 3.6 points behind. The broader top-10 range is 29.4 points, so the table still separates the published systems.
12 models have been evaluated on VITA-Bench. The benchmark falls in the Agentic category. VITA-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About VITA-Bench
Year
2025
Tasks
Interactive consumer-service agent tasks
Format
End-to-end interactive agent evaluation
Difficulty
Long-horizon real-world workflows
VITA-Bench is built to test realistic interactive agent behavior rather than toy tool calls. It stresses long-horizon coordination, tool selection, changing user intent, and domain switching across daily-life applications.
Freshness and provenance
Version
VITA-Bench 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does VITA-Bench measure?
An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.
Which model scores highest on VITA-Bench?
Qwen3.7 Max by Alibaba currently leads with a score of 47.9% on VITA-Bench.
How many models are evaluated on VITA-Bench?
12 AI models have been evaluated on VITA-Bench on BenchLM.
Compare top models on VITA-Bench
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.