# VITA-Bench

> An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.

Canonical page: https://benchlm.ai/benchmarks/vitabench

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About VITA-Bench

- Year: 2025
- Tasks: Interactive consumer-service agent tasks
- Format: End-to-end interactive agent evaluation
- Difficulty: Long-horizon real-world workflows
- Paper: [VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications](https://vitabench.github.io/)

VITA-Bench is built to test realistic interactive agent behavior rather than toy tool calls. It stresses long-horizon coordination, tool selection, changing user intent, and domain switching across daily-life applications.

VITA-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (12 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 47.9% |
| 2 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 45.6% |
| 3 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 44.3% |
| 4 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 43.7% |
| 5 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 40.3% |
| 6 | [Agents-A1](/models/agents-a1) | InternScience | 38.8% |
| 7 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 35.6% |
| 8 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 23.3% |
| 9 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 21.7% |
| 10 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 18.5% |
| 11 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 17.0% |
| 12 | [GLM-4.7](/models/glm-4-7) | Z.AI | 15.5% |

## FAQ

### What does VITA-Bench measure?

An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.

### Which model scores highest on VITA-Bench?

Qwen3.7 Max by Alibaba currently leads with a score of 47.9% on VITA-Bench.

### How many models are evaluated on VITA-Bench?

12 AI models have been evaluated on VITA-Bench on BenchLM.

### Does VITA-Bench affect BenchLM's overall score?

Not directly. VITA-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on VITA-Bench

- [Qwen3.7 Max vs Qwen3.7 Plus](/compare/qwen3-7-max-vs-qwen3-7-plus)
- [Qwen3.7 Plus vs Qwen3.6 Plus](/compare/qwen3-6-plus-vs-qwen3-7-plus)
- [Qwen3.6 Plus vs Qwen3.5 397B](/compare/qwen3-5-397b-vs-qwen3-6-plus)
- [Qwen3.5 397B vs Agents-A1-4B](/compare/agents-a1-4b-vs-qwen3-5-397b)
