# MLVU mean average (MLVU (M-Avg))

> A multi-task video understanding benchmark averaged across MLVU categories.

Canonical page: https://benchlm.ai/benchmarks/mlvuavg

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About MLVU (M-Avg)

- Year: 2026
- Tasks: General video understanding
- Format: Video QA and understanding
- Difficulty: Broad multimodal video reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

MLVU captures general-purpose video understanding rather than a single narrow skill. BenchLM tracks the mean-average summary row so provider comparison tables can be compared directly.

MLVU (M-Avg) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (4 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 90.8% |
| 2 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 87.4% |
| 3 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 86.6% |
| 4 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 86.2% |

## FAQ

### What does MLVU (M-Avg) measure?

A multi-task video understanding benchmark averaged across MLVU categories.

### Which model scores highest on MLVU (M-Avg)?

Qwen3.8 Max by Alibaba currently leads with a score of 90.8% on MLVU (M-Avg).

### How many models are evaluated on MLVU (M-Avg)?

4 AI models have been evaluated on MLVU (M-Avg) on BenchLM.

### Does MLVU (M-Avg) affect BenchLM's overall score?

Not directly. MLVU (M-Avg) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MLVU (M-Avg)

- [Qwen3.8 Max vs Qwen3.7 Plus](/compare/qwen3-7-plus-vs-qwen3-8-max)
- [Qwen3.7 Plus vs Qwen3.6-27B](/compare/qwen3-6-27b-vs-qwen3-7-plus)
- [Qwen3.6-27B vs Qwen3.6-35B-A3B](/compare/qwen3-6-27b-vs-qwen3-6-35b-a3b)
