Skip to main content
BenchLM

Massive Multitask Language Understanding Professional (MMLU-Pro)

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Top models on MMLU-Pro — September 27, 2026

As of September 27, 2026, Qwen3.7 Max leads the MMLU-Pro leaderboard with 89.6% , followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%).

52 modelsKnowledge6% of Knowledge reference weightRefreshingUpdated September 27, 2026

Leaderboard (52 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
89.6%
2
Claude Opus 4.5Anthropic · Closed
89.5%
3
Qwen3.7 PlusAlibaba · Closed
88.5%
4
Qwen3.6 PlusAlibaba · Closed
88.5%
5
Qwen3.5 397BAlibaba · Open weight
87.8%
6
DeepSeek V4 Pro 0813DeepSeek · Open weight
87.5%
7
Kimi K2.5Moonshot AI · Open weight
87.1%
8
Kimi K2.5 (Reasoning)Moonshot AI · Closed
87.1%
9
DeepSeek V4 Pro (High)DeepSeek · Open weight
87.1%
10
Nemotron 3 UltraNVIDIA · Open weight
86.8%
11
Qwen3.5-122B-A10BAlibaba · Open weight
86.7%
12
DeepSeek V4 Flash (High)DeepSeek · Open weight
86.4%
13
Solar Pro 4Upstage · Closed
86.3%
14
Qwen3.6-27BAlibaba · Open weight
86.2%
15
DeepSeek V4 Flash 0731DeepSeek · Open weight
86.2%
16
Solar Open 2Upstage · Open weight
86.2%
17
Qwen3.5-27BAlibaba · Open weight
86.1%
18
GLM-5Z.AI · Open weight
85.7%
19
Qwen3.5-35B-A3BAlibaba · Open weight
85.3%
20
Gemma 4 31BGoogle · Open weight
85.2%
21
Qwen3.6-35B-A3BAlibaba · Open weight
85.2%
22
MAI-Thinking-1Microsoft · Closed
85%
23
MiMo-V2-FlashXiaomi · Open weight
84.9%
24
GLM-4.7Z.AI · Open weight
84.3%
25
K-EXAONE 2.0LG AI Research · Open weight
83.5%
26
Qwen3 235B 2507Alibaba · Open weight
83%
27
DeepSeek V4 FlashDeepSeek · Open weight
83%
28
DeepSeek V4 ProDeepSeek · Open weight
82.9%
29
Gemma 4 26B A4BGoogle · Open weight
82.6%
30
Claude Opus 4.6Anthropic · Closed
82%
31
Exaone 4.0 32BLG AI Research · Open weight
81.8%
32
81.6%
33
LongCat-Flash-Lite-SparseMeituan · Open weight
79.2%
34
Claude Sonnet 4.6Anthropic · Closed
79.2%
35
Granite 4.2 30BIBM · Open weight
77.6%
36
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
77.3%
37
Gemma 4 12BGoogle · Open weight
77.2%
38
Celeris-1Celeris · Closed
75.9%
39
DeepSeek V3DeepSeek · Open weight
75.9%
40
ZAYA1-8BZyphra · Open weight
74.2%
41
Granite 4.2 8BIBM · Open weight
74.0%
42
DeepSeek V4 Pro BaseDeepSeek · Open weight
73.5%
43
MiniCPM5-2BOpenBMB · Open weight
70.8%
44
Gemma 4 E4BGoogle · Open weight
69.4%
45
DeepSeek V4 Flash BaseDeepSeek · Open weight
68.3%
46
ZAYA1-74B-PreviewZyphra · Open weight
68.1%
47
Granite 4.2 3BIBM · Open weight
67.8%
48
Gemma 4 E2BGoogle · Open weight
60%
49
Soofi S 30B-A3BSoofi Project · Open weight
51.4%
50
MiniCPM5-1BOpenBMB · Open weight
48.9%
51
LFM2.5-230MLiquidAI · Open weight
20.3%
52
LFM2.5-VL-450MLiquidAI · Open weight
19.3%

According to BenchLM.ai, Qwen3.7 Max leads the MMLU-Pro benchmark with a score of 89.6%, followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%). The top models are clustered within 1.1 points, suggesting this benchmark is nearing saturation for frontier models.

52 models have been evaluated on MMLU-Pro. The benchmark falls in the Knowledge category. BenchAlign v5.7 gives MMLU-Pro 6% of the Knowledge reference weight, so it moves the Knowledge leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About MMLU-Pro

Year

2024

Tasks

Multiple subjects

Format

10-way multiple choice

Difficulty

Professional level

MMLU-Pro increases the number of choices from 4 to 10 and integrates more reasoning-focused problems, reducing the chance of correct guessing and better evaluating true understanding. It serves as a more robust discriminator of model capabilities.

Freshness and provenance

Version

MMLU-Pro

Refresh cadence

Static

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MMLU-Pro measure?

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Which model scores highest on MMLU-Pro?

Qwen3.7 Max by Alibaba currently leads with a score of 89.6% on MMLU-Pro.

How many models are evaluated on MMLU-Pro?

52 AI models have been evaluated on MMLU-Pro on BenchLM.

Last updated: September 27, 2026 · BenchLM version MMLU-Pro

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.