Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

MMLU-Pro, Vals AI run (MMLU-Pro (Vals))

Vals AI’s independent run of MMLU-Pro across fourteen academic subjects.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Vals MMLU-Pro mirror

BenchLM mirrors the public Vals AI Vals MMLU-Pro mirror leaderboard captured from https://www.vals.ai/benchmarks/mmlu_pro and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals MMLU-Pro mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

138 Vals rows15 task viewspublic datasetTasks: Overall, Biology, Business, Chemistry, Computer ScienceDisplay only

Vals MMLU-Pro mirror score on MMLU-Pro (Vals) — September 1, 2026

We mirror the published vals mmlu-pro mirror score view for MMLU-Pro (Vals). Claude Fable 5.1 leads the public snapshot at 92.4%, followed by Claude Opus 5 (91.6%) and Claude Fable 5 (91.5%). We do not use these results to rank models overall.

138 modelsKnowledge10% of category scoreCurrentUpdated September 1, 2026

Vals MMLU-Pro mirror score table (138 models)

Score
1
Claude Fable 5.1Anthropic · Closed
92.4%
2
Claude Opus 5Anthropic · Closed
91.6%
3
Claude Fable 5Anthropic · Closed
91.5%
4
Gemini 3.1 Pro PreviewGooglehigh reasoning
91.0%
5
Gemini 3.8 FlashGoogle · Closedhigh reasoning
90.2%
6
Gemini 3.7 FlashGoogle · Closedhigh reasoning
90.1%
8
Claude Opus 4.7Anthropic · Closed
89.9%
9
Claude Opus 4.8Anthropic · Closed
89.6%
10
Gemini 3.5 FlashGoogle · Closedhigh reasoning
89.5%
11
Grok 4.6xAI · Closedhigh reasoning
89.4%
12
Qwen3.7 MaxAlibaba · Closed
89.3%
13
Gemini 3.6 FlashGoogle · Closedhigh reasoning
89.3%
14
Grok 4.5xAI · Closedhigh reasoning
89.2%
15
Claude Opus 4.6 (Adaptive)Anthropic · Closed
89.1%
16
GPT-5.6 SolOpenAI · Closedmax reasoning
89.1%
17
Muse Spark 1.1Meta · Closedxhigh reasoning
88.7%
18
Qwen3.8 MaxAlibaba · Open weight
88.6%
19
Gemini 3 Flash PreviewGooglehigh reasoning
88.6%
20
Muse Spark 1.2Meta · Closedxhigh reasoning
88.3%
21
GPT-5.5OpenAI · Closedxhigh reasoning
88.1%
22
Kimi K3Moonshot AI · Closedmax reasoning
88.0%
24
Qwen3.6 PlusAlibaba · Closed
87.7%
25
Kimi K2.6Moonshot AI · Open weight
87.6%
26
Claude Sonnet 5Anthropic · Closed
87.5%
27
GPT-5.4OpenAI · Closedxhigh reasoning
87.5%
28
Claude Sonnet 4.5 ThinkingAnthropic · Closed
87.4%
29
Claude Sonnet 4.6Anthropic · Closed
87.3%
30
Muse SparkMeta · Closed
87.3%
31
Claude Opus 4.5 ThinkingAnthropic · Closed
87.3%
32
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
87.2%
33
87.2%
34
87.2%
35
87.0%
36
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
87.0%
37
GLM-5.1Z.AI · Open weight
86.9%
38
GLM-5.3Z.AI · Open weightmax reasoning
86.8%
39
GLM-5.2Z.AI · Open weight
86.7%
40
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
86.7%
41
GPT-5OpenAIhigh reasoning
86.5%
42
GPT-5.1OpenAI · Closedhigh reasoning
86.4%
43
InklingThinking Machines Lab · Open weight0.99 reasoning
86.3%
46
GPT-5.2OpenAI · Closedxhigh reasoning
86.2%
47
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
86.2%
48
Claude Opus 4Anthropic
86.2%
49
GLM-5.3-FlashZ.AI · Open weightmax reasoning
86.1%
50
GPT-5.6 LunaOpenAI · Closedmax reasoning
86.0%
51
86.0%
52
85.9%
53
Grok 4.3xAI · Closedhigh reasoning
85.8%
54
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
85.8%
56
o3OpenAI · Closedhigh reasoning
85.6%
57
Claude Opus 4.5Anthropic · Closed
85.6%
58
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
85.6%
59
85.3%
60
Qwen3 MaxAlibaba · Closed
85.0%
61
84.9%
62
MiMo-V2.5-ProXiaomi · Closed
84.6%
63
GPT-5.4 miniOpenAI · Closedxhigh reasoning
84.6%
64
Qwen3 MaxAlibaba · Closed
84.4%
65
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
84.3%
66
MiniMax M3MiniMax · Open weight
84.2%
67
84.2%
68
Qwen3.5 FlashAlibaba · Closed
84.1%
73
83.5%
74
o1OpenAI · Closedhigh reasoning
83.5%
75
DeepSeek-R1DeepSeek · Open weight
83.2%
76
DeepSeek V3p2Fireworks AI
83.1%
77
MiMo-V2.5Xiaomi · Closed
82.9%
78
GLM-4.7Z.AI · Open weight
82.7%
80
GPT-5 miniOpenAI · Closedhigh reasoning
82.2%
81
GLM-4.6Z.AI · Open weight
82.2%
83
81.4%
84
Qwen3 235b A22bFireworks AI
81.2%
85
GLM-4.5Z.AI · Closed
81.2%
86
Kimi K2 ThinkingMoonshot AI
81.1%
87
80.7%
88
O4 MiniOpenAIhigh reasoning
80.6%
89
GPT-4.1OpenAI · Closedhigh reasoning
80.5%
90
MiniMax M2.7MiniMax · Open weight
80.4%
91
MiniMax M2.5MiniMax · Closed
80.1%
92
80.0%
93
79.9%
94
79.8%
95
79.7%
96
DeepSeek V3 0324Fireworks AI
79.5%
97
79.4%
99
79.4%
100
GPT-OSS 120BOpenAI · Open weight
79.2%
102
Claude Haiku 4.5 ThinkingAnthropic · Closed
78.7%
103
o3-miniOpenAI · Closedhigh reasoning
78.7%
105
Claude 3.5 SonnetAnthropic · Closed
78.4%
106
77.4%
107
GPT-4.1 miniOpenAI · Closedhigh reasoning
77.2%
108
GPT-5.4 nanoOpenAI · Closedhigh reasoning
77.2%
109
GPT-5 nanoOpenAI · Closedhigh reasoning
76.1%
110
75.5%
111
Mistral Medium 3.5Mistral AIhigh reasoning
75.3%
112
75.3%
113
75.3%
115
GPT-4oOpenAI · Closedhigh reasoning
74.1%
116
DeepSeek V3DeepSeek · Open weight
73.8%
117
GPT-4oOpenAI · Closedhigh reasoning
72.6%
118
GPT-OSS 20BOpenAI · Open weight
71.6%
122
69.7%
125
69.2%
126
Laguna XS.2Poolside · Open weight
69.0%
127
Laguna M.1Poolside · Closed
68.8%
128
68.7%
129
66.0%
130
65.6%
131
64.4%
132
64.1%
133
GPT-4.1 nanoOpenAI · Closedhigh reasoning
63.5%
134
GPT-4o miniOpenAI · Closedhigh reasoning
62.7%
135
62.1%
136
49.8%
137
44.0%
138
30.3%

The published MMLU-Pro (Vals) snapshot places Claude Fable 5.1 first at 92.4%. The third row is 0.9 points behind. The broader top-10 range is 2.9 points, so many of the published results sit in a relatively narrow band.

138 models have been evaluated on MMLU-Pro (Vals). The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMLU-Pro (Vals) contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.

About MMLU-Pro (Vals)

Year

2026

Tasks

Academic multiple-choice questions

Format

Accuracy

Difficulty

Broad academic knowledge

BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).

BenchLM freshness & provenance

Version

MMLU-Pro (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMLU-Pro (Vals) measure?

Vals AI’s independent run of MMLU-Pro across fourteen academic subjects.

Which model leads the published MMLU-Pro (Vals) snapshot?

Claude Fable 5.1 currently leads the published MMLU-Pro (Vals) snapshot with 92.4% vals mmlu-pro mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MMLU-Pro (Vals)?

The September 1, 2026 snapshot contains 138 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.