Skip to main content

Benchmark profile

Vals-hosted MMLU-Pro mirror (Vals MMLU-Pro mirror)

Vals AI hosted MMLU-Pro view with subject-level task splits.

Data verified

How BenchLM shows Vals MMLU-Pro mirror

BenchLM mirrors the public Vals AI Vals MMLU-Pro mirror leaderboard captured from https://www.vals.ai/benchmarks/mmlu_pro and updated by Vals on July 23, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals MMLU-Pro mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

125 Vals rows15 task viewspublic datasetTasks: Overall, Biology, Business, Chemistry, Computer ScienceDisplay only

Vals MMLU-Pro mirror score on Vals MMLU-Pro mirror — July 23, 2026

BenchLM mirrors the published vals mmlu-pro mirror score view for Vals MMLU-Pro mirror. Claude Opus 5 leads the public snapshot at 91.59% , followed by Claude Fable 5 (91.50%) and Gemini 3.1 Pro Preview (90.99%). BenchLM does not use these results to rank models overall.

125 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 23, 2026

Vals MMLU-Pro mirror score table (125 models)

Score
1
Claude Opus 5Anthropic
91.59%
2
91.50%
5
89.87%
6
89.58%
7
89.52%
8
89.31%
9
89.28%
10
89.22%
11
89.11%
12
89.10%
13
88.73%
15
GPT-5.5OpenAI
88.14%
16
Kimi K3Moonshot AI
87.97%
18
87.67%
19
Kimi K2.6Moonshot AI
87.57%
20
87.55%
21
GPT-5.4OpenAI
87.48%
23
87.34%
24
87.32%
26
87.25%
27
87.21%
28
87.18%
29
87.05%
30
GLM 5.1Zhipu AI
86.90%
31
GLM 5.2Zhipu AI
86.71%
32
86.66%
33
GPT-5OpenAI
86.54%
34
GPT-5.1OpenAI
86.38%
35
InklingThinkingmachines
86.30%
38
GPT-5.2OpenAI
86.23%
39
Claude Opus 4Anthropic
86.17%
40
86.04%
41
86.03%
42
85.91%
43
85.84%
44
85.84%
46
O3OpenAI
85.59%
47
85.59%
48
85.30%
49
Qwen3 MaxAlibaba
84.98%
50
84.92%
51
84.59%
52
84.55%
53
Qwen3 MaxAlibaba
84.36%
54
MiniMax M3MiniMax
84.22%
56
84.06%
61
83.54%
62
O1OpenAI
83.49%
63
DeepSeek R1Fireworks AI
83.18%
64
DeepSeek V3p2Fireworks AI
83.06%
65
Mimo V2.5Xiaomi
82.93%
66
GLM 4.7Zhipu AI
82.74%
68
82.23%
69
GLM 4.6Zhipu AI
82.20%
71
Qwen3 235b A22bFireworks AI
81.25%
72
GLM 4.5Zhipu AI
81.22%
73
Kimi K2 ThinkingMoonshot AI
81.07%
74
80.66%
75
O4 MiniOpenAI
80.56%
76
GPT-4.1OpenAI
80.50%
77
80.43%
78
80.09%
80
79.95%
81
79.82%
83
DeepSeek V3 0324Fireworks AI
79.47%
84
79.43%
85
79.42%
86
79.39%
87
GPT Oss 120bFireworks AI
79.17%
90
O3 MiniOpenAI
78.69%
92
78.40%
93
77.38%
94
77.22%
95
77.17%
96
76.07%
97
75.47%
98
75.33%
99
75.29%
100
75.29%
102
GPT-4oOpenAI
74.13%
103
DeepSeek V3Fireworks AI
73.82%
104
GPT-4oOpenAI
72.56%
105
GPT Oss 20bFireworks AI
71.64%
109
69.71%
112
Laguna Xs.2Poolside
69.41%
113
69.17%
114
68.66%
115
66.02%
116
65.61%
117
64.44%
118
64.12%
119
63.48%
120
Laguna M.1Poolside
63.48%
121
62.73%
122
62.13%
123
49.78%
124
44.00%
125
30.28%

The published Vals MMLU-Pro mirror snapshot places Claude Opus 5 first at 91.59%. The third row is 0.60 points behind. The broader top-10 range is 2.38 points, so many of the published results sit in a relatively narrow band.

125 models have been evaluated on Vals MMLU-Pro mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals MMLU-Pro mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals MMLU-Pro mirror

Year

2026

Tasks

MMLU-Pro subject splits

Format

Accuracy score

Difficulty

Professional academic reasoning

BenchLM keeps this Vals-hosted MMLU-Pro table separate from canonical MMLU-Pro source records.

BenchLM freshness & provenance

Version

Vals MMLU-Pro mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals MMLU-Pro mirror measure?

Vals AI hosted MMLU-Pro view with subject-level task splits.

Which model leads the published Vals MMLU-Pro mirror snapshot?

Claude Opus 5 currently leads the published Vals MMLU-Pro mirror snapshot with 91.59% vals mmlu-pro mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals MMLU-Pro mirror?

125 AI models are included in BenchLM's mirrored Vals MMLU-Pro mirror snapshot, based on the public leaderboard captured on July 23, 2026.

Last updated: July 23, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.