Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

GPQA Diamond, Vals AI run (GPQA Diamond (Vals))

Vals AI’s independent run of the GPQA Diamond graduate-level science questions.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Vals GPQA Diamond mirror

BenchLM mirrors the public Vals AI Vals GPQA Diamond mirror leaderboard captured from https://www.vals.ai/benchmarks/gpqa and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals GPQA Diamond mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

138 Vals rows3 task viewspublic datasetTasks: Overall, Few-Shot CoT, Zero-Shot CoTDisplay only

Vals GPQA Diamond mirror score on GPQA Diamond (Vals) — September 1, 2026

We mirror the published vals gpqa diamond mirror score view for GPQA Diamond (Vals). Gemini 3.1 Pro Preview leads the public snapshot at 95.5%, followed by GPT-5.6 Sol (95.2%) and Grok 4.6 (94.7%). We do not use these results to rank models overall.

138 modelsKnowledgeCurrentDisplay onlyUpdated September 1, 2026

Vals GPQA Diamond mirror score table (138 models)

Score
1
Gemini 3.1 Pro PreviewGooglehigh reasoning
95.5%
2
GPT-5.6 SolOpenAI · Closedmax reasoning
95.2%
3
Grok 4.6xAI · Closedhigh reasoning
94.7%
4
Gemini 3.8 FlashGoogle · Closedhigh reasoning
94.4%
5
Gemini 3.7 FlashGoogle · Closedhigh reasoning
93.9%
6
Qwen3.8 MaxAlibaba · Open weight
93.7%
7
Gemini 3.6 FlashGoogle · Closedhigh reasoning
93.4%
8
Claude Opus 5Anthropic · Closed
93.4%
9
Claude Fable 5.1Anthropic · Closed
93.4%
10
Claude Fable 5Anthropic · Closed
93.2%
11
GPT-5.5OpenAI · Closedxhigh reasoning
93.2%
12
Kimi K3Moonshot AI · Closed
92.9%
13
Grok 4.5xAI · Closedhigh reasoning
92.9%
14
Gemini 3.5 FlashGoogle · Closedhigh reasoning
92.7%
15
MiniMax M3MiniMax · Open weight
92.7%
16
Claude Opus 4.8Anthropic · Closed
92.4%
17
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
92.4%
18
GPT-5.6 LunaOpenAI · Closedmax reasoning
91.7%
19
Gemini 3 Pro PreviewGooglehigh reasoning
91.7%
20
GPT-5.2OpenAI · Closedxhigh reasoning
91.7%
21
GPT-5.4OpenAI · Closedxhigh reasoning
91.7%
22
Grok 4.3xAI · Closed
91.4%
23
Muse Spark 1.1Meta · Closedxhigh reasoning
91.2%
24
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
90.9%
25
Qwen3.7 MaxAlibaba · Closed
90.2%
26
Claude Opus 4.7Anthropic · Closed
90.2%
27
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
89.9%
28
Muse SparkMeta · Closed
89.6%
29
Claude Opus 4.6 (Adaptive)Anthropic · Closed
89.6%
30
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
89.4%
31
Kimi K2.6Moonshot AI · Open weight
89.1%
32
Claude Sonnet 5Anthropic · Closed
88.9%
33
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
88.9%
35
88.1%
36
GLM-5.3Z.AI · Open weightmax reasoning
88.1%
37
Gemini 3 Flash PreviewGooglehigh reasoning
87.9%
38
87.4%
39
Qwen3.6 PlusAlibaba · Closed
87.4%
40
InklingThinking Machines Lab · Open weight0.99 reasoning
87.1%
41
GPT-5.1OpenAI · Closedhigh reasoning
86.6%
42
MiniMax M2.7MiniMax · Open weight
86.6%
43
GLM-5.3-FlashZ.AI · Open weightmax reasoning
86.4%
45
Claude Opus 4.5 ThinkingAnthropic · Closed
85.9%
46
GPT-5OpenAIhigh reasoning
85.6%
47
GLM-5.2Z.AI · Open weight
85.6%
48
Claude Sonnet 4.6Anthropic · Closed
85.6%
49
85.4%
51
Qwen3 MaxAlibaba · Closed
84.8%
52
GLM-5.1Z.AI · Open weight
84.5%
53
84.3%
54
o3OpenAI · Closedhigh reasoning
84.1%
55
84.1%
56
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
83.8%
57
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
83.6%
58
83.3%
59
GPT-5.4 miniOpenAI · Closedxhigh reasoning
83.1%
60
Qwen3.5 FlashAlibaba · Closed
82.8%
61
MiMo-V2.5-ProXiaomi · Closed
82.6%
62
MiniMax M2.5MiniMax · Closed
82.1%
63
Claude Sonnet 4.5 ThinkingAnthropic · Closed
81.6%
65
MiMo-V2.5Xiaomi · Closed
81.6%
66
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
81.1%
68
GPT-5 miniOpenAI · Closedhigh reasoning
80.3%
69
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
80.3%
70
GLM-4.7Z.AI · Open weight
80.0%
71
Claude Opus 4.5Anthropic · Closed
79.5%
72
Qwen3 MaxAlibaba · Closed
79.5%
73
79.3%
74
GPT-OSS 120BOpenAI · Open weight
78.5%
75
78.5%
76
Kimi K2 ThinkingMoonshot AI
78.5%
77
77.8%
78
GPT-5.4 nanoOpenAI · Closedhigh reasoning
77.5%
81
DeepSeek V3p2Fireworks AInone reasoning
76.3%
82
o3-miniOpenAI · Closedhigh reasoning
75.5%
85
O4 MiniOpenAIhigh reasoning
74.5%
86
GLM-4.6Z.AI · Open weight
74.5%
87
74.2%
88
o1OpenAI · Closedhigh reasoning
73.2%
89
73.0%
90
Claude Haiku 4.5 ThinkingAnthropic · Closed
72.2%
91
GLM-4.5Z.AI · Closed
72.2%
92
Claude Opus 4Anthropic
71.7%
93
71.5%
95
Qwen3 235b A22bFireworks AI
70.2%
96
70.0%
98
69.4%
99
GPT-OSS 20BOpenAI · Open weight
68.9%
100
68.4%
101
GPT-4.1 miniOpenAI · Closedhigh reasoning
67.9%
102
67.4%
103
GPT-4.1OpenAI · Closedhigh reasoning
65.4%
104
65.2%
107
GPT-5 nanoOpenAI · Closedhigh reasoning
63.4%
109
62.4%
111
DeepSeek V3 0324Fireworks AI
61.6%
112
Claude 3.5 SonnetAnthropic · Closed
59.3%
113
MiMo-V2-FlashXiaomi · Open weight
59.3%
114
58.3%
115
58.3%
117
Laguna XS.2Poolside · Open weight
55.0%
118
DeepSeek V3DeepSeek · Open weight
54.5%
119
GPT-4oOpenAI · Closedhigh reasoning
53.8%
120
GPT-4.1 nanoOpenAI · Closedhigh reasoning
50.8%
121
50.8%
122
50.8%
123
GPT-4oOpenAI · Closedhigh reasoning
50.3%
125
48.5%
126
47.7%
129
46.0%
130
44.2%
131
GPT-4o miniOpenAI · Closedhigh reasoning
44.2%
132
37.9%
133
36.1%
134
Mistral Medium 3.5Mistral AIhigh reasoning
34.8%
135
32.1%
136
31.1%
137
GPT-3.5 TurboOpenAIhigh reasoning
30.6%
138
Laguna M.1Poolside · Closed
27.0%

The published GPQA Diamond (Vals) snapshot places Gemini 3.1 Pro Preview first at 95.5%. The third row is 0.8 points behind. The broader top-10 range is 2.3 points, so many of the published results sit in a relatively narrow band.

138 models have been evaluated on GPQA Diamond (Vals). The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. GPQA Diamond (Vals) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GPQA Diamond (Vals)

Year

2026

Tasks

Graduate-level science questions

Format

Accuracy

Difficulty

Expert reasoning

BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).

BenchLM freshness & provenance

Version

GPQA Diamond (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GPQA Diamond (Vals) measure?

Vals AI’s independent run of the GPQA Diamond graduate-level science questions.

Which model leads the published GPQA Diamond (Vals) snapshot?

Gemini 3.1 Pro Preview currently leads the published GPQA Diamond (Vals) snapshot with 95.5% vals gpqa diamond mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on GPQA Diamond (Vals)?

The September 1, 2026 snapshot contains 138 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.