Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Artificial Analysis GPQA Diamond (AA-GPQA Diamond)

A display-only Artificial Analysis GPQA Diamond score.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on AA-GPQA Diamond — September 10, 2026

We mirror the published score view for AA-GPQA Diamond. GPT-6 Astra leads the public snapshot at 96.1%, followed by Gemini 3.8 Flash (95.3%) and Grok 4.6 (94.9%). We do not use these results to rank models overall.

200 modelsKnowledgeCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (200 models)

Score
1
GPT-6 AstraOpenAI · Closed
96.1%
2
Gemini 3.8 FlashGoogle · Closed
95.3%
3
Grok 4.6xAI · Closed
94.9%
4
Gemini 3.7 FlashGoogle · Closed
94.5%
5
GPT-5.6 SolOpenAI · Closed
94.1%
6
Gemini 3.1 ProGoogle · Closed
94.1%
7
Claude Fable 5.1Anthropic · Closed
93.7%
8
Kimi K3Moonshot AI · Closed
93.5%
9
GPT-5.5OpenAI · Closed
93.5%
10
Muse Spark 1.3Meta · Closed
93.5%
11
Claude Opus 5Anthropic · Closed
93.2%
12
Grok 4.5xAI · Closed
93.1%
13
MiniMax M3MiniMax · Open weight
92.9%
14
DeepSeek V4 Pro 0813DeepSeek · Closed
92.8%
15
Gemini 3.6 FlashGoogle · Closed
92.8%
16
Qwen3.8 Max PreviewAlibaba · Closed
92.7%
17
Claude Fable 5Anthropic · Closed
92.6%
18
GPT-5.6 TerraOpenAI · Closed
92.5%
19
Qwen3.7 MaxAlibaba · Closed
92.3%
20
Qwen3.8-Flash-NextAlibaba · Open weight
92.3%
21
Gemini 3.5 FlashGoogle · Closed
92.2%
22
Claude Opus 4.8Anthropic · Closed
92.0%
23
GPT-5.4OpenAI · Closed
92.0%
24
GLM-5.3Z.AI · Open weight
91.7%
25
GPT-5.3 CodexOpenAI · Closed
91.5%
26
GPT-5.3-Codex-SparkOpenAI · Closed
91.5%
27
Claude Opus 4.7 (Adaptive)Anthropic · Closed
91.4%
28
GLM-5.3-FlashZ.AI · Open weight
91.2%
29
Claude Sonnet 5Anthropic · Closed
91.1%
30
GPT-5.6 LunaOpenAI · Closed
91.1%
31
Kimi K2.6Moonshot AI · Open weight
91.1%
32
Gemini 3 ProGoogle · Closed
90.8%
33
DeepSeek V4 Flash 0731DeepSeek · Closed
90.8%
34
Qwen3.8-27BAlibaba · Open weight
90.5%
35
DeepSeek V4 Pro (High)DeepSeek · Open weight
90.5%
36
Muse Spark 1.2Meta · Closed
90.4%
37
GPT-5.2OpenAI · Closed
90.3%
38
Grok 4.3xAI · Closed
90.1%
39
Qwen3.7 PlusAlibaba · Closed
90.0%
40
GPT-5.2-CodexOpenAI · Closed
89.9%
41
Muse Spark 1.1Meta · Closed
89.8%
42
Hy3Tencent · Open weight
89.7%
43
Hy3 PreviewTencent · Open weight
89.7%
44
Kimi K2.7 CodeMoonshot AI · Open weight
89.6%
45
Claude Opus 4.6 (Adaptive)Anthropic · Closed
89.6%
46
GLM-5.2Z.AI · Open weight
89.5%
47
Inkling-SmallThinking Machines Lab · Open weight
89.5%
48
Solar Pro 4Upstage · Closed
89.1%
49
Qwen 3.6 Max (preview)Alibaba · Closed
88.8%
50
Claude Opus 4.7Anthropic · Closed
88.5%
51
Muse SparkMeta · Closed
88.4%
52
Qwen3.6 PlusAlibaba · Closed
88.2%
53
Kimi K2.5Moonshot AI · Open weight
87.9%
54
Kimi K2.5 (Reasoning)Moonshot AI · Closed
87.9%
55
Grok 4xAI · Closed
87.7%
56
GPT-5.4 miniOpenAI · Closed
87.5%
57
MiniMax M2.7MiniMax · Open weight
87.4%
58
GPT-5.1OpenAI · Closed
87.3%
59
InklingThinking Machines Lab · Open weight
87.2%
60
MiMo-V2-ProXiaomi · Closed
87.0%
61
GLM-5.1Z.AI · Open weight
86.8%
62
Nemotron 3 UltraNVIDIA · Open weight
86.7%
63
Claude Opus 4.5 ThinkingAnthropic · Closed
86.6%
64
MiMo-V2.5-ProXiaomi · Closed
86.6%
65
Apodex 1.1Apodex · Closed
86.4%
66
Apodex 1.1 MiniApodex · Open weight
86.4%
67
Ling 3.0 Flash VLInclusionAI · Open weight
86.2%
68
Qwen3.5 397BAlibaba · Open weight
86.1%
69
Qwen3.5 397B (Reasoning)Alibaba · Open weight
86.1%
70
GPT-5.1-CodexOpenAI · Closed
86.0%
71
GPT-5.1-Codex-MaxOpenAI · Closed
86.0%
72
GLM-4.7Z.AI · Open weight
85.9%
73
Qwen3.5-27BAlibaba · Open weight
85.8%
74
Gemma 4 31BGoogle · Open weight
85.7%
75
Qwen3.5-122B-A10BAlibaba · Open weight
85.7%
76
A.X K2SK Telecom · Open weight
85.7%
77
Ling 3.0 FlashInclusionAI · Open weight
85.5%
78
Ling 3.0 Flash FP8InclusionAI · Open weight
85.5%
79
GPT-5 (high)OpenAI · Closed
85.4%
80
85.3%
81
MiniMax M2.5MiniMax · Closed
84.8%
82
GLM-5-TurboZ.AI · Closed
84.7%
83
84.7%
84
Qwen3.5-35B-A3BAlibaba · Open weight
84.5%
85
o3-proOpenAI · Closed
84.5%
86
Gemini 2.5 ProGoogle · Closed
84.4%
87
Qwen3.6-27BAlibaba · Open weight
84.2%
88
GPT-5 (medium)OpenAI · Closed
84.2%
89
Qwen3.6-35B-A3BAlibaba · Open weight
84.1%
90
Claude Opus 4.6Anthropic · Closed
84.0%
91
Gemini 3.5 Flash-LiteGoogle · Closed
83.8%
92
Muse Glimmer 30BMeta · Open weight
83.5%
93
K-EXAONE 2.0LG AI Research · Open weight
82.9%
94
MiMo-V2-OmniXiaomi · Closed
82.8%
95
GPT-5 miniOpenAI · Closed
82.8%
96
o3OpenAI · Closed
82.7%
97
Step 3.5 FlashStepFun · Open weight
82.6%
98
GLM-5Z.AI · Open weight
82.0%
99
GPT-5.4 nanoOpenAI · Closed
81.7%
100
DeepSeek-R1DeepSeek · Open weight
81.3%
101
Gemini 3 FlashGoogle · Closed
81.2%
102
Claude Opus 4.5Anthropic · Closed
81.0%
103
Claude 4.1 Opus ThinkingAnthropic · Closed
80.9%
104
GLM-5V-TurboZ.AI · Closed
80.9%
105
Step 3.7 FlashStepFun · Open weight
80.9%
106
Nemotron 3 Super 100BNVIDIA · Open weight
80.0%
107
Nemotron 3 Super 120B A12BNVIDIA · Open weight
80.0%
108
Claude Sonnet 4.6Anthropic · Closed
79.9%
109
Gemma 4 26B A4BGoogle · Open weight
79.2%
110
K-ExaoneLG AI Research · Closed
78.3%
111
GPT-OSS 120BOpenAI · Open weight
78.2%
112
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
77.9%
113
Mercury 2Inception · Closed
77.0%
114
Mistral Small 4Mistral · Open weight
76.9%
115
Mistral Small 4 (Reasoning)Mistral · Open weight
76.9%
116
Kimi K2Moonshot AI · Closed
76.6%
117
o1-previewOpenAI · Closed
76.5%
118
Qwen3 MaxAlibaba · Closed
76.4%
119
Command A+Cohere · Open weight
76.1%
120
Nemotron 3 Nano 30BNVIDIA · Open weight
75.7%
121
North Mini CodeCohere · Open weight
75.7%
122
Gemma 4 12BGoogle · Open weight
75.3%
123
Trinity-Large-ThinkingArcee AI · Open weight
75.2%
124
Trinity-Large-PreviewArcee AI · Open weight
75.2%
125
DeepSeek V3.2DeepSeek · Open weight
75.1%
126
Mistral Medium 3.5 128BMistral · Open weight
74.8%
127
o3-miniOpenAI · Closed
74.8%
128
o1OpenAI · Closed
74.7%
129
74.3%
130
Sarvam 105BSarvam · Open weight
73.8%
131
DeepSeek V3.1DeepSeek · Open weight
73.5%
132
Ling 3.0 TinyInclusionAI · Open weight
73.4%
133
GLM-4.5-AirZ.AI · Closed
73.3%
134
Quasar 438BMultiverse Computing · Closed
73.2%
135
Nemotron Ultra 253BNVIDIA · Open weight
72.8%
136
Grok Code Fast 1xAI · Closed
72.7%
137
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weight
72.6%
138
Solar Pro 3Upstage · Closed
72.4%
139
MiniCPM5-2BOpenBMB · Open weight
70.2%
140
MiniMax M1 80kMiniMax · Closed
69.7%
141
GPT-OSS 20BOpenAI · Open weight
68.8%
142
Claude 4 SonnetAnthropic · Closed
68.3%
143
Gemini 2.5 FlashGoogle · Closed
68.3%
144
Mistral Large 3Mistral · Closed
68.0%
145
GPT-5 nanoOpenAI · Closed
67.6%
146
Llama 4 MaverickMeta · Open weight
67.1%
147
GPT-4.1OpenAI · Closed
66.6%
148
GPT-4.1 miniOpenAI · Closed
66.4%
149
MiMo-V2-FlashXiaomi · Open weight
65.6%
150
Granite 4.2 30BIBM · Open weight
64.4%
151
Grok 4.1 FastxAI · Closed
63.7%
152
Sarvam 30BSarvam · Open weight
63.3%
153
GLM-4.6Z.AI · Open weight
63.2%
154
Celeris-1Celeris · Closed
63.1%
155
Granite 4.2 8BIBM · Open weight
63.1%
156
Exaone 4.0 32BLG AI Research · Open weight
62.8%
157
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weight
62.0%
158
DeepSeek R1 Distill Qwen 32BDeepSeek · Open weight
61.5%
159
Ling 2.6 FlashInclusionAI · Open weight
59.3%
160
Gemini 1.5 ProGoogle · Closed
58.9%
161
Llama 4 ScoutMeta · Open weight
58.7%
162
GLM-4.7-FlashZ.AI · Open weight
58.1%
163
Mistral Medium 3Mistral · Closed
57.8%
164
Gemma 4 E4BGoogle · Open weight
57.6%
165
Phi-4Microsoft · Open weight
57.5%
166
Ministral 3 14B (Reasoning)Mistral · Open weight
57.2%
167
Ministral 3 14BMistral · Open weight
57.2%
168
Solar Pro 2Upstage · Closed
56.1%
169
Granite 4.2 3BIBM · Open weight
55.9%
170
LFM2.5-2.6BLiquidAI · Open weight
55.8%
171
DeepSeek V3DeepSeek · Open weight
55.7%
172
GPT-4oOpenAI · Closed
54.3%
173
Llama 3.1 405BMeta · Open weight
51.5%
174
LFM2.5-8B-A1BLiquidAI · Open weight
51.3%
175
GPT-4.1 nanoOpenAI · Closed
51.2%
176
Nova ProAmazon · Closed
49.9%
177
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
49.8%
178
Claude 3 OpusAnthropic · Closed
48.9%
179
Mistral Large 2Mistral · Closed
48.6%
180
LFM2-24B-A2BLiquidAI · Closed
47.4%
181
Ministral 3 8B (Reasoning)Mistral · Open weight
47.1%
182
Ministral 3 8BMistral · Open weight
47.1%
183
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
46.9%
184
Gemma 4 E2BGoogle · Open weight
43.3%
185
Gemma 3 27BGoogle · Open weight
42.8%
186
GPT-4o miniOpenAI · Closed
42.6%
187
Exaone 4.0 1.2BLG AI Research · Open weight
42.4%
188
Qwen2.5 Coder 32B InstructAlibaba · Open weight
41.7%
189
Claude 3 HaikuAnthropic · Closed
37.4%
190
Ministral 3 3B (Reasoning)Mistral · Open weight
35.8%
191
Ministral 3 3BMistral · Open weight
35.8%
192
LFM2.5-1.2B-ThinkingLiquidAI · Closed
33.9%
193
LFM2.5-1.2B-InstructLiquidAI · Closed
32.6%
194
Phi-4 Multimodal InstructMicrosoft · Open weight
31.5%
195
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
28.9%
196
Granite-4.0-1BIBM · Open weight
28.1%
197
Gemini 1.0 ProGoogle · Closed
27.7%
198
Granite-4.0-H-1BIBM · Open weight
26.3%
199
Granite-4.0-350MIBM · Open weight
26.1%
200
Granite-4.0-H-350MIBM · Open weight
25.7%

The published AA-GPQA Diamond snapshot places GPT-6 Astra first at 96.1%. The third row is 1.2 points behind. The broader top-10 range is 2.6 points, so many of the published results sit in a relatively narrow band.

200 models have been evaluated on AA-GPQA Diamond. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-GPQA Diamond is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-GPQA Diamond

Year

2026

Tasks

Graduate-level science questions

Format

Accuracy

Difficulty

Graduate-level science reasoning

BenchLM stores the Artificial Analysis GPQA Diamond result separately from the weighted GPQA lane so AA refreshes remain display-only.

BenchLM freshness & provenance

Version

AA-GPQA Diamond 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-GPQA Diamond measure?

A display-only Artificial Analysis GPQA Diamond score.

Which model scores highest on AA-GPQA Diamond?

GPT-6 Astra by OpenAI currently leads with a score of 96.1% on AA-GPQA Diamond.

How many models are evaluated on AA-GPQA Diamond?

200 AI models have been evaluated on AA-GPQA Diamond on BenchLM.

Last updated: September 10, 2026 · BenchLM version AA-GPQA Diamond 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.