Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Artificial Analysis GPQA Diamond (AA-GPQA Diamond)

A display-only Artificial Analysis GPQA Diamond score.

Data verified 24 confirmed releases in the last 30 daysStart free brief

Benchmark score on AA-GPQA Diamond — August 22, 2026

We mirror the published score view for AA-GPQA Diamond. Grok 4.6 leads the public snapshot at 94.9%, followed by Gemini 3.7 Flash (94.5%) and GPT-5.6 Sol (94.1%). We do not use these results to rank models overall.

177 modelsKnowledgeCurrentDisplay onlyUpdated August 22, 2026

Benchmark score table (177 models)

Score
1
Grok 4.6xAI · Closed
94.9%
2
Gemini 3.7 FlashGoogle · Closed
94.5%
3
GPT-5.6 SolOpenAI · Closed
94.1%
4
Gemini 3.1 ProGoogle · Closed
94.1%
5
Kimi K3Moonshot AI · Closed
93.5%
6
GPT-5.5OpenAI · Closed
93.5%
7
Claude Opus 5Anthropic · Closed
93.2%
8
Grok 4.5xAI · Closed
93.1%
9
MiniMax M3MiniMax · Open weight
92.9%
10
DeepSeek V4 Pro 0813DeepSeek · Closed
92.8%
11
Gemini 3.6 FlashGoogle · Closed
92.8%
12
Qwen3.8 Max PreviewAlibaba · Closed
92.7%
13
Claude Fable 5Anthropic · Closed
92.6%
14
GPT-5.6 TerraOpenAI · Closed
92.5%
15
Qwen3.7 MaxAlibaba · Closed
92.3%
16
Gemini 3.5 FlashGoogle · Closed
92.2%
17
Claude Opus 4.8Anthropic · Closed
92.0%
18
GPT-5.4OpenAI · Closed
92.0%
19
GLM-5.3Z.AI · Closed
91.7%
20
GPT-5.3 CodexOpenAI · Closed
91.5%
21
GPT-5.3-Codex-SparkOpenAI · Closed
91.5%
22
Claude Opus 4.7 (Adaptive)Anthropic · Closed
91.4%
23
Claude Sonnet 5Anthropic · Closed
91.1%
24
GPT-5.6 LunaOpenAI · Closed
91.1%
25
Kimi K2.6Moonshot AI · Open weight
91.1%
26
Gemini 3 ProGoogle · Closed
90.8%
27
DeepSeek V4 Flash 0731DeepSeek · Closed
90.8%
28
Qwen3.8-27BAlibaba · Open weight
90.5%
29
DeepSeek V4 Pro (High)DeepSeek · Open weight
90.5%
30
Muse Spark 1.2Meta · Closed
90.4%
31
GPT-5.2OpenAI · Closed
90.3%
32
Grok 4.3xAI · Closed
90.1%
33
Qwen3.7 PlusAlibaba · Closed
90.0%
34
GPT-5.2-CodexOpenAI · Closed
89.9%
35
Muse Spark 1.1Meta · Closed
89.8%
36
Hy3 PreviewTencent · Open weight
89.7%
37
Hy3Tencent · Open weight
89.7%
38
Claude Opus 4.6 (Adaptive)Anthropic · Closed
89.6%
39
Kimi K2.7 CodeMoonshot AI · Open weight
89.6%
40
GLM-5.2Z.AI · Open weight
89.5%
41
Inkling-SmallThinking Machines Lab · Open weight
89.5%
42
Qwen3.5 397BAlibaba · Open weight
89.3%
43
Qwen3.5 397B (Reasoning)Alibaba · Open weight
89.3%
44
Qwen 3.6 Max (preview)Alibaba · Closed
88.8%
45
Claude Opus 4.7Anthropic · Closed
88.5%
46
Muse SparkMeta · Closed
88.4%
47
Qwen3.6 PlusAlibaba · Closed
88.2%
48
Kimi K2.5Moonshot AI · Open weight
87.9%
49
Kimi K2.5 (Reasoning)Moonshot AI · Closed
87.9%
50
Grok 4xAI · Closed
87.7%
51
GPT-5.4 miniOpenAI · Closed
87.5%
52
MiniMax M2.7MiniMax · Open weight
87.4%
53
GPT-5.1OpenAI · Closed
87.3%
54
InklingThinking Machines Lab · Open weight
87.2%
55
MiMo-V2-ProXiaomi · Closed
87.0%
56
GLM-5.1Z.AI · Open weight
86.8%
57
Nemotron 3 UltraNVIDIA · Open weight
86.7%
58
MiMo-V2.5-ProXiaomi · Closed
86.6%
59
Claude Opus 4.5 ThinkingAnthropic · Closed
86.6%
60
GPT-5.1-Codex-MaxOpenAI · Closed
86.0%
61
GPT-5.1-CodexOpenAI · Closed
86.0%
62
GLM-4.7Z.AI · Open weight
85.9%
63
Qwen3.5-27BAlibaba · Open weight
85.8%
64
Qwen3.5-122B-A10BAlibaba · Open weight
85.7%
65
Gemma 4 31BGoogle · Open weight
85.7%
66
Ling 3.0 FlashInclusionAI · Open weight
85.5%
67
Ling 3.0 Flash FP8InclusionAI · Open weight
85.5%
68
GPT-5 (high)OpenAI · Closed
85.4%
69
85.3%
70
MiniMax M2.5MiniMax · Closed
84.8%
71
GLM-5-TurboZ.AI · Closed
84.7%
72
84.7%
73
Qwen3.5-35B-A3BAlibaba · Open weight
84.5%
74
o3-proOpenAI · Closed
84.5%
75
Gemini 2.5 ProGoogle · Closed
84.4%
76
Qwen3.6-27BAlibaba · Open weight
84.2%
77
GPT-5 (medium)OpenAI · Closed
84.2%
78
Qwen3.6-35B-A3BAlibaba · Open weight
84.1%
79
Claude Opus 4.6Anthropic · Closed
84.0%
80
Gemini 3.5 Flash-LiteGoogle · Closed
83.8%
81
Muse Glimmer 30BMeta · Open weight
83.5%
82
MiMo-V2-OmniXiaomi · Closed
82.8%
83
GPT-5 miniOpenAI · Closed
82.8%
84
o3OpenAI · Closed
82.7%
85
Step 3.5 FlashStepFun · Open weight
82.6%
86
GLM-5Z.AI · Open weight
82.0%
87
GPT-5.4 nanoOpenAI · Closed
81.7%
88
DeepSeek-R1DeepSeek · Open weight
81.3%
89
Gemini 3 FlashGoogle · Closed
81.2%
90
Claude Opus 4.5Anthropic · Closed
81.0%
91
Step 3.7 FlashStepFun · Open weight
80.9%
92
GLM-5V-TurboZ.AI · Closed
80.9%
93
Claude 4.1 Opus ThinkingAnthropic · Closed
80.9%
94
Nemotron 3 Super 100BNVIDIA · Open weight
80.0%
95
Nemotron 3 Super 120B A12BNVIDIA · Open weight
80.0%
96
Claude Sonnet 4.6Anthropic · Closed
79.9%
97
Gemma 4 26B A4BGoogle · Open weight
79.2%
98
K-ExaoneLG AI Research · Closed
78.3%
99
GPT-OSS 120BOpenAI · Open weight
78.2%
100
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
77.9%
101
Mercury 2Inception · Closed
77.0%
102
Mistral Small 4Mistral · Open weight
76.9%
103
Mistral Small 4 (Reasoning)Mistral · Open weight
76.9%
104
Kimi K2Moonshot AI · Closed
76.6%
105
o1-previewOpenAI · Closed
76.5%
106
Qwen3 MaxAlibaba · Closed
76.4%
107
Command A+Cohere · Open weight
76.1%
108
Nemotron 3 Nano 30BNVIDIA · Open weight
75.7%
109
Gemma 4 12BGoogle · Open weight
75.3%
110
Trinity-Large-PreviewArcee AI · Open weight
75.2%
111
Trinity-Large-ThinkingArcee AI · Open weight
75.2%
112
DeepSeek V3.2DeepSeek · Open weight
75.1%
113
Mistral Medium 3.5 128BMistral · Open weight
74.8%
114
o3-miniOpenAI · Closed
74.8%
115
o1OpenAI · Closed
74.7%
116
74.3%
117
Sarvam 105BSarvam · Open weight
73.8%
118
DeepSeek V3.1DeepSeek · Open weight
73.5%
119
GLM-4.5-AirZ.AI · Closed
73.3%
120
Nemotron Ultra 253BNVIDIA · Open weight
72.8%
121
Grok Code Fast 1xAI · Closed
72.7%
122
MiniMax M1 80kMiniMax · Closed
69.7%
123
GPT-OSS 20BOpenAI · Open weight
68.8%
124
Claude 4 SonnetAnthropic · Closed
68.3%
125
Gemini 2.5 FlashGoogle · Closed
68.3%
126
Mistral Large 3Mistral · Closed
68.0%
127
GPT-5 nanoOpenAI · Closed
67.6%
128
Llama 4 MaverickMeta · Open weight
67.1%
129
GPT-4.1OpenAI · Closed
66.6%
130
GPT-4.1 miniOpenAI · Closed
66.4%
131
MiMo-V2-FlashXiaomi · Open weight
65.6%
132
Grok 4.1 FastxAI · Closed
63.7%
133
Sarvam 30BSarvam · Open weight
63.3%
134
GLM-4.6Z.AI · Open weight
63.2%
135
Celeris-1Celeris · Closed
63.1%
136
Exaone 4.0 32BLG AI Research · Open weight
62.8%
137
DeepSeek R1 Distill Qwen 32BDeepSeek · Open weight
61.5%
138
Ling 2.6 FlashInclusionAI · Open weight
59.3%
139
Gemini 1.5 ProGoogle · Closed
58.9%
140
Llama 4 ScoutMeta · Open weight
58.7%
141
GLM-4.7-FlashZ.AI · Open weight
58.1%
142
Mistral Medium 3Mistral · Closed
57.8%
143
Gemma 4 E4BGoogle · Open weight
57.6%
144
Phi-4Microsoft · Open weight
57.5%
145
Ministral 3 14B (Reasoning)Mistral · Open weight
57.2%
146
Ministral 3 14BMistral · Open weight
57.2%
147
Solar Pro 2Upstage · Closed
56.1%
148
DeepSeek V3DeepSeek · Open weight
55.7%
149
GPT-4oOpenAI · Closed
54.3%
150
Llama 3.1 405BMeta · Open weight
51.5%
151
LFM2.5-8B-A1BLiquidAI · Open weight
51.3%
152
GPT-4.1 nanoOpenAI · Closed
51.2%
153
Nova ProAmazon · Closed
49.9%
154
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
49.8%
155
Claude 3 OpusAnthropic · Closed
48.9%
156
Mistral Large 2Mistral · Closed
48.6%
157
LFM2-24B-A2BLiquidAI · Closed
47.4%
158
Ministral 3 8B (Reasoning)Mistral · Open weight
47.1%
159
Ministral 3 8BMistral · Open weight
47.1%
160
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
46.9%
161
Gemma 4 E2BGoogle · Open weight
43.3%
162
Gemma 3 27BGoogle · Open weight
42.8%
163
GPT-4o miniOpenAI · Closed
42.6%
164
Exaone 4.0 1.2BLG AI Research · Open weight
42.4%
165
Qwen2.5 Coder 32B InstructAlibaba · Open weight
41.7%
166
Claude 3 HaikuAnthropic · Closed
37.4%
167
Ministral 3 3B (Reasoning)Mistral · Open weight
35.8%
168
Ministral 3 3BMistral · Open weight
35.8%
169
LFM2.5-1.2B-ThinkingLiquidAI · Closed
33.9%
170
LFM2.5-1.2B-InstructLiquidAI · Closed
32.6%
171
Phi-4 Multimodal InstructMicrosoft · Open weight
31.5%
172
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
28.9%
173
Granite-4.0-1BIBM · Open weight
28.1%
174
Gemini 1.0 ProGoogle · Closed
27.7%
175
Granite-4.0-H-1BIBM · Open weight
26.3%
176
Granite-4.0-350MIBM · Open weight
26.1%
177
Granite-4.0-H-350MIBM · Open weight
25.7%

The published AA-GPQA Diamond snapshot places Grok 4.6 first at 94.9%. The third row is 0.8 points behind. The broader top-10 range is 2.1 points, so many of the published results sit in a relatively narrow band.

177 models have been evaluated on AA-GPQA Diamond. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-GPQA Diamond is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-GPQA Diamond

Year

2026

Tasks

Graduate-level science questions

Format

Accuracy

Difficulty

Graduate-level science reasoning

BenchLM stores the Artificial Analysis GPQA Diamond result separately from the weighted GPQA lane so AA refreshes remain display-only.

BenchLM freshness & provenance

Version

AA-GPQA Diamond 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-GPQA Diamond measure?

A display-only Artificial Analysis GPQA Diamond score.

Which model scores highest on AA-GPQA Diamond?

Grok 4.6 by xAI currently leads with a score of 94.9% on AA-GPQA Diamond.

How many models are evaluated on AA-GPQA Diamond?

177 AI models have been evaluated on AA-GPQA Diamond on BenchLM.

Last updated: August 22, 2026 · BenchLM version AA-GPQA Diamond 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.