Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Artificial Analysis Humanity's Last Exam (AA-HLE)

A display-only Artificial Analysis Humanity's Last Exam score.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on AA-HLE — September 10, 2026

We mirror the published score view for AA-HLE. Claude Fable 5.1 leads the public snapshot at 59.1%, followed by Claude Fable 5 (55.5%) and Claude Opus 5 (54.9%). We do not use these results to rank models overall.

200 modelsKnowledgeCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (200 models)

Score
1
Claude Fable 5.1Anthropic · Closed
59.1%
2
Claude Fable 5Anthropic · Closed
55.5%
3
Claude Opus 5Anthropic · Closed
54.9%
4
GPT-6 AstraOpenAI · Closed
54.7%
5
GPT-5.6 SolOpenAI · Closed
49.5%
6
Claude Opus 4.8Anthropic · Closed
48.7%
7
Muse Spark 1.3Meta · Closed
48.7%
8
Gemini 3.7 FlashGoogle · Closed
47.9%
9
Gemini 3.8 FlashGoogle · Closed
47.8%
10
Gemini 3.1 ProGoogle · Closed
47.0%
11
Kimi K3Moonshot AI · Closed
46.9%
12
Muse Spark 1.1Meta · Closed
46.2%
13
GPT-5.5OpenAI · Closed
45.8%
14
Muse Spark 1.2Meta · Closed
45.5%
15
GPT-5.4OpenAI · Closed
43.7%
16
Qwen3.8 Max PreviewAlibaba · Closed
43.0%
17
GPT-5.6 TerraOpenAI · Closed
42.9%
18
Grok 4.6xAI · Closed
42.9%
19
Gemini 3.5 FlashGoogle · Closed
42.7%
20
Grok 4.5xAI · Closed
42.7%
21
GPT-5.3 CodexOpenAI · Closed
42.5%
22
GPT-5.3-Codex-SparkOpenAI · Closed
42.5%
23
Claude Opus 4.7 (Adaptive)Anthropic · Closed
42.3%
24
GLM-5.3Z.AI · Open weight
42.3%
25
Claude Sonnet 5Anthropic · Closed
41.3%
26
GLM-5.2Z.AI · Open weight
41.1%
27
DeepSeek V4 Pro 0813DeepSeek · Closed
41.0%
28
Gemini 3.6 FlashGoogle · Closed
40.8%
29
Muse SparkMeta · Closed
40.7%
30
Qwen3.7 MaxAlibaba · Closed
40.5%
31
GLM-5.3-FlashZ.AI · Open weight
39.9%
32
Claude Opus 4.6 (Adaptive)Anthropic · Closed
39.9%
33
Gemini 3 ProGoogle · Closed
39.7%
34
GPT-5.6 LunaOpenAI · Closed
39.5%
35
DeepSeek V4.1 FlashDeepSeek · Open weight
39.2%
36
MiniMax M3MiniMax · Open weight
39.0%
37
DeepSeek V4 Flash 0731DeepSeek · Closed
38.6%
38
Qwen3.8-Flash-NextAlibaba · Open weight
38.0%
39
GPT-5.2OpenAI · Closed
37.7%
40
Kimi K2.6Moonshot AI · Open weight
37.5%
41
Grok 4.3xAI · Closed
37.2%
42
GPT-5.2-CodexOpenAI · Closed
35.7%
43
MiMo-V2.5-ProXiaomi · Closed
35.7%
44
Qwen3.7 PlusAlibaba · Closed
35.6%
45
DeepSeek V4 Pro (High)DeepSeek · Open weight
35.2%
46
Kimi K2.7 CodeMoonshot AI · Open weight
35.0%
47
Apodex 1.1Apodex · Closed
34.1%
48
Apodex 1.1 MiniApodex · Open weight
34.1%
49
Qwen3.8-27BAlibaba · Open weight
33.9%
50
Hy3Tencent · Open weight
33.5%
51
Hy3 PreviewTencent · Open weight
33.5%
52
Inkling-SmallThinking Machines Lab · Open weight
33.3%
53
Claude Opus 4.7Anthropic · Closed
33.3%
54
InklingThinking Machines Lab · Open weight
31.9%
55
Qwen 3.6 Max (preview)Alibaba · Closed
30.8%
56
Kimi K2.5Moonshot AI · Open weight
30.7%
57
Kimi K2.5 (Reasoning)Moonshot AI · Closed
30.7%
58
MiMo-V2-ProXiaomi · Closed
30.4%
59
GLM-5.1Z.AI · Open weight
30.1%
60
Claude Opus 4.5 ThinkingAnthropic · Closed
30.1%
61
MiniMax M2.7MiniMax · Open weight
29.6%
62
A.X K2SK Telecom · Open weight
29.6%
63
GLM-5Z.AI · Open weight
29.3%
64
Solar Pro 4Upstage · Closed
29.2%
65
GPT-5.1OpenAI · Closed
28.5%
66
GPT-5 (high)OpenAI · Closed
28.5%
67
Nemotron 3 UltraNVIDIA · Open weight
28.4%
68
GPT-5.4 nanoOpenAI · Closed
28.3%
69
GPT-5.4 miniOpenAI · Closed
28.1%
70
Qwen3.6 PlusAlibaba · Closed
27.8%
71
GLM-5-TurboZ.AI · Closed
27.8%
72
GLM-4.7Z.AI · Open weight
27.4%
73
Grok 4xAI · Closed
26.7%
74
GPT-5.1-CodexOpenAI · Closed
25.7%
75
GPT-5.1-Codex-MaxOpenAI · Closed
25.7%
76
GPT-5 (medium)OpenAI · Closed
25.4%
77
Qwen3.5-122B-A10BAlibaba · Open weight
25.2%
78
Step 3.5 FlashStepFun · Open weight
24.5%
79
Qwen3.5-27BAlibaba · Open weight
23.9%
80
Ling 3.0 FlashInclusionAI · Open weight
23.7%
81
Ling 3.0 Flash FP8InclusionAI · Open weight
23.7%
82
Gemma 4 31BGoogle · Open weight
23.6%
83
Qwen3.6-27BAlibaba · Open weight
23.1%
84
Gemini 2.5 ProGoogle · Closed
22.5%
85
Qwen3.6-35B-A3BAlibaba · Open weight
22.2%
86
MiMo-V2-OmniXiaomi · Closed
22.1%
87
Muse Glimmer 30BMeta · Open weight
22.0%
88
Ling 3.0 Flash VLInclusionAI · Open weight
22.0%
89
GPT-5 miniOpenAI · Closed
21.5%
90
Step 3.7 FlashStepFun · Open weight
21.4%
91
Qwen3.5-35B-A3BAlibaba · Open weight
21.0%
92
Nemotron 3 Super 100BNVIDIA · Open weight
20.8%
93
Nemotron 3 Super 120B A12BNVIDIA · Open weight
20.8%
94
MiniMax M2.5MiniMax · Closed
20.5%
95
o3OpenAI · Closed
20.1%
96
Qwen3.5 397BAlibaba · Open weight
19.8%
97
Qwen3.5 397B (Reasoning)Alibaba · Open weight
19.8%
98
GPT-OSS 120BOpenAI · Open weight
19.6%
99
Gemma 4 26B A4BGoogle · Open weight
19.3%
100
19.3%
101
Claude Opus 4.6Anthropic · Closed
19.1%
102
19.1%
103
Gemini 3.5 Flash-LiteGoogle · Closed
18.8%
104
Quasar 438BMultiverse Computing · Closed
18.7%
105
K-EXAONE 2.0LG AI Research · Open weight
18.6%
106
GLM-5V-TurboZ.AI · Closed
17.1%
107
Mercury 2Inception · Closed
17.1%
108
DeepSeek-R1DeepSeek · Open weight
15.8%
109
Trinity-Large-ThinkingArcee AI · Open weight
15.8%
110
Trinity-Large-PreviewArcee AI · Open weight
15.8%
111
Gemma 4 12BGoogle · Open weight
15.7%
112
Gemini 3 FlashGoogle · Closed
15.0%
113
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
14.3%
114
K-ExaoneLG AI Research · Closed
13.9%
115
Mistral Medium 3.5 128BMistral · Open weight
13.8%
116
Claude Sonnet 4.6Anthropic · Closed
13.3%
117
Claude Opus 4.5Anthropic · Closed
13.2%
118
Claude 4.1 Opus ThinkingAnthropic · Closed
12.5%
119
Command A+Cohere · Open weight
12.0%
120
Qwen3 MaxAlibaba · Closed
11.9%
121
Nemotron 3 Nano 30BNVIDIA · Open weight
11.4%
122
DeepSeek V3.2DeepSeek · Open weight
11.2%
123
Granite 4.2 30BIBM · Open weight
11.2%
124
North Mini CodeCohere · Open weight
11.1%
125
GPT-OSS 20BOpenAI · Open weight
11.0%
126
Sarvam 105BSarvam · Open weight
11.0%
127
10.6%
128
Solar Pro 3Upstage · Closed
10.3%
129
Mistral Small 4Mistral · Open weight
9.9%
130
Mistral Small 4 (Reasoning)Mistral · Open weight
9.9%
131
Granite 4.2 8BIBM · Open weight
9.7%
132
GPT-5 nanoOpenAI · Closed
9.5%
133
Ling 3.0 TinyInclusionAI · Open weight
9.3%
134
MiniCPM5-2BOpenBMB · Open weight
8.9%
135
MiniMax M1 80kMiniMax · Closed
8.9%
136
MiMo-V2-FlashXiaomi · Open weight
8.6%
137
Grok Code Fast 1xAI · Closed
8.0%
138
o3-miniOpenAI · Closed
7.9%
139
GLM-4.7-FlashZ.AI · Open weight
7.6%
140
Sarvam 30BSarvam · Open weight
7.5%
141
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weight
7.5%
142
Kimi K2Moonshot AI · Closed
7.4%
143
Nemotron Ultra 253BNVIDIA · Open weight
7.4%
144
o1OpenAI · Closed
7.0%
145
GLM-4.5-AirZ.AI · Closed
7.0%
146
LFM2.5-8B-A1BLiquidAI · Open weight
6.9%
147
Celeris-1Celeris · Closed
6.8%
148
DeepSeek V3.1DeepSeek · Open weight
6.7%
149
LFM2.5-1.2B-InstructLiquidAI · Closed
6.7%
150
Granite 4.2 3BIBM · Open weight
6.6%
151
Granite-4.0-H-350MIBM · Open weight
6.4%
152
Ling 2.6 FlashInclusionAI · Open weight
6.3%
153
LFM2.5-2.6BLiquidAI · Open weight
6.2%
154
LFM2.5-1.2B-ThinkingLiquidAI · Closed
6.2%
155
Exaone 4.0 1.2BLG AI Research · Open weight
5.7%
156
GLM-4.6Z.AI · Open weight
5.5%
157
Granite-4.0-350MIBM · Open weight
5.5%
158
Ministral 3 3B (Reasoning)Mistral · Open weight
5.4%
159
Ministral 3 3BMistral · Open weight
5.4%
160
Grok 4.1 FastxAI · Closed
5.1%
161
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
5.1%
162
Exaone 4.0 32BLG AI Research · Open weight
5.0%
163
GPT-4.1 miniOpenAI · Closed
5.0%
164
Granite-4.0-H-1BIBM · Open weight
5.0%
165
Phi-4 Multimodal InstructMicrosoft · Open weight
5.0%
166
Llama 4 MaverickMeta · Open weight
4.9%
167
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
4.8%
168
Gemma 4 E2BGoogle · Open weight
4.8%
169
Granite-4.0-1BIBM · Open weight
4.8%
170
Gemini 2.5 FlashGoogle · Closed
4.7%
171
Gemini 1.5 ProGoogle · Closed
4.6%
172
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weight
4.6%
173
DeepSeek R1 Distill Qwen 32BDeepSeek · Open weight
4.6%
174
Ministral 3 14B (Reasoning)Mistral · Open weight
4.6%
175
Ministral 3 14BMistral · Open weight
4.6%
176
Gemma 3 27BGoogle · Open weight
4.4%
177
Claude 4 SonnetAnthropic · Closed
4.3%
178
Ministral 3 8B (Reasoning)Mistral · Open weight
4.3%
179
Ministral 3 8BMistral · Open weight
4.3%
180
Mistral Large 3Mistral · Closed
4.2%
181
GPT-4.1OpenAI · Closed
4.2%
182
GPT-4o miniOpenAI · Closed
4.2%
183
Gemini 1.0 ProGoogle · Closed
4.2%
184
LFM2-24B-A2BLiquidAI · Closed
4.2%
185
Mistral Medium 3Mistral · Closed
4.1%
186
Claude 3 HaikuAnthropic · Closed
4.1%
187
Llama 3.1 405BMeta · Open weight
4.0%
188
Gemma 4 E4BGoogle · Open weight
3.8%
189
Llama 4 ScoutMeta · Open weight
3.8%
190
GPT-4.1 nanoOpenAI · Closed
3.8%
191
Phi-4Microsoft · Open weight
3.8%
192
Solar Pro 2Upstage · Closed
3.7%
193
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
3.6%
194
Qwen2.5 Coder 32B InstructAlibaba · Open weight
3.5%
195
Mistral Large 2Mistral · Closed
3.3%
196
Nova ProAmazon · Closed
3.2%
197
GPT-4 TurboOpenAI · Closed
3.1%
198
DeepSeek V3DeepSeek · Open weight
2.9%
199
Claude 3 OpusAnthropic · Closed
2.8%
200
GPT-4oOpenAI · Closed
2.4%

The published AA-HLE snapshot places Claude Fable 5.1 first at 59.1%. The third row is 4.2 points behind. The broader top-10 range is 12.1 points, so the table still separates the published systems.

200 models have been evaluated on AA-HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-HLE is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-HLE

Year

2026

Tasks

Expert-level questions

Format

Accuracy

Difficulty

Frontier expert reasoning

BenchLM stores the Artificial Analysis HLE result separately from the weighted HLE lane so AA refreshes remain display-only.

BenchLM freshness & provenance

Version

AA-HLE 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-HLE measure?

A display-only Artificial Analysis Humanity's Last Exam score.

Which model scores highest on AA-HLE?

Claude Fable 5.1 by Anthropic currently leads with a score of 59.1% on AA-HLE.

How many models are evaluated on AA-HLE?

200 AI models have been evaluated on AA-HLE on BenchLM.

Last updated: September 10, 2026 · BenchLM version AA-HLE 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.