Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

Artificial Analysis Humanity's Last Exam (AA-HLE)

A display-only Artificial Analysis Humanity's Last Exam score.

Data verified 25 confirmed releases in the last 30 daysStart the free Radar Brief

Benchmark score on AA-HLE — August 29, 2026

We mirror the published score view for AA-HLE. Claude Fable 5 leads the public snapshot at 55.5%, followed by Claude Opus 5 (54.9%) and GPT-5.6 Sol (49.5%). We do not use these results to rank models overall.

179 modelsKnowledgeCurrentDisplay onlyUpdated August 29, 2026

Benchmark score table (179 models)

Score
1
Claude Fable 5Anthropic · Closed
55.5%
2
Claude Opus 5Anthropic · Closed
54.9%
3
GPT-5.6 SolOpenAI · Closed
49.5%
4
Claude Opus 4.8Anthropic · Closed
48.7%
5
Gemini 3.7 FlashGoogle · Closed
47.9%
6
Gemini 3.1 ProGoogle · Closed
47.0%
7
Kimi K3Moonshot AI · Closed
46.9%
8
Muse Spark 1.1Meta · Closed
46.2%
9
GPT-5.5OpenAI · Closed
45.8%
10
Muse Spark 1.2Meta · Closed
45.5%
11
GPT-5.4OpenAI · Closed
43.7%
12
Qwen3.8 Max PreviewAlibaba · Closed
43.0%
13
GPT-5.6 TerraOpenAI · Closed
42.9%
14
Grok 4.6xAI · Closed
42.9%
15
Gemini 3.5 FlashGoogle · Closed
42.7%
16
Grok 4.5xAI · Closed
42.7%
17
GPT-5.3 CodexOpenAI · Closed
42.5%
18
GPT-5.3-Codex-SparkOpenAI · Closed
42.5%
19
Claude Opus 4.7 (Adaptive)Anthropic · Closed
42.3%
20
GLM-5.3Z.AI · Open weight
42.3%
21
Claude Sonnet 5Anthropic · Closed
41.3%
22
GLM-5.2Z.AI · Open weight
41.1%
23
DeepSeek V4 Pro 0813DeepSeek · Closed
41.0%
24
Gemini 3.6 FlashGoogle · Closed
40.8%
25
Muse SparkMeta · Closed
40.7%
26
Qwen3.7 MaxAlibaba · Closed
40.5%
27
GLM-5.3-FlashZ.AI · Open weight
39.9%
28
Claude Opus 4.6 (Adaptive)Anthropic · Closed
39.9%
29
Gemini 3 ProGoogle · Closed
39.7%
30
GPT-5.6 LunaOpenAI · Closed
39.5%
31
MiniMax M3MiniMax · Open weight
39.0%
32
DeepSeek V4 Flash 0731DeepSeek · Closed
38.6%
33
Qwen3.8-Flash-NextAlibaba · Open weight
38.0%
34
GPT-5.2OpenAI · Closed
37.7%
35
Kimi K2.6Moonshot AI · Open weight
37.5%
36
Grok 4.3xAI · Closed
37.2%
37
MiMo-V2.5-ProXiaomi · Closed
35.7%
38
GPT-5.2-CodexOpenAI · Closed
35.7%
39
Qwen3.7 PlusAlibaba · Closed
35.6%
40
DeepSeek V4 Pro (High)DeepSeek · Open weight
35.2%
41
Kimi K2.7 CodeMoonshot AI · Open weight
35.0%
42
Qwen3.8-27BAlibaba · Open weight
33.9%
43
Hy3 PreviewTencent · Open weight
33.5%
44
Hy3Tencent · Open weight
33.5%
45
Inkling-SmallThinking Machines Lab · Open weight
33.3%
46
Claude Opus 4.7Anthropic · Closed
33.3%
47
InklingThinking Machines Lab · Open weight
31.9%
48
Qwen 3.6 Max (preview)Alibaba · Closed
30.8%
49
Kimi K2.5Moonshot AI · Open weight
30.7%
50
Kimi K2.5 (Reasoning)Moonshot AI · Closed
30.7%
51
MiMo-V2-ProXiaomi · Closed
30.4%
52
GLM-5.1Z.AI · Open weight
30.1%
53
Claude Opus 4.5 ThinkingAnthropic · Closed
30.1%
54
MiniMax M2.7MiniMax · Open weight
29.6%
55
GLM-5Z.AI · Open weight
29.3%
56
Qwen3.5 397BAlibaba · Open weight
29.0%
57
Qwen3.5 397B (Reasoning)Alibaba · Open weight
29.0%
58
GPT-5.1OpenAI · Closed
28.5%
59
GPT-5 (high)OpenAI · Closed
28.5%
60
Nemotron 3 UltraNVIDIA · Open weight
28.4%
61
GPT-5.4 nanoOpenAI · Closed
28.3%
62
GPT-5.4 miniOpenAI · Closed
28.1%
63
Qwen3.6 PlusAlibaba · Closed
27.8%
64
GLM-5-TurboZ.AI · Closed
27.8%
65
GLM-4.7Z.AI · Open weight
27.4%
66
Grok 4xAI · Closed
26.7%
67
GPT-5.1-Codex-MaxOpenAI · Closed
25.7%
68
GPT-5.1-CodexOpenAI · Closed
25.7%
69
GPT-5 (medium)OpenAI · Closed
25.4%
70
Qwen3.5-122B-A10BAlibaba · Open weight
25.2%
71
Step 3.5 FlashStepFun · Open weight
24.5%
72
Qwen3.5-27BAlibaba · Open weight
23.9%
73
Ling 3.0 FlashInclusionAI · Open weight
23.7%
74
Ling 3.0 Flash FP8InclusionAI · Open weight
23.7%
75
Gemma 4 31BGoogle · Open weight
23.6%
76
Qwen3.6-27BAlibaba · Open weight
23.1%
77
Gemini 2.5 ProGoogle · Closed
22.5%
78
Qwen3.6-35B-A3BAlibaba · Open weight
22.2%
79
MiMo-V2-OmniXiaomi · Closed
22.1%
80
Muse Glimmer 30BMeta · Open weight
22.0%
81
GPT-5 miniOpenAI · Closed
21.5%
82
Step 3.7 FlashStepFun · Open weight
21.4%
83
Qwen3.5-35B-A3BAlibaba · Open weight
21.0%
84
Nemotron 3 Super 100BNVIDIA · Open weight
20.8%
85
Nemotron 3 Super 120B A12BNVIDIA · Open weight
20.8%
86
MiniMax M2.5MiniMax · Closed
20.5%
87
o3OpenAI · Closed
20.1%
88
GPT-OSS 120BOpenAI · Open weight
19.6%
89
Gemma 4 26B A4BGoogle · Open weight
19.3%
90
19.3%
91
Claude Opus 4.6Anthropic · Closed
19.1%
92
19.1%
93
Gemini 3.5 Flash-LiteGoogle · Closed
18.8%
94
GLM-5V-TurboZ.AI · Closed
17.1%
95
Mercury 2Inception · Closed
17.1%
96
DeepSeek-R1DeepSeek · Open weight
15.8%
97
Trinity-Large-PreviewArcee AI · Open weight
15.8%
98
Trinity-Large-ThinkingArcee AI · Open weight
15.8%
99
Gemma 4 12BGoogle · Open weight
15.7%
100
Gemini 3 FlashGoogle · Closed
15.0%
101
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
14.3%
102
K-ExaoneLG AI Research · Closed
13.9%
103
Mistral Medium 3.5 128BMistral · Open weight
13.8%
104
Claude Sonnet 4.6Anthropic · Closed
13.3%
105
Claude Opus 4.5Anthropic · Closed
13.2%
106
Claude 4.1 Opus ThinkingAnthropic · Closed
12.5%
107
Command A+Cohere · Open weight
12.0%
108
Qwen3 MaxAlibaba · Closed
11.9%
109
Nemotron 3 Nano 30BNVIDIA · Open weight
11.4%
110
DeepSeek V3.2DeepSeek · Open weight
11.2%
111
GPT-OSS 20BOpenAI · Open weight
11.0%
112
Sarvam 105BSarvam · Open weight
11.0%
113
10.6%
114
Mistral Small 4Mistral · Open weight
9.9%
115
Mistral Small 4 (Reasoning)Mistral · Open weight
9.9%
116
GPT-5 nanoOpenAI · Closed
9.5%
117
MiniMax M1 80kMiniMax · Closed
8.9%
118
MiMo-V2-FlashXiaomi · Open weight
8.6%
119
Grok Code Fast 1xAI · Closed
8.0%
120
o3-miniOpenAI · Closed
7.9%
121
GLM-4.7-FlashZ.AI · Open weight
7.6%
122
Sarvam 30BSarvam · Open weight
7.5%
123
Kimi K2Moonshot AI · Closed
7.4%
124
Nemotron Ultra 253BNVIDIA · Open weight
7.4%
125
o1OpenAI · Closed
7.0%
126
GLM-4.5-AirZ.AI · Closed
7.0%
127
LFM2.5-8B-A1BLiquidAI · Open weight
6.9%
128
Celeris-1Celeris · Closed
6.8%
129
DeepSeek V3.1DeepSeek · Open weight
6.7%
130
LFM2.5-1.2B-InstructLiquidAI · Closed
6.7%
131
Granite-4.0-H-350MIBM · Open weight
6.4%
132
Ling 2.6 FlashInclusionAI · Open weight
6.3%
133
LFM2.5-2.6BLiquidAI · Open weight
6.2%
134
LFM2.5-1.2B-ThinkingLiquidAI · Closed
6.2%
135
Exaone 4.0 1.2BLG AI Research · Open weight
5.7%
136
GLM-4.6Z.AI · Open weight
5.5%
137
Granite-4.0-350MIBM · Open weight
5.5%
138
Ministral 3 3B (Reasoning)Mistral · Open weight
5.4%
139
Ministral 3 3BMistral · Open weight
5.4%
140
Grok 4.1 FastxAI · Closed
5.1%
141
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
5.1%
142
Exaone 4.0 32BLG AI Research · Open weight
5.0%
143
GPT-4.1 miniOpenAI · Closed
5.0%
144
Granite-4.0-H-1BIBM · Open weight
5.0%
145
Phi-4 Multimodal InstructMicrosoft · Open weight
5.0%
146
Llama 4 MaverickMeta · Open weight
4.9%
147
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
4.8%
148
Gemma 4 E2BGoogle · Open weight
4.8%
149
Granite-4.0-1BIBM · Open weight
4.8%
150
Gemini 2.5 FlashGoogle · Closed
4.7%
151
Gemini 1.5 ProGoogle · Closed
4.6%
152
DeepSeek R1 Distill Qwen 32BDeepSeek · Open weight
4.6%
153
Ministral 3 14B (Reasoning)Mistral · Open weight
4.6%
154
Ministral 3 14BMistral · Open weight
4.6%
155
Gemma 3 27BGoogle · Open weight
4.4%
156
Claude 4 SonnetAnthropic · Closed
4.3%
157
Ministral 3 8B (Reasoning)Mistral · Open weight
4.3%
158
Ministral 3 8BMistral · Open weight
4.3%
159
GPT-4.1OpenAI · Closed
4.2%
160
Mistral Large 3Mistral · Closed
4.2%
161
GPT-4o miniOpenAI · Closed
4.2%
162
Gemini 1.0 ProGoogle · Closed
4.2%
163
LFM2-24B-A2BLiquidAI · Closed
4.2%
164
Claude 3 HaikuAnthropic · Closed
4.1%
165
Mistral Medium 3Mistral · Closed
4.1%
166
Llama 3.1 405BMeta · Open weight
4.0%
167
GPT-4.1 nanoOpenAI · Closed
3.8%
168
Llama 4 ScoutMeta · Open weight
3.8%
169
Phi-4Microsoft · Open weight
3.8%
170
Gemma 4 E4BGoogle · Open weight
3.8%
171
Solar Pro 2Upstage · Closed
3.7%
172
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
3.6%
173
Qwen2.5 Coder 32B InstructAlibaba · Open weight
3.5%
174
Mistral Large 2Mistral · Closed
3.3%
175
Nova ProAmazon · Closed
3.2%
176
GPT-4 TurboOpenAI · Closed
3.1%
177
DeepSeek V3DeepSeek · Open weight
2.9%
178
Claude 3 OpusAnthropic · Closed
2.8%
179
GPT-4oOpenAI · Closed
2.4%

The published AA-HLE snapshot places Claude Fable 5 first at 55.5%. The third row is 6.0 points behind. The broader top-10 range is 10.0 points, so the table still separates the published systems.

179 models have been evaluated on AA-HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-HLE is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-HLE

Year

2026

Tasks

Expert-level questions

Format

Accuracy

Difficulty

Frontier expert reasoning

BenchLM stores the Artificial Analysis HLE result separately from the weighted HLE lane so AA refreshes remain display-only.

BenchLM freshness & provenance

Version

AA-HLE 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-HLE measure?

A display-only Artificial Analysis Humanity's Last Exam score.

Which model scores highest on AA-HLE?

Claude Fable 5 by Anthropic currently leads with a score of 55.5% on AA-HLE.

How many models are evaluated on AA-HLE?

179 AI models have been evaluated on AA-HLE on BenchLM.

Last updated: August 29, 2026 · BenchLM version AA-HLE 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.