Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

Humanity's Last Exam (HLE)

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Data verified 23 confirmed releases in the last 30 daysSee the free Radar Brief

Claude Fable 5.1 leads the HLE leaderboard on BenchLM's September 2026 update with 65%, ahead of Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%), across 57 models.

How to read this leaderboard

Use HLE to compare frontier expert-work capability only after checking the protocol attached to each row. A model with search, browsing, or code execution is not taking the same test as a closed-book model.

Operator receipt: 57 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 65%.

Honest limit: BenchLM can preserve published protocol notes and per-row provenance, but it cannot make differently disclosed HLE configurations fully comparable. Treat small gaps across protocols as directional, not decisive.

Top models on HLE — September 4, 2026

As of September 4, 2026, Claude Fable 5.1 leads the HLE leaderboard with 65% , followed by Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%).

57 modelsKnowledge35% of category scoreCurrentUpdated September 4, 2026

Leaderboard (57 models)

Score
1
Claude Fable 5.1Anthropic · Closed
65%
2
Claude Opus 5Anthropic · Closed
64.7%
3
Claude Mythos 5Anthropic · Closed
64.5%
4
Muse Spark 1.1Meta · Closed
62.1%
5
GPT-5.4 ProOpenAI · Closed
58.7%
6
Claude Opus 4.8Anthropic · Closed
57.9%
7
Claude Sonnet 5Anthropic · Closed
57.4%
8
GPT-5.5 ProOpenAI · Closed
57.2%
9
Apodex 1.1Apodex · Closed
56.1%
10
Kimi K3Moonshot AI · Closed
56%
11
Hy4 previewTencent · Open weight
55.4%
12
GLM-5.2Z.AI · Open weight
54.7%
13
Claude Opus 4.7 (Adaptive)Anthropic · Closed
54.7%
14
Claude Opus 4.6Anthropic · Closed
53%
15
dots3-note PreviewDots Studio · Open weight
52.6%
16
GLM-5.1Z.AI · Open weight
52.3%
17
GPT-5.5OpenAI · Closed
52.2%
18
GPT-5.4OpenAI · Closed
52.1%
19
Muse SparkMeta · Closed
50.4%
20
GLM-5Z.AI · Open weight
50.4%
21
Claude Sonnet 4.6Anthropic · Closed
49%
22
MiMo-V2.5-ProXiaomi · Closed
48%
23
Inkling-SmallThinking Machines Lab · Open weight
47.8%
24
Agents-A1InternScience · Open weight
47.6%
25
InklingThinking Machines Lab · Open weight
46%
26
Ornith-1.5-397BOrnith AI · Open weight
44.6%
27
Qwen3.8 MaxAlibaba · Open weight
43.6%
28
DeepSeek V4 Pro 0813DeepSeek · Closed
42.7%
29
GPT-5.4 miniOpenAI · Closed
41.5%
30
Qwen3.7 MaxAlibaba · Closed
41.4%
31
Gemini 3.5 FlashGoogle · Closed
40.2%
32
GPT-5.4 nanoOpenAI · Closed
37.7%
33
Qwen3.8-Flash-NextAlibaba · Open weight
35.9%
34
Grok 4.3xAI · Closed
35%
35
DeepSeek V4 Flash 0731DeepSeek · Closed
34.8%
36
Qwen3.7 PlusAlibaba · Closed
34.7%
37
Kimi K2.6Moonshot AI · Open weight
34.7%
38
DeepSeek V4 Pro (High)DeepSeek · Open weight
34.5%
39
Claude Opus 4.5Anthropic · Closed
30.8%
40
Qwen3.8-27BAlibaba · Open weight
30.8%
41
Kimi K2.5Moonshot AI · Open weight
30.1%
42
DeepSeek V4 Flash (High)DeepSeek · Closed
29.4%
43
Qwen3.6 PlusAlibaba · Closed
28.8%
44
Qwen3.5 397BAlibaba · Open weight
28.7%
45
Nemotron 3 UltraNVIDIA · Open weight
26.7%
46
Gemma 4 31BGoogle · Open weight
26.5%
47
Ornith-1.5-35B-A3BOrnith AI · Open weight
25.6%
48
Hy3 PreviewTencent · Open weight
25.5%
49
GLM-4.7Z.AI · Open weight
24.8%
50
Qwen3.6-27BAlibaba · Open weight
24%
51
Ling 3.0 FlashInclusionAI · Open weight
22.7%
52
Qwen3.6-35B-A3BAlibaba · Open weight
21.4%
53
Ornith-1.5-9BOrnith AI · Open weight
20.2%
54
Gemini 2.5 ProGoogle · Closed
18.8%
55
Gemma 4 26B A4BGoogle · Open weight
17.2%
56
DeepSeek V4 FlashDeepSeek · Closed
8.1%
57
DeepSeek V4 ProDeepSeek · Open weight
7.7%

According to BenchLM.ai, Claude Fable 5.1 leads the HLE benchmark with a score of 65%, followed by Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%). The top models are clustered within 0.5 points, suggesting this benchmark is nearing saturation for frontier models.

57 models have been evaluated on HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, HLE contributes 35% of the category score, so strong performance here directly affects a model's overall ranking.

About HLE

Year

2025

Tasks

Expert-level questions

Format

Open-ended and multiple choice

Difficulty

Frontier expert level

Questions span advanced mathematics, theoretical physics, philosophy, and other specialist fields. The benchmark still has substantial headroom, but its public results increasingly mix tool-assisted and closed-book-style protocols.

BenchLM freshness & provenance

Version

Humanity's Last Exam

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HLE measure?

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Which model scores highest on HLE?

Claude Fable 5.1 by Anthropic currently leads with a score of 65% on HLE.

How many models are evaluated on HLE?

57 AI models have been evaluated on HLE on BenchLM.

Last updated: September 4, 2026 · BenchLM version Humanity's Last Exam

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.