Skip to main content

Benchmark profile

Humanity's Last Exam (HLE)

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Data verified

Claude Mythos 5 leads the HLE leaderboard on BenchLM's July 2026 update with 64.5%, ahead of Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%), across 44 tracked models.

How to read this leaderboard

Use HLE to compare frontier expert-work capability only after checking the protocol attached to each row. A model with search, browsing, or code execution is not taking the same test as a closed-book model.

Operator receipt: 44 sourced rows are currently displayable on this page; the leading published row is Claude Mythos 5 at 64.5%.

Honest limit: BenchLM can preserve published protocol notes and per-row provenance, but it cannot make differently disclosed HLE configurations fully comparable. Treat small gaps across protocols as directional, not decisive.

Top models on HLE — July 20, 2026

As of July 20, 2026, Claude Mythos 5 leads the HLE leaderboard with 64.5% , followed by Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%).

44 modelsKnowledge45% of category scoreCurrentUpdated July 20, 2026

Leaderboard (44 models)

Score
1
Claude Mythos 5Anthropic · Closed
64.5%
2
Muse Spark 1.1Meta · Closed
62.1%
3
GPT-5.4 ProOpenAI · Closed
58.7%
4
Claude Opus 4.8Anthropic · Closed
57.9%
5
Claude Sonnet 5Anthropic · Closed
57.4%
6
GPT-5.5 ProOpenAI · Closed
57.2%
7
Kimi K3Moonshot AI · Closed
56%
8
Claude Opus 4.7 (Adaptive)Anthropic · Closed
54.7%
9
GLM-5.2Z.AI · Open weight
54.7%
10
Claude Opus 4.6Anthropic · Closed
53%
11
GLM-5.1Z.AI · Open weight
52.3%
12
GPT-5.5OpenAI · Closed
52.2%
13
GPT-5.4OpenAI · Closed
52.1%
14
GLM-5Z.AI · Open weight
50.4%
15
Muse SparkMeta · Closed
50.4%
16
Claude Sonnet 4.6Anthropic · Closed
49%
17
MiMo-V2.5-ProXiaomi · Closed
48%
18
Agents-A1InternScience · Open weight
47.6%
19
InklingThinking Machines Lab · Open weight
46%
20
GPT-5.4 miniOpenAI · Closed
41.5%
21
Qwen3.7 MaxAlibaba · Closed
41.4%
22
Gemini 3.5 FlashGoogle · Closed
40.2%
23
DeepSeek V4 Pro (Max)DeepSeek · Open weight
37.7%
24
GPT-5.4 nanoOpenAI · Closed
37.7%
25
Grok 4.3xAI · Closed
35%
26
DeepSeek V4 Flash (Max)DeepSeek · Open weight
34.8%
27
Qwen3.7 PlusAlibaba · Closed
34.7%
28
Kimi K2.6Moonshot AI · Open weight
34.7%
29
DeepSeek V4 Pro (High)DeepSeek · Open weight
34.5%
30
Claude Opus 4.5Anthropic · Closed
30.8%
31
Kimi K2.5Moonshot AI · Open weight
30.1%
32
DeepSeek V4 Flash (High)DeepSeek · Open weight
29.4%
33
Qwen3.6 PlusAlibaba · Closed
28.8%
34
Qwen3.5 397BAlibaba · Open weight
28.7%
35
Nemotron 3 UltraNVIDIA · Open weight
26.7%
36
Gemma 4 31BGoogle · Open weight
26.5%
37
Hy3 PreviewTencent · Open weight
25.5%
38
GLM-4.7Z.AI · Open weight
24.8%
39
Qwen3.6-27BAlibaba · Open weight
24%
40
Qwen3.6-35B-A3BAlibaba · Open weight
21.4%
41
Gemini 2.5 ProGoogle · Closed
18.8%
42
Gemma 4 26B A4BGoogle · Open weight
17.2%
43
DeepSeek V4 FlashDeepSeek · Open weight
8.1%
44
DeepSeek V4 ProDeepSeek · Open weight
7.7%

According to BenchLM.ai, Claude Mythos 5 leads the HLE benchmark with a score of 64.5%, followed by Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

44 models have been evaluated on HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, HLE contributes 45% of the category score, so strong performance here directly affects a model's overall ranking.

About HLE

Year

2025

Tasks

Expert-level questions

Format

Open-ended and multiple choice

Difficulty

Frontier expert level

Questions span advanced mathematics, theoretical physics, philosophy, and other specialist fields. The benchmark still has substantial headroom, but its public results increasingly mix tool-assisted and closed-book-style protocols.

BenchLM freshness & provenance

Version

Humanity's Last Exam

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HLE measure?

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Which model scores highest on HLE?

Claude Mythos 5 by Anthropic currently leads with a score of 64.5% on HLE.

How many models are evaluated on HLE?

44 AI models have been evaluated on HLE on BenchLM.

Last updated: July 20, 2026 · BenchLM version Humanity's Last Exam

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.