Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Humanity's Last Exam without tools (HLE w/o tools)

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

Top models on HLE w/o tools — September 18, 2026

As of September 18, 2026, Claude Fable 5.1 leads the HLE w/o tools leaderboard with 60.9% , followed by Claude Mythos 5 (59%) and Claude Opus 5 (56.3%).

38 modelsKnowledge10% of category scoreCurrentUpdated September 18, 2026

Leaderboard (38 models)

Score
1
Claude Fable 5.1Anthropic · Closed
60.9%
2
Claude Mythos 5Anthropic · Closed
59%
3
Claude Opus 5Anthropic · Closed
56.3%
4
Muse Spark 1.1Meta · Closed
52.2%
5
Sakana Fugu-UltraSakana AI · Closed
50%
6
Claude Opus 4.8Anthropic · Closed
49.8%
7
Pareto 26.9Unbiased · Closed
49%
8
Sakana FuguSakana AI · Closed
47.2%
9
Claude Opus 4.7 (Adaptive)Anthropic · Closed
46.9%
10
Gemini 3.1 ProGoogle · Closed
45.4%
11
Ornith-1.5-397BOrnith AI · Open weight
44.6%
12
Qwen3.8 MaxAlibaba · Open weight
43.6%
13
Kimi K3Moonshot AI · Closed
43.5%
14
Hy4 previewTencent · Open weight
43.4%
15
Claude Sonnet 5Anthropic · Closed
43.2%
16
GPT-5.5 ProOpenAI · Closed
43.1%
17
Muse SparkMeta · Closed
42.8%
18
GPT-5.4 ProOpenAI · Closed
42.7%
19
GPT-5.5OpenAI · Closed
41.4%
20
GLM-5.2Z.AI · Open weight
40.5%
21
Claude Opus 4.6Anthropic · Closed
40%
22
GPT-5.4OpenAI · Closed
39.8%
23
Qwen3.8-Flash-NextAlibaba · Open weight
35.9%
24
MiMo-V2.5-ProXiaomi · Closed
34%
25
Inkling-SmallThinking Machines Lab · Open weight
31.6%
26
Grok 4.20xAI · Closed
31.6%
27
Qwen3.8-27BAlibaba · Open weight
30.8%
28
InklingThinking Machines Lab · Open weight
30%
29
Solar Open 2Upstage · Open weight
28.8%
30
GPT-5.4 miniOpenAI · Closed
28.2%
31
Nemotron 3 UltraNVIDIA · Open weight
26.7%
32
Ornith-1.5-35B-A3BOrnith AI · Open weight
25.6%
33
GPT-5.4 nanoOpenAI · Closed
24.3%
34
Ornith-1.5-9BOrnith AI · Open weight
20.2%
35
Gemma 4 31BGoogle · Open weight
19.5%
36
10.5%
37
Gemma 4 26B A4BGoogle · Open weight
8.7%
38
Gemma 4 12BGoogle · Open weight
5.2%

According to BenchLM.ai, Claude Fable 5.1 leads the HLE w/o tools benchmark with a score of 60.9%, followed by Claude Mythos 5 (59%) and Claude Opus 5 (56.3%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

38 models have been evaluated on HLE w/o tools. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, HLE w/o tools contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.

About HLE w/o tools

Year

2026

Tasks

Expert-level questions

Format

Tool-free expert QA

Difficulty

Frontier expert level

This variant removes external tools so the score reflects pure model performance on frontier expert questions.

Freshness and provenance

Version

HLE w/o tools 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does HLE w/o tools measure?

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

Which model scores highest on HLE w/o tools?

Claude Fable 5.1 by Anthropic currently leads with a score of 60.9% on HLE w/o tools.

How many models are evaluated on HLE w/o tools?

38 AI models have been evaluated on HLE w/o tools on BenchLM.

Last updated: September 18, 2026 · BenchLM version HLE w/o tools 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.