Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Humanity's Last Exam with tools (HLE w/ tools)

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on HLE w/ tools — September 10, 2026

We mirror the published score view for HLE w/ tools. Claude Opus 5 leads the public snapshot at 64.7%, followed by DeepSeek V4.1 Flash (63.9%) and GLM-5.3 (62.5%). We do not use these results to rank models overall.

21 modelsAgenticCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (21 models)

Score
1
Claude Opus 5Anthropic · Closed
64.7%
2
DeepSeek V4.1 FlashDeepSeek · Open weight
63.9%
3
GLM-5.3Z.AI · Open weight
62.5%
4
DeepSeek V4 Pro 0813DeepSeek · Closed
60.0%
5
Claude Sonnet 5Anthropic · Closed
57.4%
6
GPT-6 AstraOpenAI · Closed
57.2%
7
Qwen3.8 MaxAlibaba · Open weight
56.2%
8
Apodex 1.1Apodex · Closed
56.1%
9
Ornith-1.5-397BOrnith AI · Open weight
56.1%
10
Hy4 previewTencent · Open weight
55.4%
11
GLM-5.3-FlashZ.AI · Open weight
55.3%
12
Qwen3.7 MaxAlibaba · Closed
53.5%
13
dots3-note PreviewDots Studio · Open weight
52.6%
14
Agents-A1InternScience · Open weight
47.6%
15
Step 3.7 FlashStepFun · Open weight
47.2%
16
DeepSeek V4 Flash 0731DeepSeek · Closed
45.1%
17
DeepSeek V4 Pro (High)DeepSeek · Open weight
44.7%
18
DeepSeek V4 Flash (High)DeepSeek · Closed
40.3%
19
Nemotron 3 UltraNVIDIA · Open weight
37.4%
20
Ornith-1.5-35B-A3BOrnith AI · Open weight
33.4%
21
Ornith-1.5-9BOrnith AI · Open weight
30.5%

The published HLE w/ tools snapshot places Claude Opus 5 first at 64.7%. The third row is 2.2 points behind. The broader top-10 range is 9.3 points, so many of the published results sit in a relatively narrow band.

21 models have been evaluated on HLE w/ tools. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. HLE w/ tools is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HLE w/ tools

Year

2026

Tasks

Expert questions with tool use

Format

Pass@1

Difficulty

Frontier tool-augmented reasoning

BenchLM stores HLE w/ tools as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

BenchLM freshness & provenance

Version

HLE w/ tools 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HLE w/ tools measure?

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

Which model scores highest on HLE w/ tools?

Claude Opus 5 by Anthropic currently leads with a score of 64.7% on HLE w/ tools.

How many models are evaluated on HLE w/ tools?

21 AI models have been evaluated on HLE w/ tools on BenchLM.

Last updated: September 10, 2026 · BenchLM version HLE w/ tools 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.