# Humanity's Last Exam with tools (HLE w/ tools)

> Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

Canonical page: https://benchlm.ai/benchmarks/hlewithtools

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About HLE w/ tools

- Year: 2026
- Tasks: Expert questions with tool use
- Format: Pass@1
- Difficulty: Frontier tool-augmented reasoning
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores HLE w/ tools as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

HLE w/ tools is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 64.7% |
| 2 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 63.9% |
| 3 | [GLM-5.3](/models/glm-5-3) | Z.AI | 62.5% |
| 4 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 60.0% |
| 5 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 57.4% |
| 6 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 57.2% |
| 7 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 56.2% |
| 8 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 56.1% |
| 9 | [Apodex 1.1](/models/apodex-1-1) | Apodex | 56.1% |
| 10 | [Hy4 preview](/models/hy4-preview) | Tencent | 55.4% |
| 11 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 55.3% |
| 12 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 53.5% |
| 13 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 52.6% |
| 14 | [Agents-A1](/models/agents-a1) | InternScience | 47.6% |
| 15 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 47.2% |
| 16 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 45.1% |
| 17 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 37.4% |
| 18 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 33.4% |
| 19 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 30.5% |

## FAQ

### What does HLE w/ tools measure?

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

### Which model scores highest on HLE w/ tools?

Claude Opus 5 by Anthropic currently leads with a score of 64.7% on HLE w/ tools.

### How many models are evaluated on HLE w/ tools?

19 AI models have been evaluated on HLE w/ tools on BenchLM.

### Does HLE w/ tools affect BenchLM's overall score?

Not directly. HLE w/ tools is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on HLE w/ tools

- [Claude Opus 5 vs DeepSeek V4.1 Flash](/compare/claude-opus-5-vs-deepseek-v4-1-flash)
- [DeepSeek V4.1 Flash vs GLM-5.3](/compare/deepseek-v4-1-flash-vs-glm-5-3)
- [GLM-5.3 vs DeepSeek V4 Pro 0813](/compare/deepseek-v4-pro-0813-vs-glm-5-3)
- [DeepSeek V4 Pro 0813 vs Claude Sonnet 5](/compare/claude-sonnet-5-vs-deepseek-v4-pro-0813)
