Humanity's Last Exam (HLE)
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Data verified 23 confirmed releases in the last 30 daysSee the free Radar BriefClaude Fable 5.1 leads the HLE leaderboard on BenchLM's September 2026 update with 65%, ahead of Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%), across 57 models.
How to read this leaderboard
Use HLE to compare frontier expert-work capability only after checking the protocol attached to each row. A model with search, browsing, or code execution is not taking the same test as a closed-book model.
Operator receipt: 57 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 65%.
Honest limit: BenchLM can preserve published protocol notes and per-row provenance, but it cannot make differently disclosed HLE configurations fully comparable. Treat small gaps across protocols as directional, not decisive.
Top models on HLE — September 4, 2026
As of September 4, 2026, Claude Fable 5.1 leads the HLE leaderboard with 65% , followed by Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%).
Claude Fable 5.1
Anthropic
Claude Opus 5
Anthropic
Claude Mythos 5
Anthropic
57 modelsKnowledge35% of category scoreCurrentUpdated September 4, 2026
Leaderboard (57 models)
ScoreAccording to BenchLM.ai, Claude Fable 5.1 leads the HLE benchmark with a score of 65%, followed by Claude Opus 5 (64.7%) and Claude Mythos 5 (64.5%). The top models are clustered within 0.5 points, suggesting this benchmark is nearing saturation for frontier models.
57 models have been evaluated on HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, HLE contributes 35% of the category score, so strong performance here directly affects a model's overall ranking.
About HLE
Year
2025
Tasks
Expert-level questions
Format
Open-ended and multiple choice
Difficulty
Frontier expert level
Questions span advanced mathematics, theoretical physics, philosophy, and other specialist fields. The benchmark still has substantial headroom, but its public results increasingly mix tool-assisted and closed-book-style protocols.
BenchLM freshness & provenance
Version
Humanity's Last Exam
Refresh cadence
Static
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does HLE measure?
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Which model scores highest on HLE?
Claude Fable 5.1 by Anthropic currently leads with a score of 65% on HLE.
How many models are evaluated on HLE?
57 AI models have been evaluated on HLE on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.