Benchmark profile
Humanity's Last Exam (HLE)
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Data verifiedClaude Mythos 5 leads the HLE leaderboard on BenchLM's July 2026 update with 64.5%, ahead of Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%), across 44 tracked models.
How to read this leaderboard
Use HLE to compare frontier expert-work capability only after checking the protocol attached to each row. A model with search, browsing, or code execution is not taking the same test as a closed-book model.
Operator receipt: 44 sourced rows are currently displayable on this page; the leading published row is Claude Mythos 5 at 64.5%.
Honest limit: BenchLM can preserve published protocol notes and per-row provenance, but it cannot make differently disclosed HLE configurations fully comparable. Treat small gaps across protocols as directional, not decisive.
Top models on HLE — July 20, 2026
As of July 20, 2026, Claude Mythos 5 leads the HLE leaderboard with 64.5% , followed by Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%).
Claude Mythos 5
Anthropic
claude-mythos-5
Muse Spark 1.1
Meta
muse-spark-1-1
GPT-5.4 Pro
OpenAI
gpt-5-4-pro
Leaderboard (44 models)
ScoreAccording to BenchLM.ai, Claude Mythos 5 leads the HLE benchmark with a score of 64.5%, followed by Muse Spark 1.1 (62.1%) and GPT-5.4 Pro (58.7%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.
44 models have been evaluated on HLE. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, HLE contributes 45% of the category score, so strong performance here directly affects a model's overall ranking.
About HLE
Year
2025
Tasks
Expert-level questions
Format
Open-ended and multiple choice
Difficulty
Frontier expert level
Questions span advanced mathematics, theoretical physics, philosophy, and other specialist fields. The benchmark still has substantial headroom, but its public results increasingly mix tool-assisted and closed-book-style protocols.
BenchLM freshness & provenance
Version
Humanity's Last Exam
Refresh cadence
Static
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does HLE measure?
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Which model scores highest on HLE?
Claude Mythos 5 by Anthropic currently leads with a score of 64.5% on HLE.
How many models are evaluated on HLE?
44 AI models have been evaluated on HLE on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.