Humanity's Last Exam Diamond without tools at highest reasoning effort (HLE-Diamond (highest effort))
We show this table for reference; we do not rank on it.
Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.
Accuracy without tools at highest reasoning effort on HLE-Diamond (highest effort) — September 22, 2026 release table
We mirror the published accuracy without tools at highest reasoning effort view for HLE-Diamond (highest effort). GPT-6 Astra leads the public snapshot at 66.2%, followed by Claude Opus 5.5 (61.2%) and Claude Fable 5.1 (54.1%). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
No tools · max reasoning
Claude Opus 5.5
Anthropic
No tools · max reasoning
Claude Fable 5.1
Anthropic
No tools · max reasoning
9 modelsKnowledgeCurrentDisplay onlyUpdated September 22, 2026 release table
Accuracy without tools at highest reasoning effort table (9 models)
ScoreHow to read this leaderboard
This view uses the highest reasoning effort available for each listed model on HLE-Diamond, without external tools.
Variant check: The default no-tools table uses high reasoning effort for every model. Its scores and the web-and-code scores have separate tables.
Compare these rows as highest-available-effort system results. An effort setting can change token use, latency, and score, so the numbers should not be mixed with the high-effort table.
Operator receipt: 9 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 66.2%.
Provenance: The scores mirror the reasoning-max view in the Center for AI Safety and Scale AI release. We have not rerun the gated question set.
Source freshness: The benchmark and these scores were published on September 22, 2026.
Honest limit: Highest available effort means different settings across providers, and the release does not provide uncertainty intervals for these scores.
How to read the highest-effort no-tools table
The HLE-Diamond release has a second no-tools view for the highest reasoning effort available to each model. Its nine scores come from the same 1,000-question subset as the default high-effort table.
Reasoning effort changes the evaluation setting. We keep this view separate from the default high-effort and web-and-code tables. It is display only and does not move model rankings.
Snapshot
The published HLE-Diamond (highest effort) snapshot places GPT-6 Astra first at 66.2%. The third row is 12.1 points behind. The broader top-10 range is 40.9 points, so the table still separates the published systems.
9 models have been evaluated on HLE-Diamond (highest effort). The benchmark falls in the Knowledge category. HLE-Diamond (highest effort) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About HLE-Diamond (highest effort)
Year
2026
Tasks
1,000 expert questions: 500 reasoning and 500 knowledge
Format
Accuracy without tools, highest available reasoning effort
Difficulty
Frontier expert knowledge and reasoning
The release offers this nine-model table behind its reasoning-max view. GPT, Claude, and Muse use max; Gemini uses high; Grok uses xhigh, the highest available settings listed by the benchmark owner.
Freshness and provenance
Version
HLE-Diamond, highest available reasoning effort without tools
Refresh cadence
Static
Staleness state
Current
Question availability
Gated 1,000-question dataset; public aggregate results
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does HLE-Diamond (highest effort) measure?
Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.
Which model leads the published HLE-Diamond (highest effort) snapshot?
GPT-6 Astra currently leads the published HLE-Diamond (highest effort) snapshot with 66.2% accuracy without tools at highest reasoning effort. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on HLE-Diamond (highest effort)?
The September 22, 2026 release table snapshot contains 9 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.