Skip to main content
BenchLM

Humanity's Last Exam Diamond without tools at highest reasoning effort (HLE-Diamond (highest effort))

We show this table for reference; we do not rank on it.

Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.

Accuracy without tools at highest reasoning effort on HLE-Diamond (highest effort) — September 22, 2026 release table

We mirror the published accuracy without tools at highest reasoning effort view for HLE-Diamond (highest effort). GPT-6 Astra leads the public snapshot at 66.2%, followed by Claude Opus 5.5 (61.2%) and Claude Fable 5.1 (54.1%). We do not use these results to rank models overall.

9 modelsKnowledgeCurrentDisplay onlyUpdated September 22, 2026 release table

Accuracy without tools at highest reasoning effort table (9 models)

Score
1
GPT-6 AstraOpenAI · ClosedNo tools · max reasoning
66.2%
2
Claude Opus 5.5Anthropic · ClosedNo tools · max reasoning
61.2%
3
Claude Fable 5.1Anthropic · ClosedNo tools · max reasoning
54.1%
4
GPT-6 SolOpenAI · ClosedNo tools · max reasoning
43.8%
5
GPT-5.6 SolOpenAI · ClosedNo tools · max reasoning
41.6%
6
Claude Opus 5Anthropic · ClosedNo tools · max reasoning
41.3%
7
Muse Spark 1.3Meta · ClosedNo tools · max reasoning
34.8%
8
Gemini 3.8 FlashGoogle · ClosedNo tools · high reasoning
34.3%
9
Grok 4.7xAI · ClosedNo tools · xhigh reasoning
25.3%

How to read this leaderboard

This view uses the highest reasoning effort available for each listed model on HLE-Diamond, without external tools.

Variant check: The default no-tools table uses high reasoning effort for every model. Its scores and the web-and-code scores have separate tables.

Compare these rows as highest-available-effort system results. An effort setting can change token use, latency, and score, so the numbers should not be mixed with the high-effort table.

Operator receipt: 9 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 66.2%.

Provenance: The scores mirror the reasoning-max view in the Center for AI Safety and Scale AI release. We have not rerun the gated question set.

Source freshness: The benchmark and these scores were published on September 22, 2026.

Honest limit: Highest available effort means different settings across providers, and the release does not provide uncertainty intervals for these scores.

How to read the highest-effort no-tools table

The HLE-Diamond release has a second no-tools view for the highest reasoning effort available to each model. Its nine scores come from the same 1,000-question subset as the default high-effort table.

Reasoning effort changes the evaluation setting. We keep this view separate from the default high-effort and web-and-code tables. It is display only and does not move model rankings.

Snapshot

9 model rows1000 questionsHighest available reasoning effortNo toolsDisplay only

The published HLE-Diamond (highest effort) snapshot places GPT-6 Astra first at 66.2%. The third row is 12.1 points behind. The broader top-10 range is 40.9 points, so the table still separates the published systems.

9 models have been evaluated on HLE-Diamond (highest effort). The benchmark falls in the Knowledge category. HLE-Diamond (highest effort) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HLE-Diamond (highest effort)

Year

2026

Tasks

1,000 expert questions: 500 reasoning and 500 knowledge

Format

Accuracy without tools, highest available reasoning effort

Difficulty

Frontier expert knowledge and reasoning

The release offers this nine-model table behind its reasoning-max view. GPT, Claude, and Muse use max; Gemini uses high; Grok uses xhigh, the highest available settings listed by the benchmark owner.

Freshness and provenance

Version

HLE-Diamond, highest available reasoning effort without tools

Refresh cadence

Static

Staleness state

Current

Question availability

Gated 1,000-question dataset; public aggregate results

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does HLE-Diamond (highest effort) measure?

Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.

Which model leads the published HLE-Diamond (highest effort) snapshot?

GPT-6 Astra currently leads the published HLE-Diamond (highest effort) snapshot with 66.2% accuracy without tools at highest reasoning effort. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on HLE-Diamond (highest effort)?

The September 22, 2026 release table snapshot contains 9 AI models.

Last updated: September 22, 2026 release table · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.