Skip to main content
BenchLM

Humanity's Last Exam Diamond with web and code tools (HLE-Diamond (web+code))

We show this table for reference; we do not rank on it.

Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.

Accuracy with web and code at high reasoning effort on HLE-Diamond (web+code) — September 22, 2026 release table

We mirror the published accuracy with web and code at high reasoning effort view for HLE-Diamond (web+code). GPT-6 Astra leads the public snapshot at 82.9%, followed by Claude Opus 5.5 (73.9%) and Claude Fable 5.1 (72.4%). We do not use these results to rank models overall.

8 modelsKnowledgeCurrentDisplay onlyUpdated September 22, 2026 release table

Accuracy with web and code at high reasoning effort table (8 models)

Score
1
GPT-6 AstraOpenAI · ClosedWeb + code · high reasoning
82.9%
2
Claude Opus 5.5Anthropic · ClosedWeb + code · high reasoning
73.9%
3
Claude Fable 5.1Anthropic · ClosedWeb + code · high reasoning
72.4%
4
Claude Opus 5Anthropic · ClosedWeb + code · high reasoning
69.1%
5
GPT-6 SolOpenAI · ClosedWeb + code · high reasoning
64.9%
6
Gemini 3.8 FlashGoogle · ClosedWeb + code · high reasoning
60.3%
7
GPT-5.6 SolOpenAI · ClosedWeb + code · high reasoning
56.5%
8
Muse Spark 1.3Meta · ClosedWeb + code · high reasoning
55.5%

How to read this leaderboard

The HLE-Diamond web-and-code table tests the 1,000-question Diamond set with external web and code tools. Each published model uses high reasoning effort.

Variant check: These are tool-assisted system results. The separate HLE-Diamond table without tools tests the same questions under a different protocol.

Read each score with its agent harness and tool restrictions. The published setup uses restricted search and fetch tools plus sandboxed Python; it does not isolate model-only ability.

Operator receipt: 8 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 82.9%.

Provenance: The scores mirror the Center for AI Safety and Scale AI release table. We have not rerun the gated question set or the provider harnesses.

Source freshness: The tool-assisted scores were published on September 22, 2026.

Honest limit: Different provider harnesses are part of these results. Grok 4.7 has no published web-and-code score in the release, so it is absent from this table.

How to read the HLE-Diamond web-and-code table

The September 22, 2026 HLE-Diamond release reports eight web-and-code results at high reasoning effort. The source does not publish a web-and-code result for Grok 4.7.

Web and code access changes the task. We show these runs separately from both no-tools views and keep them out of weighted model rankings.

Snapshot

8 model rows1000 questionsHigh reasoning effortWeb + codeDisplay only

The published HLE-Diamond (web+code) snapshot places GPT-6 Astra first at 82.9%. The third row is 10.5 points behind. The broader top-10 range is 27.4 points, so the table still separates the published systems.

8 models have been evaluated on HLE-Diamond (web+code). The benchmark falls in the Knowledge category. HLE-Diamond (web+code) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HLE-Diamond (web+code)

Year

2026

Tasks

1,000 expert questions with web and code tools

Format

Accuracy with web and code, reasoning high

Difficulty

Frontier expert knowledge and reasoning

The September 22, 2026 release reports eight web-and-code model results on the same 1,000-question Diamond set. The benchmark owner recommends restricted web tools, a sandboxed Python environment, and provider agent harnesses.

Freshness and provenance

Version

HLE-Diamond, reasoning high with web and code

Refresh cadence

Static

Staleness state

Current

Question availability

Gated 1,000-question dataset; public aggregate results

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does HLE-Diamond (web+code) measure?

Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.

Which model leads the published HLE-Diamond (web+code) snapshot?

GPT-6 Astra currently leads the published HLE-Diamond (web+code) snapshot with 82.9% accuracy with web and code at high reasoning effort. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on HLE-Diamond (web+code)?

The September 22, 2026 release table snapshot contains 8 AI models.

Last updated: September 22, 2026 release table · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.