Humanity's Last Exam Diamond with web and code tools (HLE-Diamond (web+code))
We show this table for reference; we do not rank on it.
Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.
Accuracy with web and code at high reasoning effort on HLE-Diamond (web+code) — September 22, 2026 release table
We mirror the published accuracy with web and code at high reasoning effort view for HLE-Diamond (web+code). GPT-6 Astra leads the public snapshot at 82.9%, followed by Claude Opus 5.5 (73.9%) and Claude Fable 5.1 (72.4%). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
Web + code · high reasoning
Claude Opus 5.5
Anthropic
Web + code · high reasoning
Claude Fable 5.1
Anthropic
Web + code · high reasoning
8 modelsKnowledgeCurrentDisplay onlyUpdated September 22, 2026 release table
Accuracy with web and code at high reasoning effort table (8 models)
ScoreHow to read this leaderboard
The HLE-Diamond web-and-code table tests the 1,000-question Diamond set with external web and code tools. Each published model uses high reasoning effort.
Variant check: These are tool-assisted system results. The separate HLE-Diamond table without tools tests the same questions under a different protocol.
Read each score with its agent harness and tool restrictions. The published setup uses restricted search and fetch tools plus sandboxed Python; it does not isolate model-only ability.
Operator receipt: 8 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 82.9%.
Provenance: The scores mirror the Center for AI Safety and Scale AI release table. We have not rerun the gated question set or the provider harnesses.
Source freshness: The tool-assisted scores were published on September 22, 2026.
Honest limit: Different provider harnesses are part of these results. Grok 4.7 has no published web-and-code score in the release, so it is absent from this table.
How to read the HLE-Diamond web-and-code table
The September 22, 2026 HLE-Diamond release reports eight web-and-code results at high reasoning effort. The source does not publish a web-and-code result for Grok 4.7.
Web and code access changes the task. We show these runs separately from both no-tools views and keep them out of weighted model rankings.
Snapshot
The published HLE-Diamond (web+code) snapshot places GPT-6 Astra first at 82.9%. The third row is 10.5 points behind. The broader top-10 range is 27.4 points, so the table still separates the published systems.
8 models have been evaluated on HLE-Diamond (web+code). The benchmark falls in the Knowledge category. HLE-Diamond (web+code) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About HLE-Diamond (web+code)
Year
2026
Tasks
1,000 expert questions with web and code tools
Format
Accuracy with web and code, reasoning high
Difficulty
Frontier expert knowledge and reasoning
The September 22, 2026 release reports eight web-and-code model results on the same 1,000-question Diamond set. The benchmark owner recommends restricted web tools, a sandboxed Python environment, and provider agent harnesses.
Freshness and provenance
Version
HLE-Diamond, reasoning high with web and code
Refresh cadence
Static
Staleness state
Current
Question availability
Gated 1,000-question dataset; public aggregate results
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does HLE-Diamond (web+code) measure?
Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.
Which model leads the published HLE-Diamond (web+code) snapshot?
GPT-6 Astra currently leads the published HLE-Diamond (web+code) snapshot with 82.9% accuracy with web and code at high reasoning effort. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on HLE-Diamond (web+code)?
The September 22, 2026 release table snapshot contains 8 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.