# Humanity's Last Exam Diamond with web and code tools (HLE-Diamond (web+code))

> Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.

Canonical page: https://benchlm.ai/benchmarks/hle-diamond-web-code

- Category: [Knowledge](/knowledge)
- Last updated: September 22, 2026 release table

## About HLE-Diamond (web+code)

- Year: 2026
- Tasks: 1,000 expert questions with web and code tools
- Format: Accuracy with web and code, reasoning high
- Difficulty: Frontier expert knowledge and reasoning
- Paper: [Introducing HLE-Diamond](https://lastexam.ai/blog/hle-diamond)

The September 22, 2026 release reports eight web-and-code model results on the same 1,000-question Diamond set. The benchmark owner recommends restricted web tools, a sandboxed Python environment, and provider agent harnesses.

HLE-Diamond (web+code) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | Web + code · high reasoning | OpenAI | 82.9% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | Web + code · high reasoning | Anthropic | 73.9% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | Web + code · high reasoning | Anthropic | 72.4% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Web + code · high reasoning | Anthropic | 69.1% |
| 5 | [GPT-6 Sol](/models/gpt-6-sol) | Web + code · high reasoning | OpenAI | 64.9% |
| 6 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Web + code · high reasoning | Google | 60.3% |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | Web + code · high reasoning | OpenAI | 56.5% |
| 8 | [Muse Spark 1.3](/models/muse-spark-1-3) | Web + code · high reasoning | Meta | 55.5% |

## FAQ

### What does HLE-Diamond (web+code) measure?

Accuracy on HLE-Diamond with constrained web search, page fetching, and code execution at high reasoning effort.

### Which model leads the published HLE-Diamond (web+code) snapshot?

GPT-6 Astra currently leads the published HLE-Diamond (web+code) snapshot with a score of 82.9%.

### How many models are evaluated on HLE-Diamond (web+code)?

The September 22, 2026 release table contains 8 AI models.

### Does HLE-Diamond (web+code) affect BenchLM's overall score?

Not directly. HLE-Diamond (web+code) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
