# Humanity's Last Exam Diamond (HLE-Diamond)

> Accuracy on the 1,000-question HLE-Diamond set without tools, using the benchmark owner's high-reasoning evaluation.

Canonical page: https://benchlm.ai/benchmarks/hle-diamond

- Category: [Knowledge](/knowledge)
- Last updated: September 22, 2026 release table

## About HLE-Diamond

- Year: 2026
- Tasks: 1,000 expert questions: 500 reasoning and 500 knowledge
- Format: Accuracy without tools, reasoning high
- Difficulty: Frontier expert knowledge and reasoning
- Paper: [Introducing HLE-Diamond](https://lastexam.ai/blog/hle-diamond)

The September 22, 2026 release reports nine model results. HLE-Diamond contains 500 reasoning and 500 knowledge questions drawn from a refined HLE subset. The no-tools results remain separate from both original HLE and the web-and-code evaluation.

HLE-Diamond is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (9 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | No tools · high reasoning | OpenAI | 60.6% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | No tools · high reasoning | Anthropic | 55.0% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | No tools · high reasoning | Anthropic | 51.3% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | No tools · high reasoning | Anthropic | 38.6% |
| 5 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | No tools · high reasoning | Google | 34.3% |
| 6 | [GPT-6 Sol](/models/gpt-6-sol) | No tools · high reasoning | OpenAI | 33.8% |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | No tools · high reasoning | OpenAI | 31.2% |
| 8 | [Muse Spark 1.3](/models/muse-spark-1-3) | No tools · high reasoning | Meta | 25.4% |
| 9 | [Grok 4.7](/models/grok-4-7) | No tools · high reasoning | xAI | 23.4% |

## FAQ

### What does HLE-Diamond measure?

Accuracy on the 1,000-question HLE-Diamond set without tools, using the benchmark owner's high-reasoning evaluation.

### Which model leads the published HLE-Diamond snapshot?

GPT-6 Astra currently leads the published HLE-Diamond snapshot with a score of 60.6%.

### How many models are evaluated on HLE-Diamond?

The September 22, 2026 release table contains 9 AI models.

### Does HLE-Diamond affect BenchLM's overall score?

Not directly. HLE-Diamond is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
