# Humanity's Last Exam Diamond without tools at highest reasoning effort (HLE-Diamond (highest effort))

> Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.

Canonical page: https://benchlm.ai/benchmarks/hle-diamond-max

- Category: [Knowledge](/knowledge)
- Last updated: September 22, 2026 release table

## About HLE-Diamond (highest effort)

- Year: 2026
- Tasks: 1,000 expert questions: 500 reasoning and 500 knowledge
- Format: Accuracy without tools, highest available reasoning effort
- Difficulty: Frontier expert knowledge and reasoning
- Paper: [Introducing HLE-Diamond](https://lastexam.ai/blog/hle-diamond)

The release offers this nine-model table behind its reasoning-max view. GPT, Claude, and Muse use max; Gemini uses high; Grok uses xhigh, the highest available settings listed by the benchmark owner.

HLE-Diamond (highest effort) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (9 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | No tools · max reasoning | OpenAI | 66.2% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | No tools · max reasoning | Anthropic | 61.2% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | No tools · max reasoning | Anthropic | 54.1% |
| 4 | [GPT-6 Sol](/models/gpt-6-sol) | No tools · max reasoning | OpenAI | 43.8% |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | No tools · max reasoning | OpenAI | 41.6% |
| 6 | [Claude Opus 5](/models/claude-opus-5) | No tools · max reasoning | Anthropic | 41.3% |
| 7 | [Muse Spark 1.3](/models/muse-spark-1-3) | No tools · max reasoning | Meta | 34.8% |
| 8 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | No tools · high reasoning | Google | 34.3% |
| 9 | [Grok 4.7](/models/grok-4-7) | No tools · xhigh reasoning | xAI | 25.3% |

## FAQ

### What does HLE-Diamond (highest effort) measure?

Accuracy on the 1,000-question HLE-Diamond set without tools at each model's highest available reasoning effort.

### Which model leads the published HLE-Diamond (highest effort) snapshot?

GPT-6 Astra currently leads the published HLE-Diamond (highest effort) snapshot with a score of 66.2%.

### How many models are evaluated on HLE-Diamond (highest effort)?

The September 22, 2026 release table contains 9 AI models.

### Does HLE-Diamond (highest effort) affect BenchLM's overall score?

Not directly. HLE-Diamond (highest effort) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
