# Scale Labs Humanity&#x27;s Last Exam (Humanity&#x27;s Last Exam)

> A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

Canonical page: https://benchlm.ai/benchmarks/scale-hle-diamond

- Category: [Knowledge](/knowledge)
- Last updated: September 29, 2026 snapshot

## About Humanity&#x27;s Last Exam

- Year: 2026
- Tasks: 12 published rows
- Format: Published Scale leaderboard score
- Difficulty: External agent and model evaluation
- Paper: [Scale Labs leaderboard](https://labs.scale.com/leaderboard/hle-diamond)

BenchLM mirrors 12 published rows from the Humanity&#x27;s Last Exam public table captured on September 29, 2026 snapshot.

Humanity&#x27;s Last Exam is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (12 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 60.60% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 55.00% |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 51.30% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 38.60% |
| 5 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 34.30% |
| 6 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 33.80% |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 31.20% |
| 8 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 25.40% |
| 9 | [Grok 4.7](/models/grok-4-7) | xAI | 23.40% |
| 10 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 22.20% |
| 11 | [GLM-5.3](/models/glm-5-3) | Z.AI | 16.40% |
| 12 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 13.40% |

## FAQ

### What does Humanity&#x27;s Last Exam measure?

A Scale Labs public leaderboard mirrored as display-only reference data. It does not affect BenchLM rankings.

### Which model leads the published Humanity&#x27;s Last Exam snapshot?

GPT-6 Astra currently leads the published Humanity&#x27;s Last Exam snapshot with a score of 60.60%.

### How many models are evaluated on Humanity&#x27;s Last Exam?

The September 29, 2026 snapshot contains 12 AI models.

### Does Humanity&#x27;s Last Exam affect BenchLM's overall score?

Not directly. Humanity&#x27;s Last Exam is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
