Skip to main content
BenchLM

Agents' Last Exam

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Top models on Agents' Last Exam — September 27, 2026

As of September 27, 2026, GPT-6 Astra leads the Agents' Last Exam leaderboard with 59.3% , followed by GPT-6 Sol (56.4%) and Qwen3.8 Max (52.4%).

15 modelsAgentic3% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (15 models)

Score
1
GPT-6 AstraOpenAI · Closed
59.3%
2
GPT-6 SolOpenAI · Closed
56.4%
3
Qwen3.8 MaxAlibaba · Open weight
52.4%
4
Qwen3.8-Flash-NextAlibaba · Open weight
51.2%
5
Qwen3.8-27BAlibaba · Open weight
42.9%
6
DeepSeek V4.1 FlashDeepSeek · Open weight
31.8%
7
MiMo-V2.6-ProXiaomi · Open weight
31.6%
8
Step 5 PreviewStepFun · Closed
29.5%
9
GLM-5.3Z.AI · Open weight
28.5%
10
MiMo-V2.6-FlashXiaomi · Open weight
27.6%
11
Gemini 3.7 FlashGoogle · Closed
26.3%
12
GLM-5.3-FlashZ.AI · Open weight
26.3%
13
DeepSeek V4 Pro 0813DeepSeek · Open weight
25.7%
14
DeepSeek V4 Flash 0731DeepSeek · Open weight
25.2%
15
Hy4 previewTencent · Open weight
22.8%

According to BenchLM.ai, GPT-6 Astra leads the Agents' Last Exam benchmark with a score of 59.3%, followed by GPT-6 Sol (56.4%) and Qwen3.8 Max (52.4%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

15 models have been evaluated on Agents' Last Exam. The benchmark falls in the Agentic category. BenchAlign v5.7 gives Agents' Last Exam 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About Agents' Last Exam

Year

2026

Tasks

Agent tasks

Format

Provider-reported task score

Difficulty

Advanced agentic work

BenchLM stores provider-run values, such as DeepSeek's max-effort launch result, under the exact published benchmark label.

Freshness and provenance

Version

Agents' Last Exam 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Agents' Last Exam measure?

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Which model scores highest on Agents' Last Exam?

GPT-6 Astra by OpenAI currently leads with a score of 59.3% on Agents' Last Exam.

How many models are evaluated on Agents' Last Exam?

15 AI models have been evaluated on Agents' Last Exam on BenchLM.

Last updated: September 27, 2026 · BenchLM version Agents' Last Exam 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.