# Agents' Last Exam

> An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Canonical page: https://benchlm.ai/benchmarks/agentslastexam

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Agents' Last Exam

- Year: 2026
- Tasks: Agent tasks
- Format: Provider-reported task score
- Difficulty: Advanced agentic work
- Paper: [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/)

BenchLM stores provider-run values, such as DeepSeek's max-effort launch result, under the exact published benchmark label.

BenchAlign v5.7 gives Agents' Last Exam 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (15 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 59.3% |
| 2 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 56.4% |
| 3 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 52.4% |
| 4 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 51.2% |
| 5 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 42.9% |
| 6 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 31.8% |
| 7 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | Xiaomi | 31.6% |
| 8 | [Step 5 Preview](/models/step-5-preview) | StepFun | 29.5% |
| 9 | [GLM-5.3](/models/glm-5-3) | Z.AI | 28.5% |
| 10 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | Xiaomi | 27.6% |
| 11 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 26.3% |
| 12 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 26.3% |
| 13 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 25.7% |
| 14 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 25.2% |
| 15 | [Hy4 preview](/models/hy4-preview) | Tencent | 22.8% |

## FAQ

### What does Agents' Last Exam measure?

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

### Which model scores highest on Agents' Last Exam?

GPT-6 Astra by OpenAI currently leads with a score of 59.3% on Agents' Last Exam.

### How many models are evaluated on Agents' Last Exam?

15 AI models have been evaluated on Agents' Last Exam on BenchLM.

### Does Agents' Last Exam affect BenchLM's overall score?

Yes. BenchAlign v5.7 gives Agents' Last Exam 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on Agents' Last Exam

- [GPT-6 Astra vs GPT-6 Sol](/compare/gpt-6-astra-vs-gpt-6-sol)
- [GPT-6 Sol vs Qwen3.8 Max](/compare/gpt-6-sol-vs-qwen3-8-max)
- [Qwen3.8 Max vs Qwen3.8-Flash-Next](/compare/qwen3-8-flash-next-vs-qwen3-8-max)
- [Qwen3.8-Flash-Next vs Qwen3.8-27B](/compare/qwen3-8-27b-vs-qwen3-8-flash-next)
