Skip to main content

Benchmark profile

Agents' Last Exam

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Data verified

Benchmark score on Agents' Last Exam — July 31, 2026

BenchLM mirrors the published score view for Agents' Last Exam. DeepSeek V4 Flash (Max) leads the public snapshot at 25.2%. BenchLM does not use these results to rank models overall.

1 modelAgenticCurrentDisplay onlyUpdated July 31, 2026

Benchmark score table (1 model)

Score
1
DeepSeek V4 Flash (Max)DeepSeek · Closed
25.2%

About Agents' Last Exam

Year

2026

Tasks

Agent tasks

Format

Provider-reported task score

Difficulty

Advanced agentic work

BenchLM stores DeepSeek's max-effort, provider-run value as display-only launch evidence under the exact published benchmark label.

BenchLM freshness & provenance

Version

Agents' Last Exam 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Agents' Last Exam measure?

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Which model scores highest on Agents' Last Exam?

DeepSeek V4 Flash (Max) by DeepSeek currently leads with a score of 25.2% on Agents' Last Exam.

How many models are evaluated on Agents' Last Exam?

1 AI models have been evaluated on Agents' Last Exam on BenchLM.

Last updated: July 31, 2026 · BenchLM version Agents' Last Exam 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.