Skip to main content
BenchLM

GDPval-AA

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

Benchmark score on GDPval-AA — September 27, 2026

We compile the GDPval-AA rows from provider self-reports and secondary reports. Claude Opus 5 leads the table at 1862, followed by Claude Opus 5.5 (1846) and GLM-5.3-Flash (1773). We do not use these results to rank models overall.

115 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (115 models)

Score
1
Claude Opus 5Anthropic · Closed
1862
2
Claude Opus 5.5Anthropic · Closed
1846
3
GLM-5.3-FlashZ.AI · Open weight
1773
4
GLM-5.3Z.AI · Open weight
1769
5
Muse Spark 1.3Meta · Closed
1754
6
Claude Fable 5Anthropic · Closed
1747
7
Claude Fable 5.1Anthropic · Closed
1735
8
GPT-5.6 SolOpenAI · Closed
1735
9
Grok 4.7xAI · Closed
1695
10
Hy4 previewTencent · Open weight
1678
11
MiMo-V2.6-ProXiaomi · Open weight
1673
12
Qwen3.8-Flash-NextAlibaba · Open weight
1648
13
Muse Spark 1.2Meta · Closed
1631
14
Qwen3.8 Max PreviewAlibaba · Closed
1630
15
Grok 4.6xAI · Closed
1605
16
Claude Sonnet 5Anthropic · Closed
1603
17
DeepSeek V4.1 FlashDeepSeek · Open weight
1600
18
Claude Opus 4.8Anthropic · Closed
1593
19
GPT-5.6 TerraOpenAI · Closed
1583
20
Atria Dawn PreviewShanghai Artificial Intelligence Laboratory · Open weight
1583
21
GPT-5.6 LunaOpenAI · Closed
1582
22
Step 5 PreviewStepFun · Closed
1566
23
Gemini 3.8 FlashGoogle · Closed
1545
24
GPT-6 AstraOpenAI · Closed
1542
25
Gemini 3.7 FlashGoogle · Closed
1525
26
Kimi K3Moonshot AI · Closed
1524
27
Qwen3.8-27BAlibaba · Open weight
1463
28
Grok 4.5xAI · Closed
1430
29
Gemini 3.6 FlashGoogle · Closed
1423
30
GLM-5.2Z.AI · Open weight
1418
31
GPT-5.5OpenAI · Closed
1396
32
Claude Opus 4.7 (Adaptive)Anthropic · Closed
1396
33
Muse Spark 1.1Meta · Closed
1375
34
Gemini 3.5 FlashGoogle · Closed
1345
35
GPT-5.4OpenAI · Closed
1307
36
DeepSeek V4 Pro 0813DeepSeek · Open weight
1306
37
MiniMax M3MiniMax · Open weight
1304
38
DeepSeek V4 Pro (High)DeepSeek · Open weight
1294
39
Apodex 1.1Apodex · Closed
1273
40
Apodex 1.1 MiniApodex · Open weight
1273
41
MiMo-V2.5-ProXiaomi · Closed
1265
42
Quasar 438BMultiverse Computing · Closed
1239
43
Inkling-SmallThinking Machines Lab · Open weight
1191
44
Qwen3.7 MaxAlibaba · Closed
1190
45
DeepSeek V4 Flash 0731DeepSeek · Open weight
1189
46
GLM-5.1Z.AI · Open weight
1181
47
InklingThinking Machines Lab · Open weight
1165
48
Gemini 3.5 Flash-LiteGoogle · Closed
1139
49
Hy3Tencent · Open weight
1136
50
Hy3 PreviewTencent · Open weight
1136
51
Kimi K2.6Moonshot AI · Open weight
1115
52
Kimi K2.7 CodeMoonshot AI · Open weight
1114
53
Ling 3.0 FlashInclusionAI · Open weight
1107
54
GLM-4.7Z.AI · Open weight
1096
55
GPT-5.4 miniOpenAI · Closed
1095
56
Nemotron 3 UltraNVIDIA · Open weight
1091
57
MiniMax M2.7MiniMax · Open weight
1087
58
Muse SparkMeta · Closed
1076
59
Qwen3.6-27BAlibaba · Open weight
1069
60
Qwen3.6 PlusAlibaba · Closed
1066
61
Ling 3.0 Flash FP8InclusionAI · Open weight
1036
62
GPT-5.4 nanoOpenAI · Closed
1035
63
Grok 4.3xAI · Closed
1018
64
GPT-5 (high)OpenAI · Closed
1015
65
Qwen3.6-35B-A3BAlibaba · Open weight
992
66
Step 3.7 FlashStepFun · Open weight
954
67
Kimi K2.5Moonshot AI · Open weight
936
68
Kimi K2.5 (Reasoning)Moonshot AI · Closed
936
69
GPT-5.1OpenAI · Closed
930
70
Qwen3.5-122B-A10BAlibaba · Open weight
925
71
Gemini 3.1 ProGoogle · Closed
904
72
Muse Glimmer 30BMeta · Open weight
893
73
Qwen3.7 PlusAlibaba · Closed
886
74
GPT-5 miniOpenAI · Closed
880
75
Mistral Medium 3.5 128BMistral · Open weight
875
76
865
77
MiMo-V2-FlashXiaomi · Open weight
788
78
Gemma 4 31BGoogle · Open weight
755
79
GPT-OSS 120BOpenAI · Open weight
745
80
Gemma 4 26B A4BGoogle · Open weight
713
81
Command A+Cohere · Open weight
658
82
Granite 4.2 8BIBM · Open weight
646
83
Mercury 2Inception · Closed
646
84
Nemotron 3 Super 100BNVIDIA · Open weight
644
85
Nemotron 3 Super 120B A12BNVIDIA · Open weight
644
86
Gemini 2.5 ProGoogle · Closed
616
87
Gemma 4 12BGoogle · Open weight
591
88
Mistral Large 3Mistral · Closed
588
89
Ling 2.6 FlashInclusionAI · Open weight
550
90
K-ExaoneLG AI Research · Closed
540
91
Mistral Small 4Mistral · Open weight
537
92
Mistral Small 4 (Reasoning)Mistral · Open weight
537
93
GPT-OSS 20BOpenAI · Open weight
514
94
Trinity-Large-ThinkingArcee AI · Open weight
512
95
Trinity-Large-PreviewArcee AI · Open weight
512
96
Celeris-1Celeris · Closed
483
97
GPT-4.1 miniOpenAI · Closed
453
98
Nemotron 3 Nano 30BNVIDIA · Open weight
438
99
Ministral 3 14B (Reasoning)Mistral · Open weight
432
100
Ministral 3 14BMistral · Open weight
432
101
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
416
102
Ministral 3 8B (Reasoning)Mistral · Open weight
405
103
Ministral 3 8BMistral · Open weight
405
104
Ministral 3 3B (Reasoning)Mistral · Open weight
234
105
Ministral 3 3BMistral · Open weight
234
106
LFM2.5-2.6BLiquidAI · Open weight
202
107
GPT-4o miniOpenAI · Closed
188
108
DeepSeek V3DeepSeek · Open weight
183
109
Gemma 4 E4BGoogle · Open weight
177
110
Llama 4 ScoutMeta · Open weight
60
111
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
50
112
Gemma 4 E2BGoogle · Open weight
36
113
GPT-4.1 nanoOpenAI · Closed
13
114
Llama 4 MaverickMeta · Open weight
-45
115
Gemma 3 27BGoogle · Open weight
-171

Among the reported GDPval-AA rows, Claude Opus 5 is first at 1862. The third row is 89 score units behind. The broader top-10 range is 184 score units, so the table still separates the published systems.

115 models have been evaluated on GDPval-AA. The benchmark falls in the Agentic category. GDPval-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GDPval-AA

Year

2026

Tasks

Agentic real-world work tasks

Format

Elo

Difficulty

Professional agentic workflows

BenchLM stores GDPval-AA as a display-only provider-table row for DeepSeek-V4 because the source reports an Elo score rather than a 0-100 percentage.

Freshness and provenance

Version

GDPval-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does GDPval-AA measure?

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

Which model scores highest on GDPval-AA?

Claude Opus 5 by Anthropic currently leads with a score of 1862 on GDPval-AA.

How many models are evaluated on GDPval-AA?

115 AI models have been evaluated on GDPval-AA on BenchLM.

Last updated: September 27, 2026 · BenchLM version GDPval-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.