Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

GDPval-AA

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

Data verified 25 confirmed releases in the last 30 daysStart the free Radar Brief

Benchmark score on GDPval-AA — August 29, 2026

We mirror the published score view for GDPval-AA. Claude Opus 5 leads the public snapshot at 1862, followed by GLM-5.3-Flash (1773) and GLM-5.3 (1769). We do not use these results to rank models overall.

103 modelsAgenticCurrentDisplay onlyUpdated August 29, 2026

Benchmark score table (103 models)

Score
1
Claude Opus 5Anthropic · Closed
1862
2
GLM-5.3-FlashZ.AI · Open weight
1773
3
GLM-5.3Z.AI · Open weight
1769
4
Claude Fable 5Anthropic · Closed
1747
5
Qwen3.8-Flash-NextAlibaba · Open weight
1739
6
GPT-5.6 SolOpenAI · Closed
1735
7
Grok 4.6xAI · Closed
1730
8
Qwen3.8 Max PreviewAlibaba · Closed
1724
9
Hy4 previewTencent · Open weight
1678
10
Kimi K3Moonshot AI · Closed
1668
11
Muse Spark 1.2Meta · Closed
1631
12
Claude Sonnet 5Anthropic · Closed
1603
13
Claude Opus 4.8Anthropic · Closed
1593
14
GPT-5.6 TerraOpenAI · Closed
1583
15
GPT-5.6 LunaOpenAI · Closed
1582
16
Qwen3.8-27BAlibaba · Open weight
1543
17
Grok 4.5xAI · Closed
1518
18
Gemini 3.7 FlashGoogle · Closed
1516
19
GLM-5.2Z.AI · Open weight
1498
20
GPT-5.5OpenAI · Closed
1485
21
Claude Opus 4.7 (Adaptive)Anthropic · Closed
1483
22
Gemini 3.6 FlashGoogle · Closed
1423
23
GPT-5.4OpenAI · Closed
1388
24
MiniMax M3MiniMax · Open weight
1380
25
Muse Spark 1.1Meta · Closed
1375
26
Gemini 3.5 FlashGoogle · Closed
1345
27
DeepSeek V4 Pro 0813DeepSeek · Closed
1306
28
DeepSeek V4 Pro (High)DeepSeek · Open weight
1294
29
Inkling-SmallThinking Machines Lab · Open weight
1269
30
Qwen3.7 MaxAlibaba · Closed
1267
31
MiMo-V2.5-ProXiaomi · Closed
1265
32
GLM-5.1Z.AI · Open weight
1256
33
InklingThinking Machines Lab · Open weight
1234
34
Hy3 PreviewTencent · Open weight
1213
35
Hy3Tencent · Open weight
1213
36
Kimi K2.6Moonshot AI · Open weight
1191
37
DeepSeek V4 Flash 0731DeepSeek · Closed
1189
38
Kimi K2.7 CodeMoonshot AI · Open weight
1188
39
GPT-5.4 miniOpenAI · Closed
1169
40
GLM-4.7Z.AI · Open weight
1168
41
Nemotron 3 UltraNVIDIA · Open weight
1162
42
MiniMax M2.7MiniMax · Open weight
1158
43
Muse SparkMeta · Closed
1146
44
Qwen3.6-27BAlibaba · Open weight
1139
45
Gemini 3.5 Flash-LiteGoogle · Closed
1139
46
Qwen3.6 PlusAlibaba · Closed
1138
47
Ling 3.0 FlashInclusionAI · Open weight
1107
48
Ling 3.0 Flash FP8InclusionAI · Open weight
1106
49
GPT-5.4 nanoOpenAI · Closed
1105
50
Grok 4.3xAI · Closed
1087
51
GPT-5 (high)OpenAI · Closed
1083
52
Qwen3.6-35B-A3BAlibaba · Open weight
1057
53
Step 3.7 FlashStepFun · Open weight
1019
54
Kimi K2.5Moonshot AI · Open weight
1006
55
Kimi K2.5 (Reasoning)Moonshot AI · Closed
1006
56
GPT-5.1OpenAI · Closed
996
57
Qwen3.5-122B-A10BAlibaba · Open weight
987
58
Qwen3.5 397BAlibaba · Open weight
966
59
Qwen3.5 397B (Reasoning)Alibaba · Open weight
966
60
Gemini 3.1 ProGoogle · Closed
965
61
Muse Glimmer 30BMeta · Open weight
955
62
Qwen3.7 PlusAlibaba · Closed
948
63
GPT-5 miniOpenAI · Closed
942
64
Mistral Medium 3.5 128BMistral · Open weight
936
65
865
66
MiMo-V2-FlashXiaomi · Open weight
844
67
Gemma 4 31BGoogle · Open weight
812
68
GPT-OSS 120BOpenAI · Open weight
802
69
Gemma 4 26B A4BGoogle · Open weight
769
70
Command A+Cohere · Open weight
715
71
Nemotron 3 Super 100BNVIDIA · Open weight
700
72
Mercury 2Inception · Closed
700
73
Nemotron 3 Super 120B A12BNVIDIA · Open weight
700
74
Gemini 2.5 ProGoogle · Closed
671
75
Gemma 4 12BGoogle · Open weight
646
76
Mistral Large 3Mistral · Closed
642
77
K-ExaoneLG AI Research · Closed
590
78
Mistral Small 4Mistral · Open weight
589
79
Mistral Small 4 (Reasoning)Mistral · Open weight
589
80
GPT-OSS 20BOpenAI · Open weight
567
81
Trinity-Large-PreviewArcee AI · Open weight
565
82
Trinity-Large-ThinkingArcee AI · Open weight
565
83
Ling 2.6 FlashInclusionAI · Open weight
550
84
Celeris-1Celeris · Closed
536
85
GPT-4.1 miniOpenAI · Closed
507
86
Nemotron 3 Nano 30BNVIDIA · Open weight
492
87
Ministral 3 14B (Reasoning)Mistral · Open weight
485
88
Ministral 3 14BMistral · Open weight
485
89
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
468
90
Ministral 3 8B (Reasoning)Mistral · Open weight
456
91
Ministral 3 8BMistral · Open weight
456
92
Ministral 3 3B (Reasoning)Mistral · Open weight
285
93
Ministral 3 3BMistral · Open weight
285
94
LFM2.5-2.6BLiquidAI · Open weight
253
95
GPT-4o miniOpenAI · Closed
239
96
DeepSeek V3DeepSeek · Open weight
233
97
Gemma 4 E4BGoogle · Open weight
230
98
Llama 4 ScoutMeta · Open weight
110
99
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
98
100
Gemma 4 E2BGoogle · Open weight
87
101
GPT-4.1 nanoOpenAI · Closed
62
102
Llama 4 MaverickMeta · Open weight
5
103
Gemma 3 27BGoogle · Open weight
-121

The published GDPval-AA snapshot places Claude Opus 5 first at 1862. The third row is 93 score units behind. The broader top-10 range is 194 score units, so the table still separates the published systems.

103 models have been evaluated on GDPval-AA. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. GDPval-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GDPval-AA

Year

2026

Tasks

Agentic real-world work tasks

Format

Elo

Difficulty

Professional agentic workflows

BenchLM stores GDPval-AA as a display-only provider-table row for DeepSeek-V4 because the source reports an Elo score rather than a 0-100 percentage.

BenchLM freshness & provenance

Version

GDPval-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GDPval-AA measure?

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

Which model scores highest on GDPval-AA?

Claude Opus 5 by Anthropic currently leads with a score of 1862 on GDPval-AA.

How many models are evaluated on GDPval-AA?

103 AI models have been evaluated on GDPval-AA on BenchLM.

Last updated: August 29, 2026 · BenchLM version GDPval-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.