Skip to main content

Benchmark profile

GDPval-AA normalized (GDPval-AA)

A display-only Artificial Analysis normalized score for economically valuable tasks.

Data verified

Benchmark score on GDPval-AA — July 29, 2026

BenchLM mirrors the published score view for GDPval-AA. Claude Opus 5 leads the public snapshot at 68.1% , followed by Claude Fable 5 (62.3%) and GPT-5.6 Sol (61.8%). BenchLM does not use these results to rank models overall.

87 modelsAgenticCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (87 models)

Score
1
Claude Opus 5Anthropic · Closed
68.1%
2
Claude Fable 5Anthropic · Closed
62.3%
3
GPT-5.6 SolOpenAI · Closed
61.8%
4
Kimi K3Moonshot AI · Closed
59.4%
5
Claude Sonnet 5Anthropic · Closed
55.2%
6
Claude Opus 4.8Anthropic · Closed
54.6%
7
GPT-5.6 TerraOpenAI · Closed
54.1%
8
GPT-5.6 LunaOpenAI · Closed
54.1%
9
Grok 4.5xAI · Closed
51.4%
10
GLM-5.2Z.AI · Open weight
50.5%
11
Claude Opus 4.7 (Adaptive)Anthropic · Closed
49.7%
12
GPT-5.5OpenAI · Closed
49.6%
13
Gemini 3.6 FlashGoogle · Closed
46.2%
14
GPT-5.4OpenAI · Closed
44.6%
15
MiniMax M3MiniMax · Open weight
44.5%
16
Muse Spark 1.1Meta · Closed
43.8%
17
Gemini 3.5 FlashGoogle · Closed
42.2%
18
DeepSeek V4 Pro (Max)DeepSeek · Open weight
40.3%
19
DeepSeek V4 Pro (High)DeepSeek · Open weight
40.0%
20
Qwen3.7 MaxAlibaba · Closed
38.6%
21
MiMo-V2.5-ProXiaomi · Closed
38.3%
22
GLM-5.1Z.AI · Open weight
37.8%
23
InklingThinking Machines Lab · Open weight
36.8%
24
Hy3 PreviewTencent · Open weight
35.8%
25
Hy3Tencent · Open weight
35.8%
26
Kimi K2.6Moonshot AI · Open weight
34.4%
27
DeepSeek V4 Flash (Max)DeepSeek · Open weight
34.4%
28
Kimi K2.7 CodeMoonshot AI · Open weight
34.3%
29
GPT-5.4 miniOpenAI · Closed
33.5%
30
GLM-4.7Z.AI · Open weight
33.3%
31
Nemotron 3 UltraNVIDIA · Open weight
33.1%
32
MiniMax M2.7MiniMax · Open weight
32.9%
33
DeepSeek V4 Flash (High)DeepSeek · Open weight
32.4%
34
Muse SparkMeta · Closed
32.2%
35
Qwen3.6 PlusAlibaba · Closed
31.9%
36
Qwen3.6-27BAlibaba · Open weight
31.9%
37
Gemini 3.5 Flash-LiteGoogle · Closed
31.9%
38
GPT-5.4 nanoOpenAI · Closed
30.1%
39
Grok 4.3xAI · Closed
29.2%
40
GPT-5 (high)OpenAI · Closed
28.9%
41
Qwen3.6-35B-A3BAlibaba · Open weight
27.6%
42
Step 3.7 FlashStepFun · Open weight
25.8%
43
Kimi K2.5Moonshot AI · Open weight
25.1%
44
Kimi K2.5 (Reasoning)Moonshot AI · Closed
25.1%
45
GPT-5.1OpenAI · Closed
24.4%
46
Qwen3.5-122B-A10BAlibaba · Open weight
24.1%
47
Gemini 3.1 ProGoogle · Closed
23.2%
48
Qwen3.5 397BAlibaba · Open weight
23.1%
49
Qwen3.5 397B (Reasoning)Alibaba · Open weight
23.1%
50
Qwen3.7 PlusAlibaba · Closed
22.1%
51
Mistral Medium 3.5 128BMistral · Open weight
21.6%
52
GPT-5 miniOpenAI · Closed
21.6%
53
MiMo-V2-FlashXiaomi · Open weight
16.9%
54
Gemma 4 31BGoogle · Open weight
15.5%
55
GPT-OSS 120BOpenAI · Open weight
15.1%
56
Gemma 4 26B A4BGoogle · Open weight
13.5%
57
Command A+Cohere · Open weight
10.9%
58
Mercury 2Inception · Closed
9.9%
59
Nemotron 3 Super 120B A12BNVIDIA · Open weight
9.9%
60
Gemini 2.5 ProGoogle · Closed
8.6%
61
Gemma 4 12BGoogle · Open weight
7.6%
62
Gemini 3.1 Flash-LiteGoogle · Closed
7.3%
63
Mistral Large 3Mistral · Closed
7.0%
64
K-ExaoneLG AI Research · Closed
4.9%
65
Mistral Small 4Mistral · Open weight
4.6%
66
Mistral Small 4 (Reasoning)Mistral · Open weight
4.6%
67
GPT-OSS 20BOpenAI · Open weight
3.4%
68
Trinity-Large-PreviewArcee AI · Open weight
3.2%
69
Trinity-Large-ThinkingArcee AI · Open weight
3.2%
70
Ling 2.6 FlashInclusionAI · Open weight
2.5%
71
GPT-4.1 miniOpenAI · Closed
0.4%
72
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
0.0%
73
GPT-4.1 nanoOpenAI · Closed
0.0%
74
DeepSeek V3DeepSeek · Open weight
0.0%
75
Llama 4 ScoutMeta · Open weight
0.0%
76
Llama 4 MaverickMeta · Open weight
0.0%
77
Gemma 3 27BGoogle · Open weight
0.0%
78
Nemotron 3 Nano 30BNVIDIA · Open weight
0.0%
79
GPT-4o miniOpenAI · Closed
0.0%
80
Gemma 4 E2BGoogle · Open weight
0.0%
81
Gemma 4 E4BGoogle · Open weight
0.0%
82
Ministral 3 14B (Reasoning)Mistral · Open weight
0.0%
83
Ministral 3 14BMistral · Open weight
0.0%
84
Ministral 3 8B (Reasoning)Mistral · Open weight
0.0%
85
Ministral 3 8BMistral · Open weight
0.0%
86
Ministral 3 3B (Reasoning)Mistral · Open weight
0.0%
87
Ministral 3 3BMistral · Open weight
0.0%

The published GDPval-AA snapshot places Claude Opus 5 first at 68.1%. The third row is 6.3 points behind. The broader top-10 range is 17.6 points, so the table still separates the published systems.

87 models have been evaluated on GDPval-AA. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. GDPval-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GDPval-AA

Year

2026

Tasks

Economically valuable tasks

Format

Normalized score

Difficulty

Professional agentic workflows

OpenRouter's Grok 4.3 benchmark card displays GDPval-AA as a normalized percentage. BenchLM stores it separately from the Elo-style GDPval-AA rows used in provider comparison tables.

BenchLM freshness & provenance

Version

GDPval-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GDPval-AA measure?

A display-only Artificial Analysis normalized score for economically valuable tasks.

Which model scores highest on GDPval-AA?

Claude Opus 5 by Anthropic currently leads with a score of 68.1% on GDPval-AA.

How many models are evaluated on GDPval-AA?

87 AI models have been evaluated on GDPval-AA on BenchLM.

Last updated: July 29, 2026 · BenchLM version GDPval-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.