Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

GDPval-AA normalized (GDPval-AA)

A display-only Artificial Analysis normalized score for economically valuable tasks.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on GDPval-AA — September 15, 2026

We mirror the published score view for GDPval-AA. Claude Fable 5.1 leads the public snapshot at 63.2%, followed by Claude Opus 5 (61.8%) and Muse Spark 1.3 (60.2%). We do not use these results to rank models overall.

118 modelsAgenticCurrentDisplay onlyUpdated September 15, 2026

Benchmark score table (118 models)

Score
1
Claude Fable 5.1Anthropic · Closed
63.2%
2
Claude Opus 5Anthropic · Closed
61.8%
3
Muse Spark 1.3Meta · Closed
60.2%
4
GLM-5.3Z.AI · Open weight
57.8%
5
Qwen3.8-Flash-NextAlibaba · Open weight
57.4%
6
Grok 4.6xAI · Closed
57.1%
7
Claude Fable 5Anthropic · Closed
56.6%
8
DeepSeek V4.1 FlashDeepSeek · Open weight
56.6%
9
Qwen3.8 Max PreviewAlibaba · Closed
56.5%
10
GPT-5.6 SolOpenAI · Closed
56.2%
11
DeepSeek V4 Pro 0813DeepSeek · Closed
54.5%
12
GPT-6 AstraOpenAI · Closed
54.0%
13
Kimi K3Moonshot AI · Closed
52.6%
14
Muse Spark 1.2Meta · Closed
51.2%
15
Claude Sonnet 5Anthropic · Closed
50.0%
16
Claude Opus 4.8Anthropic · Closed
49.5%
17
GPT-5.6 TerraOpenAI · Closed
48.8%
18
GPT-5.6 LunaOpenAI · Closed
48.2%
19
Gemini 3.8 FlashGoogle · Closed
48.2%
20
Qwen3.8-27BAlibaba · Open weight
48.2%
21
DeepSeek V4 Flash 0731DeepSeek · Closed
47.8%
22
Gemini 3.7 FlashGoogle · Closed
46.7%
23
Grok 4.5xAI · Closed
46.5%
24
GLM-5.2Z.AI · Open weight
45.3%
25
GPT-5.5OpenAI · Closed
44.8%
26
Claude Opus 4.7 (Adaptive)Anthropic · Closed
44.8%
27
Gemini 3.5 FlashGoogle · Closed
42.2%
28
Gemini 3.6 FlashGoogle · Closed
41.6%
29
GPT-5.4OpenAI · Closed
40.3%
30
MiniMax M3MiniMax · Open weight
40.2%
31
Muse Spark 1.1Meta · Closed
39.7%
32
DeepSeek V4 Pro (High)DeepSeek · Open weight
39.7%
33
Apodex 1.1Apodex · Closed
38.7%
34
Apodex 1.1 MiniApodex · Open weight
38.7%
35
Quasar 438BMultiverse Computing · Closed
37.0%
36
Ling 3.0 Flash VLInclusionAI · Open weight
36.2%
37
Hy3 PreviewTencent · Open weight
35.8%
38
Inkling-SmallThinking Machines Lab · Open weight
34.6%
39
Qwen3.7 MaxAlibaba · Closed
34.5%
40
MiMo-V2.5-ProXiaomi · Closed
34.3%
41
GLM-5.1Z.AI · Open weight
34.0%
42
Solar Pro 4Upstage · Closed
33.6%
43
InklingThinking Machines Lab · Open weight
33.3%
44
Nemotron 3 UltraNVIDIA · Open weight
33.1%
45
Hy3Tencent · Open weight
31.8%
46
Kimi K2.7 CodeMoonshot AI · Open weight
30.7%
47
Kimi K2.6Moonshot AI · Open weight
29.8%
48
GLM-4.7Z.AI · Open weight
29.8%
49
GPT-5.4 miniOpenAI · Closed
29.7%
50
MiniMax M2.7MiniMax · Open weight
29.4%
51
Grok 4.3xAI · Closed
29.2%
52
Muse SparkMeta · Closed
28.8%
53
Qwen3.6-27BAlibaba · Open weight
28.4%
54
Qwen3.6 PlusAlibaba · Closed
28.3%
55
Gemini 3.5 Flash-LiteGoogle · Closed
28.2%
56
A.X K2SK Telecom · Open weight
27.2%
57
GPT-5.4 nanoOpenAI · Closed
26.8%
58
Ling 3.0 FlashInclusionAI · Open weight
25.9%
59
Ling 3.0 Flash FP8InclusionAI · Open weight
25.9%
60
Step 3.7 FlashStepFun · Open weight
25.8%
61
GPT-5 (high)OpenAI · Closed
25.7%
62
Qwen3.6-35B-A3BAlibaba · Open weight
24.5%
63
Kimi K2.5Moonshot AI · Open weight
21.8%
64
Kimi K2.5 (Reasoning)Moonshot AI · Closed
21.8%
65
GPT-5.1OpenAI · Closed
21.5%
66
Qwen3.5-122B-A10BAlibaba · Open weight
21.3%
67
K-EXAONE 2.0LG AI Research · Open weight
20.9%
68
Gemini 3.1 ProGoogle · Closed
20.2%
69
Muse Glimmer 30BMeta · Open weight
19.6%
70
Qwen3.7 PlusAlibaba · Closed
19.3%
71
GPT-5 miniOpenAI · Closed
19.0%
72
Mistral Medium 3.5 128BMistral · Open weight
18.8%
73
MiniCPM5-2BOpenBMB · Open weight
16.4%
74
MiMo-V2-FlashXiaomi · Open weight
14.4%
75
13.3%
76
Gemma 4 31BGoogle · Open weight
12.8%
77
GPT-OSS 120BOpenAI · Open weight
12.2%
78
Granite 4.2 30BIBM · Open weight
11.1%
79
Ling 3.0 TinyInclusionAI · Open weight
10.8%
80
Gemma 4 26B A4BGoogle · Open weight
10.7%
81
Command A+Cohere · Open weight
7.9%
82
Granite 4.2 8BIBM · Open weight
7.3%
83
Mercury 2Inception · Closed
7.3%
84
Nemotron 3 Super 100BNVIDIA · Open weight
7.2%
85
Nemotron 3 Super 120B A12BNVIDIA · Open weight
7.2%
86
Gemini 2.5 ProGoogle · Closed
5.8%
87
Gemma 4 12BGoogle · Open weight
4.5%
88
Mistral Large 3Mistral · Closed
4.4%
89
K-ExaoneLG AI Research · Closed
2.0%
90
Mistral Small 4Mistral · Open weight
1.8%
91
Mistral Small 4 (Reasoning)Mistral · Open weight
1.8%
92
GPT-OSS 20BOpenAI · Open weight
0.7%
93
Trinity-Large-ThinkingArcee AI · Open weight
0.6%
94
Trinity-Large-PreviewArcee AI · Open weight
0.6%
95
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
0.0%
96
Celeris-1Celeris · Closed
0.0%
97
Solar Pro 3Upstage · Closed
0.0%
98
Llama 4 MaverickMeta · Open weight
0.0%
99
DeepSeek V3DeepSeek · Open weight
0.0%
100
Gemma 4 E4BGoogle · Open weight
0.0%
101
Granite 4.2 3BIBM · Open weight
0.0%
102
Llama 4 ScoutMeta · Open weight
0.0%
103
GPT-4.1 miniOpenAI · Closed
0.0%
104
Ling 2.6 FlashInclusionAI · Open weight
0.0%
105
GPT-4o miniOpenAI · Closed
0.0%
106
LFM2.5-2.6BLiquidAI · Open weight
0.0%
107
Gemma 4 E2BGoogle · Open weight
0.0%
108
GPT-4.1 nanoOpenAI · Closed
0.0%
109
Gemma 3 27BGoogle · Open weight
0.0%
110
Nemotron 3 Nano 30BNVIDIA · Open weight
0.0%
111
North Mini CodeCohere · Open weight
0.0%
112
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
0.0%
113
Ministral 3 14B (Reasoning)Mistral · Open weight
0.0%
114
Ministral 3 14BMistral · Open weight
0.0%
115
Ministral 3 8B (Reasoning)Mistral · Open weight
0.0%
116
Ministral 3 8BMistral · Open weight
0.0%
117
Ministral 3 3B (Reasoning)Mistral · Open weight
0.0%
118
Ministral 3 3BMistral · Open weight
0.0%

The published GDPval-AA snapshot places Claude Fable 5.1 first at 63.2%. The third row is 3.0 points behind. The broader top-10 range is 7.0 points, so many of the published results sit in a relatively narrow band.

118 models have been evaluated on GDPval-AA. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. GDPval-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GDPval-AA

Year

2026

Tasks

Economically valuable tasks

Format

Normalized score

Difficulty

Professional agentic workflows

OpenRouter's Grok 4.3 benchmark card displays GDPval-AA as a normalized percentage. BenchLM stores it separately from the Elo-style GDPval-AA rows used in provider comparison tables.

BenchLM freshness & provenance

Version

GDPval-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GDPval-AA measure?

A display-only Artificial Analysis normalized score for economically valuable tasks.

Which model scores highest on GDPval-AA?

Claude Fable 5.1 by Anthropic currently leads with a score of 63.2% on GDPval-AA.

How many models are evaluated on GDPval-AA?

118 AI models have been evaluated on GDPval-AA on BenchLM.

Last updated: September 15, 2026 · BenchLM version GDPval-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.