Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Artificial Analysis Coding Index (AA Coding Index)

A display-only Artificial Analysis coding index.

Data verified 32 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on AA Coding Index — September 14, 2026

We mirror the published score view for AA Coding Index. Claude Fable 5.1 leads the public snapshot at 81.6%, followed by Claude Opus 5 (78.0%) and GPT-5.6 Sol (77.4%). We do not use these results to rank models overall.

112 modelsCodingCurrentDisplay onlyUpdated September 14, 2026

Benchmark score table (112 models)

Score
1
Claude Fable 5.1Anthropic · Closed
81.6%
2
Claude Opus 5Anthropic · Closed
78.0%
3
GPT-5.6 SolOpenAI · Closed
77.4%
4
GPT-6 AstraOpenAI · Closed
76.9%
5
Grok 4.6xAI · Closed
76.8%
6
GPT-5.6 TerraOpenAI · Closed
76.7%
7
Claude Fable 5Anthropic · Closed
76.5%
8
Gemini 3.8 FlashGoogle · Closed
76.3%
9
Kimi K3Moonshot AI · Closed
76.2%
10
Gemini 3.7 FlashGoogle · Closed
76.1%
11
Muse Spark 1.3Meta · Closed
75.8%
12
GPT-5.5OpenAI · Closed
74.9%
13
GLM-5.3Z.AI · Open weight
74.8%
14
Claude Opus 4.8Anthropic · Closed
74.3%
15
Claude Opus 4.7 (Adaptive)Anthropic · Closed
73.6%
16
Qwen3.8-Flash-NextAlibaba · Open weight
73.0%
17
Grok 4.5xAI · Closed
72.5%
18
Muse Spark 1.2Meta · Closed
72.2%
19
Qwen3.8 Max PreviewAlibaba · Closed
71.8%
20
Claude Sonnet 5Anthropic · Closed
71.5%
21
GPT-5.6 LunaOpenAI · Closed
71.5%
22
Muse Spark 1.1Meta · Closed
71.3%
23
GPT-5.4OpenAI · Closed
71.0%
24
Gemini 3.5 FlashGoogle · Closed
70.1%
25
Gemini 3.6 FlashGoogle · Closed
69.2%
26
DeepSeek V4 Flash 0731DeepSeek · Closed
69.1%
27
Gemini 3.1 ProGoogle · Closed
68.8%
28
DeepSeek V4 Pro 0813DeepSeek · Closed
68.8%
29
GLM-5.2Z.AI · Open weight
68.8%
30
Qwen3.8-27BAlibaba · Open weight
68.1%
31
Qwen3.7 MaxAlibaba · Closed
66.0%
32
Kimi K2.6Moonshot AI · Open weight
61.8%
33
Quasar 438BMultiverse Computing · Closed
61.2%
34
Kimi K2.7 CodeMoonshot AI · Open weight
60.8%
35
Apodex 1.1Apodex · Closed
60.8%
36
Apodex 1.1 MiniApodex · Open weight
60.8%
37
MiMo-V2.5-ProXiaomi · Closed
60.2%
38
Hy3Tencent · Open weight
58.8%
39
Hy3 PreviewTencent · Open weight
58.8%
40
DeepSeek V4 Pro (High)DeepSeek · Open weight
58.7%
41
Muse SparkMeta · Closed
58.6%
42
MiniMax M3MiniMax · Open weight
58.6%
43
GPT-5.4 miniOpenAI · Closed
56.1%
44
GPT-5.4 nanoOpenAI · Closed
56.1%
45
Qwen3.7 PlusAlibaba · Closed
55.9%
46
GLM-5.1Z.AI · Open weight
55.8%
47
Qwen3.6 PlusAlibaba · Closed
54.5%
48
Qwen3.6-27BAlibaba · Open weight
53.7%
49
Inkling-SmallThinking Machines Lab · Open weight
53.0%
50
MiniMax M2.7MiniMax · Open weight
52.6%
51
InklingThinking Machines Lab · Open weight
52.1%
52
Ling 3.0 FlashInclusionAI · Open weight
50.6%
53
Ling 3.0 Flash FP8InclusionAI · Open weight
50.6%
54
MiMo-V2-FlashXiaomi · Open weight
49.8%
55
GPT-5.1OpenAI · Closed
49.4%
56
Gemini 3.5 Flash-LiteGoogle · Closed
49.3%
57
Nemotron 3 UltraNVIDIA · Open weight
49.3%
58
Muse Glimmer 30BMeta · Open weight
49.0%
59
Mistral Medium 3.5 128BMistral · Open weight
46.9%
60
Kimi K2.5Moonshot AI · Open weight
46.8%
61
Kimi K2.5 (Reasoning)Moonshot AI · Closed
46.8%
62
Qwen3.5-122B-A10BAlibaba · Open weight
45.7%
63
GLM-4.7Z.AI · Open weight
45.3%
64
Gemma 4 31BGoogle · Open weight
43.4%
65
Grok 4.3xAI · Closed
42.3%
66
Qwen3.6-35B-A3BAlibaba · Open weight
41.9%
67
o1OpenAI · Closed
39.7%
68
Step 3.7 FlashStepFun · Open weight
39.6%
69
Gemma 4 26B A4BGoogle · Open weight
39.3%
70
GPT-5 (high)OpenAI · Closed
37.8%
71
Nemotron 3 Super 100BNVIDIA · Open weight
37.7%
72
Nemotron 3 Super 120B A12BNVIDIA · Open weight
37.7%
73
o1-previewOpenAI · Closed
34.0%
74
Gemini 2.5 ProGoogle · Closed
33.3%
75
K-ExaoneLG AI Research · Closed
32.1%
76
Mercury 2Inception · Closed
31.1%
77
Gemma 4 12BGoogle · Open weight
31.0%
78
GPT-OSS 120BOpenAI · Open weight
30.4%
79
Command A+Cohere · Open weight
27.9%
80
26.8%
81
Mistral Small 4Mistral · Open weight
26.6%
82
Mistral Small 4 (Reasoning)Mistral · Open weight
26.6%
83
Trinity-Large-ThinkingArcee AI · Open weight
25.8%
84
Trinity-Large-PreviewArcee AI · Open weight
25.8%
85
Ling 2.6 FlashInclusionAI · Open weight
25.3%
86
Gemini 1.5 ProGoogle · Closed
23.6%
87
DeepSeek V3DeepSeek · Open weight
23.0%
88
Granite 4.2 8BIBM · Open weight
22.4%
89
GPT-4 TurboOpenAI · Closed
21.5%
90
GPT-OSS 20BOpenAI · Open weight
20.7%
91
GPT-4.1 miniOpenAI · Closed
20.2%
92
Mistral Large 3Mistral · Closed
20.1%
93
Claude 3 OpusAnthropic · Closed
19.5%
94
Llama 4 MaverickMeta · Open weight
16.3%
95
GPT-5 miniOpenAI · Closed
15.6%
96
Celeris-1Celeris · Closed
14.4%
97
Nemotron 3 Nano 30BNVIDIA · Open weight
14.4%
98
Ministral 3 14B (Reasoning)Mistral · Open weight
14.4%
99
Ministral 3 14BMistral · Open weight
14.4%
100
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
13.8%
101
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
11.9%
102
GPT-4o miniOpenAI · Closed
11.4%
103
GPT-4.1 nanoOpenAI · Closed
11.1%
104
Gemma 3 27BGoogle · Open weight
10.1%
105
Ministral 3 8B (Reasoning)Mistral · Open weight
9.7%
106
Ministral 3 8BMistral · Open weight
9.7%
107
Gemma 4 E4BGoogle · Open weight
9.4%
108
Llama 4 ScoutMeta · Open weight
8.2%
109
LFM2.5-2.6BLiquidAI · Open weight
7.7%
110
Gemma 4 E2BGoogle · Open weight
7.2%
111
Ministral 3 3B (Reasoning)Mistral · Open weight
4.8%
112
Ministral 3 3BMistral · Open weight
4.8%

The published AA Coding Index snapshot places Claude Fable 5.1 first at 81.6%. The third row is 4.2 points behind. The broader top-10 range is 5.5 points, so many of the published results sit in a relatively narrow band.

112 models have been evaluated on AA Coding Index. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. AA Coding Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Coding Index

Year

2026

Tasks

Cross-benchmark coding index

Format

Aggregated model score

Difficulty

Display-only external reference

BenchLM mirrors this coding index for comparison, but does not use it as a weighted coding benchmark row.

BenchLM freshness & provenance

Version

AA Coding Index 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA Coding Index measure?

A display-only Artificial Analysis coding index.

Which model scores highest on AA Coding Index?

Claude Fable 5.1 by Anthropic currently leads with a score of 81.6% on AA Coding Index.

How many models are evaluated on AA Coding Index?

112 AI models have been evaluated on AA Coding Index on BenchLM.

Last updated: September 14, 2026 · BenchLM version AA Coding Index 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.