Skip to main content
BenchLM

Artificial Analysis SciCode (AA-SciCode)

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An Artificial Analysis SciCode score.

Top models on AA-SciCode — September 27, 2026

As of September 27, 2026, Claude Opus 5.5 leads the AA-SciCode leaderboard with 66.9% , followed by Claude Fable 5.1 (63.1%) and Claude Fable 5 (61.0%).

104 modelsCoding5% of Coding reference weightCurrentUpdated September 27, 2026

Leaderboard (104 models)

Score
1
Claude Opus 5.5Anthropic · Closed
66.9%
2
Claude Fable 5.1Anthropic · Closed
63.1%
3
Claude Fable 5Anthropic · Closed
61.0%
4
MiMo-V2.6-ProXiaomi · Open weight
60.9%
5
Kimi K3Moonshot AI · Closed
59.5%
6
GLM-5.3Z.AI · Open weight
59.0%
7
Step 5 PreviewStepFun · Closed
58.9%
8
Muse Spark 1.1Meta · Closed
58.8%
9
Muse Spark 1.3Meta · Closed
58.8%
10
Gemini 3.1 ProGoogle · Closed
58.7%
11
GPT-6 SolOpenAI · Closed
57.6%
12
Muse Spark 1.2Meta · Closed
57.4%
13
Grok 4.7xAI · Closed
57.4%
14
Gemini 3.7 FlashGoogle · Closed
57.2%
15
GPT-5.6 SolOpenAI · Closed
57.1%
16
Gemini 3.8 FlashGoogle · Closed
56.6%
17
GPT-6 AstraOpenAI · Closed
56.5%
18
Grok 4.6xAI · Closed
56.5%
19
Claude Opus 5Anthropic · Closed
56.4%
20
GPT-5.5OpenAI · Closed
55.8%
21
GPT-5.6 TerraOpenAI · Closed
55.0%
22
Grok 4.5xAI · Closed
55.0%
23
GPT-6 LunaOpenAI · Closed
54.6%
24
Claude Opus 4.8Anthropic · Closed
54.4%
25
Claude Sonnet 5Anthropic · Closed
54.3%
26
Gemini 3.5 FlashGoogle · Closed
53.9%
27
GPT-5.6 LunaOpenAI · Closed
53.6%
28
Gemini 3.6 FlashGoogle · Closed
53.4%
29
GPT-5.4 miniOpenAI · Closed
52.1%
30
Qwen3.8 Max PreviewAlibaba · Closed
52.1%
31
DeepSeek V4.1 FlashDeepSeek · Open weight
51.9%
32
Kimi K2.6Moonshot AI · Open weight
51.5%
33
GLM-5.2Z.AI · Open weight
51.2%
34
DeepSeek V4 Pro 0813DeepSeek · Open weight
51.0%
35
MiMo-V2.5-ProXiaomi · Closed
50.6%
36
Qwen3.8-Flash-NextAlibaba · Open weight
50.6%
37
DeepSeek V4 Flash 0731DeepSeek · Open weight
50.3%
38
MiniMax M2.7MiniMax · Open weight
50.1%
39
Inkling-SmallThinking Machines Lab · Open weight
49.7%
40
Qwen3.7 MaxAlibaba · Closed
49.5%
41
Hy3Tencent · Open weight
48.6%
42
Hy3 PreviewTencent · Open weight
48.6%
43
Grok 4.3xAI · Closed
48.3%
44
Quasar 438BMultiverse Computing · Closed
48.1%
45
Kimi K2.7 CodeMoonshot AI · Open weight
47.8%
46
GPT-5.4 nanoOpenAI · Closed
47.2%
47
MiniMax M3MiniMax · Open weight
47.1%
48
InklingThinking Machines Lab · Open weight
47.0%
49
Qwen3.8-27BAlibaba · Open weight
46.6%
50
DeepSeek V4 Pro (High)DeepSeek · Open weight
46.4%
51
Gemini 2.5 ProGoogle · Closed
46.3%
52
Qwen3.7 PlusAlibaba · Closed
46.1%
53
Gemma 4 31BGoogle · Open weight
45.5%
54
Apodex 1.1Apodex · Closed
45.5%
55
Apodex 1.1 MiniApodex · Open weight
45.5%
56
Muse Glimmer 30BMeta · Open weight
44.9%
57
GLM-5.1Z.AI · Open weight
44.8%
58
Solar Pro 4Upstage · Closed
44.6%
59
Ling 3.0 Flash VLInclusionAI · Open weight
44.2%
60
Step 3.7 FlashStepFun · Open weight
43.9%
61
Qwen3.6-27BAlibaba · Open weight
42.8%
62
Ling 3.0 FlashInclusionAI · Open weight
42.0%
63
Ling 3.0 Flash FP8InclusionAI · Open weight
42.0%
64
K-EXAONE 2.0LG AI Research · Open weight
42.0%
65
Gemini 3.5 Flash-LiteGoogle · Closed
41.3%
66
A.X K2SK Telecom · Open weight
41.0%
67
Trinity-Large-ThinkingArcee AI · Open weight
40.6%
68
Trinity-Large-PreviewArcee AI · Open weight
40.6%
69
Nemotron 3 UltraNVIDIA · Open weight
40.3%
70
Mistral Medium 3.5 128BMistral · Open weight
40.2%
71
Gemma 4 26B A4BGoogle · Open weight
40.0%
72
Qwen3.5-122B-A10BAlibaba · Open weight
39.7%
73
GPT-5 miniOpenAI · Closed
39.0%
74
GPT-OSS 20BOpenAI · Open weight
38.9%
75
Mistral Small 4Mistral · Open weight
38.8%
76
North Mini CodeCohere · Open weight
38.8%
77
Mistral Small 4 (Reasoning)Mistral · Open weight
38.8%
78
Command A+Cohere · Open weight
38.5%
79
Granite 4.2 30BIBM · Open weight
37.8%
80
Mercury 2Inception · Closed
37.7%
81
Qwen3.6-35B-A3BAlibaba · Open weight
36.6%
82
Mistral Large 3Mistral · Closed
36.6%
83
Nemotron 3 Super 100BNVIDIA · Open weight
36.2%
84
Nemotron 3 Super 120B A12BNVIDIA · Open weight
36.2%
85
DeepSeek V3DeepSeek · Open weight
35.8%
86
GPT-OSS 120BOpenAI · Open weight
34.0%
87
32.1%
88
Llama 4 MaverickMeta · Open weight
31.7%
89
Granite 4.2 8BIBM · Open weight
31.5%
90
Nemotron 3 Nano 30BNVIDIA · Open weight
30.6%
91
MiniCPM5-2BOpenBMB · Open weight
26.3%
92
Solar Pro 3Upstage · Closed
25.5%
93
Granite 4.2 3BIBM · Open weight
25.3%
94
Ling 3.0 TinyInclusionAI · Open weight
24.2%
95
Ministral 3 14B (Reasoning)Mistral · Open weight
23.8%
96
Ministral 3 14BMistral · Open weight
23.8%
97
Gemma 3 27BGoogle · Open weight
23.3%
98
Celeris-1Celeris · Closed
21.6%
99
Llama 4 ScoutMeta · Open weight
21.3%
100
Ministral 3 8B (Reasoning)Mistral · Open weight
20.7%
101
Ministral 3 8BMistral · Open weight
20.7%
102
Ministral 3 3B (Reasoning)Mistral · Open weight
15.3%
103
Ministral 3 3BMistral · Open weight
15.3%
104
LFM2.5-2.6BLiquidAI · Open weight
14.4%

According to BenchLM.ai, Claude Opus 5.5 leads the AA-SciCode benchmark with a score of 66.9%, followed by Claude Fable 5.1 (63.1%) and Claude Fable 5 (61.0%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

104 models have been evaluated on AA-SciCode. The benchmark falls in the Coding category. BenchAlign v5.7 gives AA-SciCode 5% of the Coding reference weight, so it moves the Coding leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About AA-SciCode

Year

2026

Tasks

Scientific coding subproblems

Format

Task success rate

Difficulty

Scientific programming

BenchLM stores the Artificial Analysis SciCode result separately from the core SciCode lane because Artificial Analysis controls the evaluation configuration.

Freshness and provenance

Version

AA-SciCode 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA-SciCode measure?

An Artificial Analysis SciCode score.

Which model scores highest on AA-SciCode?

Claude Opus 5.5 by Anthropic currently leads with a score of 66.9% on AA-SciCode.

How many models are evaluated on AA-SciCode?

104 AI models have been evaluated on AA-SciCode on BenchLM.

Last updated: September 27, 2026 · BenchLM version AA-SciCode 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.