Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Artificial Analysis Long Context Reasoning (AA-LCR)

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A display-only Artificial Analysis long-context reasoning evaluation.

Top models on AA-LCR — September 18, 2026

As of September 18, 2026, Kimi K3 leads the AA-LCR leaderboard with 88.7% , followed by Claude Fable 5.1 (85.3%) and GPT-5.5 (84.3%).

190 modelsReasoning15% of category scoreCurrentUpdated September 18, 2026

Leaderboard (190 models)

Score
1
Kimi K3Moonshot AI · Closed
88.7%
2
Claude Fable 5.1Anthropic · Closed
85.3%
3
GPT-5.5OpenAI · Closed
84.3%
4
GPT-5.6 SolOpenAI · Closed
84.0%
5
DeepSeek V4.1 FlashDeepSeek · Open weight
84.0%
6
GPT-5.6 LunaOpenAI · Closed
83.7%
7
Muse Glimmer 30BMeta · Open weight
83.3%
8
GPT-5.3 CodexOpenAI · Closed
83.3%
9
GPT-5.3-Codex-SparkOpenAI · Closed
83.3%
10
GPT-5.6 TerraOpenAI · Closed
83.0%
11
MiniMax M3MiniMax · Open weight
83.0%
12
Muse Spark 1.3Meta · Closed
83.0%
13
GPT-5.2OpenAI · Closed
82.7%
14
Claude Fable 5Anthropic · Closed
82.3%
15
GPT-5.2-CodexOpenAI · Closed
82.3%
16
Claude Sonnet 5Anthropic · Closed
82.0%
17
Gemini 3.1 ProGoogle · Closed
82.0%
18
GPT-5.4OpenAI · Closed
82.0%
19
Qwen3.8-27BAlibaba · Open weight
82.0%
20
Gemini 3.7 FlashGoogle · Closed
81.7%
21
Gemini 3.8 FlashGoogle · Closed
81.3%
22
Kimi K2.6Moonshot AI · Open weight
81.0%
23
GPT-6 AstraOpenAI · Closed
80.7%
24
Qwen 3.6 Max (preview)Alibaba · Closed
80.7%
25
DeepSeek V4 Pro 0813DeepSeek · Closed
80.3%
26
Grok 4.6xAI · Closed
80.3%
27
Qwen3.8 Max PreviewAlibaba · Closed
80.3%
28
Gemini 3.6 FlashGoogle · Closed
80.0%
29
GPT-5.1OpenAI · Closed
80.0%
30
GLM-5.3Z.AI · Open weight
79.7%
31
MiMo-V2.5-ProXiaomi · Closed
79.7%
32
Qwen3.8-Flash-NextAlibaba · Open weight
79.7%
33
DeepSeek V4 Flash 0731DeepSeek · Closed
79.7%
34
Claude Opus 5Anthropic · Closed
79.3%
35
Kimi K2.7 CodeMoonshot AI · Open weight
79.3%
36
Apodex 1.1Apodex · Closed
79.3%
37
Grok 4.5xAI · Closed
79.3%
38
Apodex 1.1 MiniApodex · Open weight
79.3%
39
Qwen3.7 MaxAlibaba · Closed
79.0%
40
Muse Spark 1.2Meta · Closed
79.0%
41
Hy3Tencent · Open weight
79.0%
42
Claude Opus 4.7 (Adaptive)Anthropic · Closed
78.7%
43
GLM-5.2Z.AI · Open weight
78.3%
44
Qwen3.6 PlusAlibaba · Closed
78.3%
45
MiniMax M2.7MiniMax · Open weight
78.3%
46
Ling 3.0 Flash VLInclusionAI · Open weight
78.3%
47
GPT-5 (high)OpenAI · Closed
78.2%
48
Muse SparkMeta · Closed
78.0%
49
Kimi K2.5Moonshot AI · Open weight
78.0%
50
Kimi K2.5 (Reasoning)Moonshot AI · Closed
78.0%
51
Claude Opus 4.6 (Adaptive)Anthropic · Closed
78.0%
52
Claude Opus 4.8Anthropic · Closed
77.7%
53
Muse Spark 1.1Meta · Closed
77.7%
54
Qwen3.5-27BAlibaba · Open weight
77.7%
55
InklingThinking Machines Lab · Open weight
77.3%
56
Qwen3.6-27BAlibaba · Open weight
77.3%
57
Claude Opus 4.5 ThinkingAnthropic · Closed
77.3%
58
GPT-5.4 miniOpenAI · Closed
77.0%
59
Ternary Bonsai 2 27BPrism ML · Open weight
77.0%
60
GPT-5.4 nanoOpenAI · Closed
76.7%
61
Qwen3.5-122B-A10BAlibaba · Open weight
76.3%
62
Quasar 438BMultiverse Computing · Closed
76.3%
63
Gemini 3 ProGoogle · Closed
76.0%
64
GPT-5 (medium)OpenAI · Closed
76.0%
65
Claude 4.1 Opus ThinkingAnthropic · Closed
76.0%
66
Gemini 3.5 Flash-LiteGoogle · Closed
76.0%
67
Inkling-SmallThinking Machines Lab · Open weight
75.7%
68
GLM-5Z.AI · Open weight
75.7%
69
Claude Opus 4.7Anthropic · Closed
75.7%
70
MiMo-V2-OmniXiaomi · Closed
75.0%
71
o3OpenAI · Closed
74.7%
72
74.0%
73
GLM-5.1Z.AI · Open weight
73.7%
74
73.7%
75
Step 3.7 FlashStepFun · Open weight
73.7%
76
MiniMax M2.5MiniMax · Closed
73.3%
77
Qwen3.7 PlusAlibaba · Closed
73.0%
78
Ling 3.0 FlashInclusionAI · Open weight
73.0%
79
Ling 3.0 Flash FP8InclusionAI · Open weight
73.0%
80
GPT-5 miniOpenAI · Closed
72.3%
81
Qwen3.5-35B-A3BAlibaba · Open weight
72.0%
82
Qwen3.6-35B-A3BAlibaba · Open weight
71.7%
83
GLM-5-TurboZ.AI · Closed
71.7%
84
GLM-4.7Z.AI · Open weight
71.0%
85
Solar Pro 4Upstage · Closed
71.0%
86
Claude Opus 4.5Anthropic · Closed
70.7%
87
GLM-5V-TurboZ.AI · Closed
70.3%
88
Gemma 4 31BGoogle · Open weight
69.7%
89
Gemini 3.5 FlashGoogle · Closed
69.3%
90
Mistral Medium 3.5 128BMistral · Open weight
69.3%
91
GPT-5.1-CodexOpenAI · Closed
69.3%
92
GPT-5.1-Codex-MaxOpenAI · Closed
69.3%
93
Gemini 2.5 ProGoogle · Closed
69.0%
94
Claude Sonnet 4.6Anthropic · Closed
68.3%
95
MiMo-V2-ProXiaomi · Closed
68.3%
96
GPT-4.1OpenAI · Closed
68.3%
97
Grok 4xAI · Closed
68.0%
98
Mercury 2.5Inception · Closed
68.0%
99
Claude Opus 4.6Anthropic · Closed
67.0%
100
Nemotron 3 UltraNVIDIA · Open weight
67.0%
101
DeepSeek V4 Pro (High)DeepSeek · Open weight
67.0%
102
Hy3 PreviewTencent · Open weight
66.7%
103
A.X K2SK Telecom · Open weight
66.0%
104
Gemma 4 26B A4BGoogle · Open weight
65.7%
105
Nemotron 3 Super 100BNVIDIA · Open weight
65.7%
106
Nemotron 3 Super 120B A12BNVIDIA · Open weight
65.7%
107
o1OpenAI · Closed
65.0%
108
Qwen3.5 397BAlibaba · Open weight
64.3%
109
Grok 4.3xAI · Closed
64.3%
110
Qwen3.5 397B (Reasoning)Alibaba · Open weight
64.3%
111
Gemma 4 12BGoogle · Open weight
63.7%
112
Step 3.5 FlashStepFun · Open weight
63.0%
113
Solar Open 2Upstage · Open weight
62.3%
114
K-ExaoneLG AI Research · Closed
61.3%
115
Ling 3.0 TinyInclusionAI · Open weight
60.3%
116
MiniCPM5-2BOpenBMB · Open weight
59.0%
117
MiniMax M1 80kMiniMax · Closed
57.7%
118
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
56.7%
119
K-EXAONE 2.0LG AI Research · Open weight
56.2%
120
DeepSeek-R1DeepSeek · Open weight
55.7%
121
Gemini 3 FlashGoogle · Closed
55.3%
122
Grok Code Fast 1xAI · Closed
53.0%
123
Kimi K2Moonshot AI · Closed
53.0%
124
Command A+Cohere · Open weight
52.7%
125
GPT-OSS 120BOpenAI · Open weight
52.0%
126
Qwen3 MaxAlibaba · Closed
50.0%
127
Llama 4 MaverickMeta · Open weight
50.0%
128
Gemini 2.5 FlashGoogle · Closed
49.9%
129
Mistral Small 4Mistral · Open weight
49.7%
130
Mistral Small 4 (Reasoning)Mistral · Open weight
49.7%
131
GPT-4oOpenAI · Closed
49.3%
132
49.2%
133
Granite 4.2 30BIBM · Open weight
49.0%
134
DeepSeek V3.1DeepSeek · Open weight
47.0%
135
GLM-4.5-AirZ.AI · Closed
46.7%
136
DeepSeek V3.2DeepSeek · Open weight
45.7%
137
Granite 4.2 8BIBM · Open weight
45.0%
138
GPT-5 nanoOpenAI · Closed
45.0%
139
Claude 4 SonnetAnthropic · Closed
44.0%
140
GPT-4.1 miniOpenAI · Closed
44.0%
141
Mercury 2Inception · Closed
43.7%
142
GLM-4.7-FlashZ.AI · Open weight
41.7%
143
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
39.7%
144
Celeris-1Celeris · Closed
38.0%
145
Trinity-Large-ThinkingArcee AI · Open weight
38.0%
146
Trinity-Large-PreviewArcee AI · Open weight
38.0%
147
Nemotron 3 Nano 30BNVIDIA · Open weight
38.0%
148
North Mini CodeCohere · Open weight
37.3%
149
MiMo-V2-FlashXiaomi · Open weight
36.3%
150
Mistral Large 3Mistral · Closed
36.0%
151
GPT-OSS 20BOpenAI · Open weight
34.7%
152
Solar Pro 3Upstage · Closed
32.3%
153
Gemma 4 E4BGoogle · Open weight
32.0%
154
Mistral Medium 3Mistral · Closed
31.3%
155
Grok 4.1 FastxAI · Closed
31.3%
156
Ling 2.6 FlashInclusionAI · Open weight
31.3%
157
DeepSeek V3DeepSeek · Open weight
29.3%
158
Llama 4 ScoutMeta · Open weight
27.7%
159
Claude 3 HaikuAnthropic · Closed
27.7%
160
GLM-4.6Z.AI · Open weight
26.3%
161
Ministral 3 14B (Reasoning)Mistral · Open weight
26.3%
162
Ministral 3 14BMistral · Open weight
26.3%
163
Ministral 3 8B (Reasoning)Mistral · Open weight
25.7%
164
Ministral 3 8BMistral · Open weight
25.7%
165
Llama 3.1 405BMeta · Open weight
25.3%
166
Granite 4.2 3BIBM · Open weight
24.3%
167
Nova ProAmazon · Closed
21.0%
168
GPT-4.1 nanoOpenAI · Closed
20.3%
169
Ministral 3 3B (Reasoning)Mistral · Open weight
17.0%
170
Ministral 3 3BMistral · Open weight
17.0%
171
Gemma 4 E2BGoogle · Open weight
16.3%
172
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
15.7%
173
DeepSeek R1 Distill Qwen 32BDeepSeek · Open weight
8.7%
174
Gemma 3 27BGoogle · Open weight
7.3%
175
Granite-4.0-H-1BIBM · Open weight
7.0%
176
LFM2.5-2.6BLiquidAI · Open weight
5.7%
177
Sarvam 105BSarvam · Open weight
0.0%
178
Solar Pro 2Upstage · Closed
0.0%
179
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weight
0.0%
180
Phi-4Microsoft · Open weight
0.0%
181
Sarvam 30BSarvam · Open weight
0.0%
182
Granite-4.0-350MIBM · Open weight
0.0%
183
Granite-4.0-H-350MIBM · Open weight
0.0%
184
Exaone 4.0 1.2BLG AI Research · Open weight
0.0%
185
LFM2.5-8B-A1BLiquidAI · Open weight
0.0%
186
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
0.0%
187
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weight
0.0%
188
LFM2-24B-A2BLiquidAI · Closed
0.0%
189
LFM2.5-1.2B-ThinkingLiquidAI · Closed
0.0%
190
LFM2.5-1.2B-InstructLiquidAI · Closed
0.0%

According to BenchLM.ai, Kimi K3 leads the AA-LCR benchmark with a score of 88.7%, followed by Claude Fable 5.1 (85.3%) and GPT-5.5 (84.3%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

190 models have been evaluated on AA-LCR. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Within that category, AA-LCR contributes 15% of the category score, so strong performance here directly affects a model's overall ranking.

About AA-LCR

Year

2026

Tasks

Long-context reasoning tasks

Format

Accuracy

Difficulty

Long-context reasoning

BenchLM stores AA-LCR as a display-only row when OpenRouter or Artificial Analysis publishes the exact long-context reasoning card value.

Freshness and provenance

Version

AA-LCR 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA-LCR measure?

A display-only Artificial Analysis long-context reasoning evaluation.

Which model scores highest on AA-LCR?

Kimi K3 by Moonshot AI currently leads with a score of 88.7% on AA-LCR.

How many models are evaluated on AA-LCR?

190 AI models have been evaluated on AA-LCR on BenchLM.

Last updated: September 18, 2026 · BenchLM version AA-LCR 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.