Skip to main content
BenchLM

LisanBench

We show this table for reference; we do not rank on it.

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

Difficulty-weighted score on LisanBench — September 23, 2026 snapshot

We mirror the published difficulty-weighted score view for LisanBench. Claude Opus 4.7 (Adaptive) leads the public snapshot at 5122.60, followed by Claude Opus 5 (5068.25) and Claude Fable 5 (4561.82). We do not use these results to rank models overall.

154 modelsReasoningCurrentDisplay onlyUpdated September 23, 2026 snapshot

Difficulty-weighted score table (154 models)

Score
1
Claude Opus 4.7 (Adaptive)Anthropic · Closed
5122.60
2
Claude Opus 5Anthropic · Closed
5068.25
3
Claude Fable 5Anthropic · Closed
4561.82
4
Claude Opus 4.6 (Adaptive)Anthropic · Closed
3526.49
5
Kimi K3Moonshot AI · Closed
3525.77
6
GPT-5.5OpenAI · Closed
3315.52
7
Claude Sonnet 4.6Anthropic · Closed
2944.27
8
GPT-5.4OpenAI · Closed
2738.16
9
Claude Opus 4.8Anthropic · Closed
2693.67
10
Claude Opus 5Anthropic · Closed
2270.40
11
Claude Opus 4.5 ThinkingAnthropic · Closed
2204.43
12
Gemini 3.1 ProGoogle · Closed
1929.11
13
Grok 4xAI · Closed
1778.40
14
Claude Sonnet 5Anthropic · Closed
1736.56
15
o3OpenAI · Closed
1523.00
17
Grok 4.20xAI · Closed
1464.63
18
GPT-5.2OpenAI · Closed
1458.80
19
GPT-5 (medium)OpenAI · Closed
1457.28
20
GPT-5.6 SolOpenAI · Closed
1403.28
21
GPT-5.6 TerraOpenAI · Closed
1196.90
22
Gemini 3 ProGoogle · Closed
1130.63
23
Gemini 3.5 FlashGoogle · Closed
1128.26
24
Claude Sonnet 4.5 ThinkingAnthropic · Closed
1090.82
25
DeepSeek V4 Flash 0731DeepSeek · Open weight
1063.47
26
DeepSeek V4 Pro 0813DeepSeek · Open weight
1059.51
27
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
925.31
28
Gemini 3.1 ProGoogle · Closed
872.67
29
Step 3.5 FlashStepFun · Open weight
811.21
30
806.48
31
GPT-5 miniOpenAI · Closed
758.84
32
GPT-5.6 LunaOpenAI · Closed
648.18
33
Kimi K2.5 (Reasoning)Moonshot AI · Closed
641.96
34
Kimi K2Moonshot AI · Closed
633.05
35
GPT-5 nanoOpenAI · Closed
626.86
36
604.43
37
602.95
38
GLM-5.2Z.AI · Open weight
591.99
39
Gemini 3 FlashGoogle · Closed
591.88
40
GPT-5.4 miniOpenAI · Closed
591.48
41
GPT-5.4 nanoOpenAI · Closed
543.21
42
o3-miniOpenAI · Closed
518.37
44
GPT-OSS 120BOpenAI · Open weight
448.33
45
Qwen3.5 397B (Reasoning)Alibaba · Open weight
387.74
46
Claude Opus 5Anthropic · Closed
366.79
47
o4-mini (high)OpenAI · Closed
352.68
48
GLM-5 (Reasoning)Z.AI · Open weight
336.34
49
GPT-5.6 SolOpenAI · Closed
332.46
50
GPT-5.5OpenAI · Closed
305.30
51
Claude Opus 4.8Anthropic · Closed
270.10
53
Opus 4Anthropic
262.56
55
MiniMax M2.5MiniMax · Closed
228.38
56
Qwen3 235B 2507 (Reasoning)Alibaba · Open weight
226.22
57
Claude Opus 4.7Anthropic · Closed
217.52
58
Claude 4.1 OpusAnthropic · Closed
215.66
59
Claude Sonnet 4.6Anthropic · Closed
208.64
60
Claude Sonnet 5Anthropic · Closed
208.26
61
Gemini 2.5 ProGoogle · Closed
197.43
62
Grok 3 MinixAI · Closed
194.03
63
Grok 3 [Beta]xAI · Closed
188.12
64
Sonnet 3.7Anthropic
162.44
65
GPT-OSS 20BOpenAI · Open weight
156.99
66
GPT-5.6 TerraOpenAI · Closed
156.14
68
Claude 4 SonnetAnthropic · Closed
150.17
69
Claude 3.5 SonnetAnthropic · Closed
149.65
70
Claude 3.5 SonnetAnthropic · Closed
129.73
71
Gemini 1.5 ProGoogle · Closed
119.72
72
DeepSeek V3.2DeepSeek · Open weight
119.04
73
DeepSeek V4 Pro 0813DeepSeek · Open weight
117.14
74
Gemini 2.5 FlashGoogle · Closed
112.75
75
DeepSeek-R1DeepSeek · Open weight
111.60
76
Qwen3.5-122B-A10BAlibaba · Open weight
109.62
77
GPT-5.4OpenAI · Closed
109.51
78
GLM-4.5Z.AI · Closed
108.32
79
Qwen3.5-35B-A3BAlibaba · Open weight
107.61
80
104.86
81
Claude Sonnet 4.5Anthropic · Closed
103.58
82
DeepSeek V3DeepSeek · Open weight
103.39
83
103.10
84
GPT-5.6 LunaOpenAI · Closed
96.63
85
GPT-4oOpenAI · Closed
94.16
86
Claude Opus 4.5Anthropic · Closed
93.49
87
Claude Opus 4.6Anthropic · Closed
91.61
88
GPT-4 TurboOpenAI · Closed
91.48
89
Kimi K2Moonshot AI · Closed
85.92
90
77.21
91
Claude 3 OpusAnthropic · Closed
75.77
92
Gemini 2.5 FlashGoogle · Closed
72.17
93
MiniMax M1 80kMiniMax · Closed
66.71
94
GLM-5.2Z.AI · Open weight
64.50
95
62.04
96
DeepSeek V4 Flash 0731DeepSeek · Open weight
56.78
97
Horizon BetaOpenRouter
55.69
98
55.20
99
GLM-4.5-AirZ.AI · Closed
54.86
100
Nova ProAmazon · Closed
54.38
101
GLM-4.7Z.AI · Open weight
54.27
102
Polaris AlphaOpenRouter
53.34
103
Claude Haiku 4.5Anthropic · Closed
52.64
104
50.96
105
50.95
106
Llama 3.1 405BMeta · Open weight
48.88
107
Grok 4.1 FastxAI · Closed
47.29
108
GLM-4.6Z.AI · Open weight
44.20
109
Gemma 3 27BGoogle · Open weight
43.79
110
Mistral Medium 3Mistral · Closed
43.31
111
GPT-5.4 miniOpenAI · Closed
42.74
112
Llama 4 MaverickMeta · Open weight
42.33
113
GPT-4.1OpenAI · Closed
42.02
114
40.99
115
40.07
116
Devstral MediumMistral AI
40.03
118
Haiku 3.5Anthropic
38.22
119
38.21
120
38.06
121
38.06
122
Claude 3 HaikuAnthropic · Closed
35.46
123
34.64
124
GPT-4.1 miniOpenAI · Closed
32.82
126
31.92
127
Llama 4 ScoutMeta · Open weight
30.85
128
MiMo-V2-FlashXiaomi · Open weight
28.75
129
25.24
130
24.88
131
24.38
132
Qwen3 32BAlibaba
23.86
133
21.56
134
GPT-4o miniOpenAI · Closed
21.21
135
18.98
136
17.77
137
Qwen3 14BAlibaba
16.68
138
Qwen3 8BAlibaba
15.50
139
GPT-4.1 nanoOpenAI · Closed
14.95
140
Devstral SmallMistral AI
13.22
141
Codestral 2508Mistral AI
13.15
142
12.29
143
11.78
144
Mistral NemoMistral AI
11.47
145
GPT-5.4 nanoOpenAI · Closed
11.16
146
9.41
147
9.24
148
Qwen3 4BAlibaba
7.90
149
6.64
150
Qwen3 1.7BAlibaba
6.27
151
3.92
152
2.85
153
0.63
154
Qwen3 0.6BAlibaba
0.06

How LisanBench is shown here

BenchLM mirrors the current public LisanBench difficulty-weighted leaderboard using the official dataset published at lisanbench.com for the September 23, 2026 snapshot. The public benchmark tests 154 model variants across 50 starting words, with 3 trials per starting word.

LisanBench is a strong reasoning reference, but BenchLM currently keeps it display only rather than weighted. The public leaderboard is highly variant-specific, strongly English-vocabulary-dependent, and not yet aligned cleanly enough with BenchLM canonical model rows to use as a ranking input.

Snapshot

154 model variants50 starting words3 trials per wordDifficulty-weighted scoresDisplay only

The published LisanBench snapshot places Claude Opus 4.7 (Adaptive) first at 5122.60. The third row is 560.78 score units behind. The broader top-10 range is 2852.20 score units, so the table still separates the published systems.

154 models have been evaluated on LisanBench. The benchmark falls in the Reasoning category. LisanBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LisanBench

Year

2026

Tasks

50 starting words × 3 trials

Format

Difficulty-weighted word-chain reasoning

Difficulty

Open-ended lexical planning

BenchLM mirrors the public difficulty-weighted LisanBench leaderboard as a display-only reasoning benchmark. The public benchmark currently evaluates 128 model variants across 50 starting words with 3 trials per word.

Freshness and provenance

Version

LisanBench 2026

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does LisanBench measure?

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

Which model leads the published LisanBench snapshot?

Claude Opus 4.7 (Adaptive) currently leads the published LisanBench snapshot with 5122.60 difficulty-weighted score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LisanBench?

The September 23, 2026 snapshot snapshot contains 154 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.