Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

LiveCodeBench, Vals AI run (LiveCodeBench (Vals))

Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Vals LiveCodeBench mirror

BenchLM mirrors the public Vals AI Vals LiveCodeBench mirror leaderboard captured from https://www.vals.ai/benchmarks/lcb and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals LiveCodeBench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

143 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals LiveCodeBench mirror score on LiveCodeBench (Vals) — September 1, 2026

We mirror the published vals livecodebench mirror score view for LiveCodeBench (Vals). Claude Fable 5.1 leads the public snapshot at 90.5%, followed by Claude Fable 5 (89.8%) and Gemini 3.8 Flash (89.5%). We do not use these results to rank models overall.

143 modelsCoding15% of category scoreCurrentUpdated September 1, 2026

Vals LiveCodeBench mirror score table (143 models)

Score
1
Claude Fable 5.1Anthropic · Closed
90.5%
2
Claude Fable 5Anthropic · Closed
89.8%
3
Gemini 3.8 FlashGoogle · Closedhigh reasoning
89.5%
4
Claude Opus 5Anthropic · Closed
89.0%
5
Gemini 3.7 FlashGoogle · Closedhigh reasoning
88.7%
6
Gemini 3.1 Pro PreviewGooglehigh reasoning
88.5%
7
Grok 4.6xAI · Closedhigh reasoning
88.2%
8
Gemini 3.6 FlashGoogle · Closedhigh reasoning
88.1%
9
GPT-5.2-CodexOpenAI · Closed
88.0%
10
Qwen3.8 MaxAlibaba · Open weight
87.9%
11
Claude Opus 4.8Anthropic · Closed
87.8%
12
Gemini 3.5 FlashGoogle · Closedhigh reasoning
87.6%
13
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
87.5%
14
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
87.5%
15
Grok 4.5xAI · Closedhigh reasoning
87.4%
16
GPT-5.3 CodexOpenAI · Closedxhigh reasoning
87.3%
17
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
87.3%
18
Kimi K3Moonshot AI · Closedmax reasoning
87.2%
19
Qwen3.7 MaxAlibaba · Closed
87.1%
20
Kimi K2.6Moonshot AI · Open weight
86.8%
21
GPT-5 miniOpenAI · Closedhigh reasoning
86.6%
22
GPT-5.1OpenAI · Closedhigh reasoning
86.5%
25
Qwen3.6 PlusAlibaba · Closed
86.0%
26
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
85.9%
27
GPT-5.6 TerraOpenAI · Closedxhigh reasoning
85.9%
28
GPT-5OpenAIhigh reasoning
85.9%
29
Muse Spark 1.1Meta · Closedxhigh reasoning
85.9%
30
Gemini 3 Flash PreviewGooglehigh reasoning
85.6%
31
GPT-5.1-CodexOpenAI · Closed
85.5%
32
InklingThinking Machines Lab · Open weight0.99 reasoning
85.5%
33
GPT-5.2OpenAI · Closedxhigh reasoning
85.4%
34
85.3%
35
GPT-5.5OpenAI · Closedxhigh reasoning
85.3%
36
Claude Opus 4.7Anthropic · Closed
85.1%
37
GPT-5 CodexOpenAIhigh reasoning
84.7%
38
Claude Opus 4.6 (Adaptive)Anthropic · Closed
84.7%
39
Grok 4.3xAI · Closedhigh reasoning
84.5%
41
GPT-5.4OpenAI · Closedxhigh reasoning
84.1%
42
GPT-5.4 nanoOpenAI · Closedhigh reasoning
84.0%
43
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
84.0%
45
o3OpenAI · Closedhigh reasoning
83.9%
46
83.9%
47
Claude Opus 4.5 ThinkingAnthropic · Closed
83.7%
48
GPT-5.1-Codex-MaxOpenAI · Closedhigh reasoning
83.6%
49
Qwen3.5 FlashAlibaba · Closed
83.3%
50
83.2%
51
GPT-OSS 120BOpenAI · Open weight
83.2%
52
GPT-5.6 SolOpenAI · Closedmax reasoning
82.6%
53
Claude Sonnet 5Anthropic · Closed
82.4%
54
GLM-4.7Z.AI · Open weight
82.2%
55
O4 MiniOpenAIhigh reasoning
82.2%
56
MiniMax M3MiniMax · Open weight
82.2%
57
Claude Sonnet 4.6Anthropic · Closed
82.1%
58
Kimi K2.7 CodeMoonshot AI · Open weight
82.0%
59
81.9%
60
81.8%
61
MiMo-V2.5Xiaomi · Closed
81.5%
62
GPT-5.4 miniOpenAI · Closedxhigh reasoning
81.5%
63
GLM-5.1Z.AI · Open weight
81.4%
64
MiMo-V2.5-ProXiaomi · Closed
81.4%
65
GLM-4.6Z.AI · Open weight
81.0%
66
80.7%
67
80.6%
68
GLM-5.3Z.AI · Open weightmax reasoning
80.5%
69
GLM-5.3-FlashZ.AI · Open weightmax reasoning
80.5%
70
GPT-OSS 20BOpenAI · Open weight
80.4%
72
MiniMax M2.7MiniMax · Open weight
79.9%
73
MiniMax M2.5MiniMax · Closed
79.2%
75
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
79.0%
76
79.0%
77
Qwen3 MaxAlibaba · Closed
78.2%
78
76.2%
81
Claude Opus 4.5Anthropic · Closed
75.0%
82
74.9%
83
Claude Sonnet 4.5 ThinkingAnthropic · Closed
73.0%
84
72.1%
85
o3-miniOpenAI · Closedhigh reasoning
71.5%
87
Qwen3 235b A22bFireworks AI
70.6%
88
70.4%
89
DeepSeek-R1DeepSeek · Open weight
70.2%
90
GPT-5 nanoOpenAI · Closedhigh reasoning
70.2%
92
DeepSeek V3p2Fireworks AI
69.9%
93
GLM-5.2Z.AI · Open weight
69.5%
94
Laguna M.1Poolside · Closed
68.1%
95
Laguna XS.2Poolside · Open weight
67.8%
97
GLM-4.5Z.AI · Closed
67.4%
98
66.9%
100
66.3%
101
DeepSeek V3 0324Fireworks AI
65.5%
102
64.6%
103
Kimi K2 ThinkingMoonshot AI
63.1%
104
Claude Opus 4Anthropic
62.6%
106
Grok Code Fast 1xAI · Closed
62.0%
108
59.7%
110
GPT-4.1 miniOpenAI · Closedhigh reasoning
58.2%
112
56.7%
113
55.3%
114
GPT-4.1OpenAI · Closedhigh reasoning
54.7%
115
52.9%
116
Devstral 2512Mistral AI
51.8%
117
o1OpenAI · Closedhigh reasoning
50.3%
118
Claude 3.5 SonnetAnthropic · Closed
49.6%
119
47.3%
122
44.8%
123
43.6%
124
GPT-4oOpenAI · Closedhigh reasoning
43.4%
125
43.2%
126
GPT-4.1 nanoOpenAI · Closedhigh reasoning
42.7%
128
41.9%
129
41.7%
130
Claude Haiku 4.5 ThinkingAnthropic · Closed
41.2%
131
38.7%
133
37.1%
134
36.9%
137
35.1%
138
31.8%
139
GPT-4o miniOpenAI · Closedhigh reasoning
26.4%
140
22.3%
141
18.2%
142
15.8%
143
9.9%

The published LiveCodeBench (Vals) snapshot places Claude Fable 5.1 first at 90.5%. The third row is 1.0 points behind. The broader top-10 range is 2.7 points, so many of the published results sit in a relatively narrow band.

143 models have been evaluated on LiveCodeBench (Vals). The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, LiveCodeBench (Vals) contributes 15% of the category score, so strong performance here directly affects a model's overall ranking.

About LiveCodeBench (Vals)

Year

2026

Tasks

Competitive programming problems (easy, medium, hard)

Format

Pass@1 accuracy

Difficulty

Frontier coding

BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).

BenchLM freshness & provenance

Version

LiveCodeBench (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LiveCodeBench (Vals) measure?

Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.

Which model leads the published LiveCodeBench (Vals) snapshot?

Claude Fable 5.1 currently leads the published LiveCodeBench (Vals) snapshot with 90.5% vals livecodebench mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LiveCodeBench (Vals)?

The September 1, 2026 snapshot contains 143 AI models.

Last updated: September 1, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.