Skip to main content
BenchLM
Data

LiveCodeBench v6

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Benchmark score on LiveCodeBench v6 — October 10, 2026

We compile the LiveCodeBench v6 rows from provider self-reports and secondary reports. Sakana Fugu-Ultra leads the table at 93.2%, followed by Sakana Fugu (92.9%) and Qwen3.8-Omni-Flash (92.6%). We do not use these results to rank models overall.

53 modelsCodingCurrentDisplay onlyUpdated October 10, 2026

Benchmark score table (53 models)

Benchmark score results for LiveCodeBench v6
RankModel / configurationScoreParameters (B)Open / closed
1Sakana Fugu-UltraSakana AI
93.2%
Not reportedClosed
2Sakana FuguSakana AI
92.9%
Not reportedClosed
3Qwen3.8-Omni-FlashAlibaba
92.6%
Not reportedClosed
4Solar Open 2Upstage
92.4%
Not reportedOpen
5Qwen3.8-Flash-NextAlibaba
91.9%
Not reportedOpen
6Qwen3.7 MaxAlibaba
91.6%
Not reportedClosed
7dots3-note PreviewDots Studio
91.5%
Not reportedOpen
8Qwen3.8-27BAlibaba
90.3%
Not reportedOpen
9Ternary Bonsai 2 27BPrism ML
90.1%
Not reportedOpen
10Qwen3.7 PlusAlibaba
89.6%
Not reportedClosed
11Kimi K2.6Moonshot AI
89.6%
Not reportedOpen
12Nemotron 3 UltraNVIDIA
89.0%
Not reportedOpen
13BTL-3Bad Theory Labs
88.1%
Not reportedOpen
14MAI-Thinking-1Microsoft
87.7%
Not reportedClosed
15Qwen3.6 PlusAlibaba
87.1%
Not reportedClosed
16Step 3.5 FlashStepFun
86.4%
Not reportedOpen
17Kimi K2.5 (Reasoning)Moonshot AI
85.0%
Not reportedClosed
18Kimi K2.5Moonshot AI
85.0%
Not reportedOpen
19GLM-4.7Z.AI
84.9%
Not reportedOpen
20Claude Opus 4.5Anthropic
84.8%
Not reportedClosed
21A.X K2SK Telecom
84.0%
Not reportedOpen
22Qwen3.6-27BAlibaba
83.9%
Not reportedOpen
23Qwen3.5 397BAlibaba
83.6%
Not reportedOpen
24Mellum2.1-12B-A2.5B-ThinkingJetBrains
82.0%
Not reportedOpen
25Qwen3.5-27BAlibaba
80.7%
Not reportedOpen
26MiMo-V2-FlashXiaomi
80.6%
Not reportedOpen
27Qwen3.6-35B-A3BAlibaba
80.4%
Not reportedOpen
28Gemma 4 31BGoogle
80.0%
Not reportedOpen
29Qwen3.5-122B-A10BAlibaba
78.9%
Not reportedOpen
30Gemma 4 26B A4BGoogle
77.1%
Not reportedOpen
31Granite 4.2 30BIBM
75.8%
Not reportedOpen
32Qwen3.5-35B-A3BAlibaba
74.6%
Not reportedOpen
33Qwen3 235B 2507 (Reasoning)Alibaba
74.1%
Not reportedOpen
34Granite 4.2 8BIBM
73.2%
Not reportedOpen
35Gemma 4 12BGoogle
72.0%
Not reportedOpen
36Mellum2-12B-A2.5B-ThinkingJetBrains
69.9%
Not reportedOpen
37Granite 4.2 3BIBM
69.7%
Not reportedOpen
38MiniCPM5-2BOpenBMB
69.1%
Not reportedOpen
39Nemotron 3 Nano 30BNVIDIA
68.3%
Not reportedOpen
40BTL-4Bad Theory Labs
66.1%
Not reportedOpen
41ZAYA1-8BZyphra
65.8%
Not reportedOpen
42ZAYA1-74B-PreviewZyphra
65.7%
Not reportedOpen
43GLM-4.7-FlashZ.AI
64.0%
Not reportedOpen
44Agents-A1-4BInternScience
59.6%
Not reportedOpen
45LFM2.5-2.6BLiquidAI
59.4%
Not reportedOpen
46Kimi K2Moonshot AI
53.7%
Not reportedClosed
47Gemma 4 E4BGoogle
52.0%
Not reportedOpen
48Qwen3 235B 2507Alibaba
51.8%
Not reportedOpen
49Gemma 4 E2BGoogle
44.0%
Not reportedOpen
50Mellum2-12B-A2.5B-InstructJetBrains
37.2%
Not reportedOpen
51MiniCPM5-1BOpenBMB
33.5%
Not reportedOpen
52Mistral Medium 3Mistral
30.3%
Not reportedClosed
53LLaDA2.2-miniInclusionAI
28.1%
Not reportedOpen

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Compare v6 rows only when the date window, code-generation scenario, pass@k or average-at-k metric, sampling count, temperature, and execution policy match. A shared v6 label does not guarantee the rest of the setup is controlled.

Operator receipt: 53 sourced rows are currently displayable on this page; the leading published row is Sakana Fugu-Ultra at 93.2%.

Honest limit: Provider tables can use different v6 windows and generation settings. The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch review, or regression safety.

Among the reported LiveCodeBench v6 rows, Sakana Fugu-Ultra is first at 93.2%. The third row is 0.6 points behind. The broader top-10 range is 3.6 points, so many of the published results sit in a relatively narrow band.

53 models have been evaluated on LiveCodeBench v6. The benchmark falls in the Coding category. LiveCodeBench v6 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LiveCodeBench v6

Year

2026

Tasks

Fresh programming problems

Format

Provider-published v6 competitive programming results

Difficulty

Competitive programming level

Providers often publish a specific release or date window instead of the rolling aggregate. This route contains rows explicitly labeled v6 by their sources and excludes them from the weighted legacy lane.

Freshness and provenance

Version

LiveCodeBench v6 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does LiveCodeBench v6 measure?

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Which model scores highest on LiveCodeBench v6?

Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.2% on LiveCodeBench v6.

How many models are evaluated on LiveCodeBench v6?

53 AI models have published results on LiveCodeBench v6 in the BenchLM catalog.

Last updated: October 10, 2026 · BenchLM version LiveCodeBench v6 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.