LiveCodeBench v6
We show this table for reference; we do not rank on it.
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
Benchmark score on LiveCodeBench v6 — October 10, 2026
We compile the LiveCodeBench v6 rows from provider self-reports and secondary reports. Sakana Fugu-Ultra leads the table at 93.2%, followed by Sakana Fugu (92.9%) and Qwen3.8-Omni-Flash (92.6%). We do not use these results to rank models overall.
Sakana Fugu-Ultra
Sakana AI
Sakana Fugu
Sakana AI
Qwen3.8-Omni-Flash
Alibaba
53 modelsCodingCurrentDisplay onlyUpdated October 10, 2026
Benchmark score table (53 models)
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Sakana Fugu-UltraSakana AI | 93.2% | Not reported | Closed |
| 2 | Sakana FuguSakana AI | 92.9% | Not reported | Closed |
| 3 | Qwen3.8-Omni-FlashAlibaba | 92.6% | Not reported | Closed |
| 4 | Solar Open 2Upstage | 92.4% | Not reported | Open |
| 5 | Qwen3.8-Flash-NextAlibaba | 91.9% | Not reported | Open |
| 6 | Qwen3.7 MaxAlibaba | 91.6% | Not reported | Closed |
| 7 | dots3-note PreviewDots Studio | 91.5% | Not reported | Open |
| 8 | Qwen3.8-27BAlibaba | 90.3% | Not reported | Open |
| 9 | Ternary Bonsai 2 27BPrism ML | 90.1% | Not reported | Open |
| 10 | Qwen3.7 PlusAlibaba | 89.6% | Not reported | Closed |
| 11 | Kimi K2.6Moonshot AI | 89.6% | Not reported | Open |
| 12 | Nemotron 3 UltraNVIDIA | 89.0% | Not reported | Open |
| 13 | BTL-3Bad Theory Labs | 88.1% | Not reported | Open |
| 14 | MAI-Thinking-1Microsoft | 87.7% | Not reported | Closed |
| 15 | Qwen3.6 PlusAlibaba | 87.1% | Not reported | Closed |
| 16 | Step 3.5 FlashStepFun | 86.4% | Not reported | Open |
| 17 | Kimi K2.5 (Reasoning)Moonshot AI | 85.0% | Not reported | Closed |
| 18 | Kimi K2.5Moonshot AI | 85.0% | Not reported | Open |
| 19 | GLM-4.7Z.AI | 84.9% | Not reported | Open |
| 20 | Claude Opus 4.5Anthropic | 84.8% | Not reported | Closed |
| 21 | A.X K2SK Telecom | 84.0% | Not reported | Open |
| 22 | Qwen3.6-27BAlibaba | 83.9% | Not reported | Open |
| 23 | Qwen3.5 397BAlibaba | 83.6% | Not reported | Open |
| 24 | Mellum2.1-12B-A2.5B-ThinkingJetBrains | 82.0% | Not reported | Open |
| 25 | Qwen3.5-27BAlibaba | 80.7% | Not reported | Open |
| 26 | MiMo-V2-FlashXiaomi | 80.6% | Not reported | Open |
| 27 | Qwen3.6-35B-A3BAlibaba | 80.4% | Not reported | Open |
| 28 | Gemma 4 31BGoogle | 80.0% | Not reported | Open |
| 29 | Qwen3.5-122B-A10BAlibaba | 78.9% | Not reported | Open |
| 30 | Gemma 4 26B A4BGoogle | 77.1% | Not reported | Open |
| 31 | Granite 4.2 30BIBM | 75.8% | Not reported | Open |
| 32 | Qwen3.5-35B-A3BAlibaba | 74.6% | Not reported | Open |
| 33 | Qwen3 235B 2507 (Reasoning)Alibaba | 74.1% | Not reported | Open |
| 34 | Granite 4.2 8BIBM | 73.2% | Not reported | Open |
| 35 | Gemma 4 12BGoogle | 72.0% | Not reported | Open |
| 36 | Mellum2-12B-A2.5B-ThinkingJetBrains | 69.9% | Not reported | Open |
| 37 | Granite 4.2 3BIBM | 69.7% | Not reported | Open |
| 38 | MiniCPM5-2BOpenBMB | 69.1% | Not reported | Open |
| 39 | Nemotron 3 Nano 30BNVIDIA | 68.3% | Not reported | Open |
| 40 | BTL-4Bad Theory Labs | 66.1% | Not reported | Open |
| 41 | ZAYA1-8BZyphra | 65.8% | Not reported | Open |
| 42 | ZAYA1-74B-PreviewZyphra | 65.7% | Not reported | Open |
| 43 | GLM-4.7-FlashZ.AI | 64.0% | Not reported | Open |
| 44 | Agents-A1-4BInternScience | 59.6% | Not reported | Open |
| 45 | LFM2.5-2.6BLiquidAI | 59.4% | Not reported | Open |
| 46 | Kimi K2Moonshot AI | 53.7% | Not reported | Closed |
| 47 | Gemma 4 E4BGoogle | 52.0% | Not reported | Open |
| 48 | Qwen3 235B 2507Alibaba | 51.8% | Not reported | Open |
| 49 | Gemma 4 E2BGoogle | 44.0% | Not reported | Open |
| 50 | Mellum2-12B-A2.5B-InstructJetBrains | 37.2% | Not reported | Open |
| 51 | MiniCPM5-1BOpenBMB | 33.5% | Not reported | Open |
| 52 | Mistral Medium 3Mistral | 30.3% | Not reported | Closed |
| 53 | LLaDA2.2-miniInclusionAI | 28.1% | Not reported | Open |
How to read this leaderboard
Editorial review by Glevd · 2026-07-15
Compare v6 rows only when the date window, code-generation scenario, pass@k or average-at-k metric, sampling count, temperature, and execution policy match. A shared v6 label does not guarantee the rest of the setup is controlled.
Operator receipt: 53 sourced rows are currently displayable on this page; the leading published row is Sakana Fugu-Ultra at 93.2%.
Honest limit: Provider tables can use different v6 windows and generation settings. The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch review, or regression safety.
Among the reported LiveCodeBench v6 rows, Sakana Fugu-Ultra is first at 93.2%. The third row is 0.6 points behind. The broader top-10 range is 3.6 points, so many of the published results sit in a relatively narrow band.
53 models have been evaluated on LiveCodeBench v6. The benchmark falls in the Coding category. LiveCodeBench v6 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About LiveCodeBench v6
Year
2026
Tasks
Fresh programming problems
Format
Provider-published v6 competitive programming results
Difficulty
Competitive programming level
Providers often publish a specific release or date window instead of the rolling aggregate. This route contains rows explicitly labeled v6 by their sources and excludes them from the weighted legacy lane.
Freshness and provenance
Version
LiveCodeBench v6 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does LiveCodeBench v6 measure?
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
Which model scores highest on LiveCodeBench v6?
Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.2% on LiveCodeBench v6.
How many models are evaluated on LiveCodeBench v6?
53 AI models have published results on LiveCodeBench v6 in the BenchLM catalog.
Compare top models on LiveCodeBench v6
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.