Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

LiveCodeBench v6

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Data verified 24 confirmed releases in the last 30 daysStart free brief

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Compare v6 rows only when the date window, code-generation scenario, pass@k or average-at-k metric, sampling count, temperature, and execution policy match. A shared v6 label does not guarantee the rest of the setup is controlled.

Operator receipt: 19 sourced rows are currently displayable on this page; the leading published row is Sakana Fugu-Ultra at 93.2%.

Honest limit: Provider tables can use different v6 windows and generation settings. The route is a sourced release ledger, not a BenchLM rerun. LiveCodeBench still measures contest-style code generation rather than repository navigation, patch review, or regression safety.

Benchmark score on LiveCodeBench v6 — August 22, 2026

We mirror the published score view for LiveCodeBench v6. Sakana Fugu-Ultra leads the public snapshot at 93.2%, followed by Sakana Fugu (92.9%) and dots3-note Preview (91.5%). We do not use these results to rank models overall.

19 modelsCodingCurrentDisplay onlyUpdated August 22, 2026

Benchmark score table (19 models)

Score
1
Sakana Fugu-UltraSakana AI · Closed
93.2%
2
Sakana FuguSakana AI · Closed
92.9%
3
dots3-note PreviewDots Studio · Open weight
91.5%
4
Qwen3.8-27BAlibaba · Open weight
90.3%
5
Kimi K2.6Moonshot AI · Open weight
89.6%
6
Nemotron 3 UltraNVIDIA · Open weight
89.0%
7
MAI-Thinking-1Microsoft · Closed
87.7%
8
Qwen3.6 PlusAlibaba · Closed
87.1%
9
Kimi K2.5Moonshot AI · Open weight
85.0%
10
Claude Opus 4.5Anthropic · Closed
84.8%
11
Qwen3.5 397BAlibaba · Open weight
83.6%
12
Gemma 4 12BGoogle · Open weight
72.0%
13
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weight
69.9%
14
BTL-4Bad Theory Labs · Open weight
66.1%
15
ZAYA1-8BZyphra · Open weight
65.8%
16
ZAYA1-74B-PreviewZyphra · Open weight
65.7%
17
LFM2.5-2.6BLiquidAI · Open weight
59.4%
18
Mellum2-12B-A2.5B-InstructJetBrains · Open weight
37.2%
19
MiniCPM5-1BOpenBMB · Open weight
33.5%

The published LiveCodeBench v6 snapshot places Sakana Fugu-Ultra first at 93.2%. The third row is 1.7 points behind. The broader top-10 range is 8.4 points, so many of the published results sit in a relatively narrow band.

19 models have been evaluated on LiveCodeBench v6. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. LiveCodeBench v6 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LiveCodeBench v6

Year

2026

Tasks

Fresh programming problems

Format

Provider-published v6 competitive programming results

Difficulty

Competitive programming level

Providers often publish a specific release or date window instead of the rolling aggregate. This route contains rows explicitly labeled v6 by their sources and excludes them from the weighted legacy lane.

BenchLM freshness & provenance

Version

LiveCodeBench v6 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LiveCodeBench v6 measure?

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Which model scores highest on LiveCodeBench v6?

Sakana Fugu-Ultra by Sakana AI currently leads with a score of 93.2% on LiveCodeBench v6.

How many models are evaluated on LiveCodeBench v6?

19 AI models have been evaluated on LiveCodeBench v6 on BenchLM.

Last updated: August 22, 2026 · BenchLM version LiveCodeBench v6 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.