Skip to main content
BenchLM

BigCodeBench

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

Benchmark score on BigCodeBench — September 27, 2026

We compile the BigCodeBench rows from provider self-reports. DeepSeek V4 Pro Base leads the table at 59.2%, followed by Ternary Bonsai 2 27B (58.1%) and DeepSeek V4 Flash Base (56.8%). We do not use these results to rank models overall.

1Open weights

DeepSeek V4 Pro Base

DeepSeek

59.2%
Context 1M
2Open weights

Ternary Bonsai 2 27B

Prism ML

58.1%
Overall 52.04Context 262K
3Open weights

DeepSeek V4 Flash Base

DeepSeek

56.8%
Context 1M

3 modelsCodingCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (3 models)

Score
1
DeepSeek V4 Pro BaseDeepSeek · Open weight
59.2%
2
Ternary Bonsai 2 27BPrism ML · Open weight
58.1%
3
DeepSeek V4 Flash BaseDeepSeek · Open weight
56.8%

Among the reported BigCodeBench rows, DeepSeek V4 Pro Base is first at 59.2%. The third row is 2.4 points behind. The broader top-10 range is 2.4 points, so many of the published results sit in a relatively narrow band.

3 models have been evaluated on BigCodeBench. The benchmark falls in the Coding category. BigCodeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About BigCodeBench

Year

2026

Tasks

Code generation tasks

Format

Pass@1

Difficulty

Software engineering

BenchLM stores BigCodeBench as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

Freshness and provenance

Version

BigCodeBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does BigCodeBench measure?

A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

Which model scores highest on BigCodeBench?

DeepSeek V4 Pro Base by DeepSeek currently leads with a score of 59.2% on BigCodeBench.

How many models are evaluated on BigCodeBench?

3 AI models have been evaluated on BigCodeBench on BenchLM.

Last updated: September 27, 2026 · BenchLM version BigCodeBench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.