# BigCodeBench

> A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

Canonical page: https://benchlm.ai/benchmarks/bigcodebench

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About BigCodeBench

- Year: 2026
- Tasks: Code generation tasks
- Format: Pass@1
- Difficulty: Software engineering
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores BigCodeBench as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

BigCodeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Prism ML | 58.1% |

## FAQ

### What does BigCodeBench measure?

A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

### Which model scores highest on BigCodeBench?

Ternary Bonsai 2 27B by Prism ML currently leads with a score of 58.1% on BigCodeBench.

### How many models are evaluated on BigCodeBench?

1 AI models have been evaluated on BigCodeBench on BenchLM.

### Does BigCodeBench affect BenchLM's overall score?

Not directly. BigCodeBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
