Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

Evaluating Large Language Models Trained on Code (HumanEval)

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

Data verified 27 confirmed releases in the last 30 daysSee the free Radar Brief

How to read this leaderboard

Read this page as a source-coverage receipt for the exact rows BenchLM can currently display, not as a ranking of the best coding models. Use the calibrated coding leaderboard and repository benchmarks for a broader decision.

Operator receipt: 3 sourced rows are currently displayable on this page; the leading published row is DeepSeek V4 Pro Base at 76.8%.

Honest limit: Only exact, displayable rows appear. Missing frontier rows are missing evidence, not zero scores, and the two visible DeepSeek base-model rows do not establish an overall coding winner.

Benchmark score on HumanEval — September 2, 2026

We mirror the published score view for HumanEval. DeepSeek V4 Pro Base leads the public snapshot at 76.8%, followed by Soofi S 30B-A3B (73.8%) and DeepSeek V4 Flash Base (69.5%). We do not use these results to rank models overall.

1Open weights

DeepSeek V4 Pro Base

DeepSeek

76.8%
Context 1M
2Open weights

Soofi S 30B-A3B

Soofi Project

73.8%
Context 1M
3Open weights

DeepSeek V4 Flash Base

DeepSeek

69.5%
Context 1M

3 modelsCodingStaleSaturatedDisplay onlyUpdated September 2, 2026

Benchmark score table (3 models)

Score
1
DeepSeek V4 Pro BaseDeepSeek · Open weight
76.8%
2
Soofi S 30B-A3BSoofi Project · Open weight
73.8%
3
DeepSeek V4 Flash BaseDeepSeek · Open weight
69.5%

The published HumanEval snapshot places DeepSeek V4 Pro Base first at 76.8%. The third row is 7.3 points behind. The broader top-10 range is 7.3 points, so many of the published results sit in a relatively narrow band.

3 models have been evaluated on HumanEval. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. HumanEval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About HumanEval

Year

2021

Tasks

164 problems

Format

Python function generation

Difficulty

Introductory to intermediate programming

HumanEval measures functional correctness for programs synthesized from docstrings. The tasks range from string manipulation to intermediate algorithms, but they do not test repository navigation, multi-file changes, tool use, or long-running agent work.

BenchLM freshness & provenance

Version

HumanEval

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleSaturatedDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HumanEval measure?

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

Which model scores highest on HumanEval?

DeepSeek V4 Pro Base by DeepSeek currently leads with a score of 76.8% on HumanEval.

How many models are evaluated on HumanEval?

3 AI models have been evaluated on HumanEval on BenchLM.

Last updated: September 2, 2026 · BenchLM version HumanEval

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.