Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

C-Eval

A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on C-Eval — September 15, 2026

We mirror the published score view for C-Eval. Qwen3.6 Plus leads the public snapshot at 93.3%, followed by DeepSeek V4 Pro Base (93.1%) and Qwen3.5 397B (93%). We do not use these results to rank models overall.

8 modelsKnowledgeStaleDisplay onlyUpdated September 15, 2026

Benchmark score table (8 models)

Score
1
Qwen3.6 PlusAlibaba · Closed
93.3%
2
DeepSeek V4 Pro BaseDeepSeek · Open weight
93.1%
3
Qwen3.5 397BAlibaba · Open weight
93%
4
Claude Opus 4.5Anthropic · Closed
92.2%
5
DeepSeek V4 Flash BaseDeepSeek · Open weight
92.1%
6
Qwen3.6-27BAlibaba · Open weight
91.4%
7
Qwen3.6-35B-A3BAlibaba · Open weight
90%
8
LongCat-Flash-Lite-SparseMeituan · Open weight
85.8%

The published C-Eval snapshot places Qwen3.6 Plus first at 93.3%. The third row is 0.3 points behind. The broader top-10 range is 7.5 points, so many of the published results sit in a relatively narrow band.

8 models have been evaluated on C-Eval. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. C-Eval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About C-Eval

Year

2023

Tasks

Chinese academic and professional exams

Format

Multiple choice questions

Difficulty

High school to professional level

C-Eval is one of the clearest public signals for non-English academic knowledge performance. It tests whether a model can sustain strong factual recall and reasoning under Chinese-language exam conditions across many domains.

BenchLM freshness & provenance

Version

C-Eval 2023

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does C-Eval measure?

A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.

Which model scores highest on C-Eval?

Qwen3.6 Plus by Alibaba currently leads with a score of 93.3% on C-Eval.

How many models are evaluated on C-Eval?

8 AI models have been evaluated on C-Eval on BenchLM.

Compare Top Models on C-Eval

Last updated: September 15, 2026 · BenchLM version C-Eval 2023

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.