# Chinese-SimpleQA

> A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.

Canonical page: https://benchlm.ai/benchmarks/chinesesimpleqa

- Category: [Knowledge](/knowledge)
- Last updated: September 15, 2026

## About Chinese-SimpleQA

- Year: 2026
- Tasks: Chinese factual questions
- Format: Short-form factual QA
- Difficulty: Factual accuracy focused
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores Chinese-SimpleQA as a display-only provider-table reference for DeepSeek-V4. It is separate from the English SimpleQA row.

Chinese-SimpleQA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 84.4% |
| 2 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 78.9% |

## FAQ

### What does Chinese-SimpleQA measure?

A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.

### Which model scores highest on Chinese-SimpleQA?

DeepSeek V4 Pro 0813 by DeepSeek currently leads with a score of 84.4% on Chinese-SimpleQA.

### How many models are evaluated on Chinese-SimpleQA?

2 AI models have been evaluated on Chinese-SimpleQA on BenchLM.

### Does Chinese-SimpleQA affect BenchLM's overall score?

Not directly. Chinese-SimpleQA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on Chinese-SimpleQA

- [DeepSeek V4 Pro 0813 vs DeepSeek V4 Flash 0731](/compare/deepseek-v4-flash-0731-vs-deepseek-v4-pro-0813)
