Skip to main content

Benchmark profile

Measuring Short-Form Factuality in Large Language Models (SimpleQA)

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Data verified

Top models on SimpleQA — July 20, 2026

As of July 20, 2026, DeepSeek V4 Pro (Max) leads the SimpleQA leaderboard with 57.9% , followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%).

9 modelsKnowledge11% of category scoreRefreshingUpdated July 20, 2026

Leaderboard (9 models)

Score
1
DeepSeek V4 Pro (Max)DeepSeek · Open weight
57.9%
2
DeepSeek V4 Pro BaseDeepSeek · Open weight
55.2%
3
DeepSeek V4 Pro (High)DeepSeek · Open weight
46.2%
4
DeepSeek V4 ProDeepSeek · Open weight
45%
5
DeepSeek V4 Flash (Max)DeepSeek · Open weight
34.1%
6
MAI-Thinking-1Microsoft · Closed
31%
7
DeepSeek V4 Flash BaseDeepSeek · Open weight
30.1%
8
DeepSeek V4 Flash (High)DeepSeek · Open weight
28.9%
9
DeepSeek V4 FlashDeepSeek · Open weight
23.1%

According to BenchLM.ai, DeepSeek V4 Pro (Max) leads the SimpleQA benchmark with a score of 57.9%, followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

9 models have been evaluated on SimpleQA. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, SimpleQA contributes 11% of the category score, so strong performance here directly affects a model's overall ranking.

About SimpleQA

Year

2024

Tasks

Factual questions

Format

Short-form Q&A

Difficulty

Factual accuracy focused

SimpleQA prioritizes two key properties: questions should have short, factual answers that can be easily verified, and questions should be diverse and challenging. It serves as a crucial test of factual knowledge and accuracy.

BenchLM freshness & provenance

Version

SimpleQA 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does SimpleQA measure?

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Which model scores highest on SimpleQA?

DeepSeek V4 Pro (Max) by DeepSeek currently leads with a score of 57.9% on SimpleQA.

How many models are evaluated on SimpleQA?

9 AI models have been evaluated on SimpleQA on BenchLM.

Last updated: July 20, 2026 · BenchLM version SimpleQA 2024

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.