Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Measuring Short-Form Factuality in Large Language Models (SimpleQA)

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Top models on SimpleQA — September 18, 2026

As of September 18, 2026, DeepSeek V4 Pro 0813 leads the SimpleQA leaderboard with 57.9% , followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%).

1Proprietary

DeepSeek V4 Pro 0813

DeepSeek

57.9%
Overall 63.69Context 1M
2Open weights

DeepSeek V4 Pro Base

DeepSeek

55.2%
Context 1M
3Open weights

DeepSeek V4 Pro (High)

DeepSeek

46.2%
Context 1M

9 modelsKnowledge5% of category scoreRefreshingUpdated September 18, 2026

Leaderboard (9 models)

Score
1
DeepSeek V4 Pro 0813DeepSeek · Closed
57.9%
2
DeepSeek V4 Pro BaseDeepSeek · Open weight
55.2%
3
DeepSeek V4 Pro (High)DeepSeek · Open weight
46.2%
4
DeepSeek V4 ProDeepSeek · Open weight
45%
5
DeepSeek V4 Flash 0731DeepSeek · Closed
34.1%
6
MAI-Thinking-1Microsoft · Closed
31%
7
DeepSeek V4 Flash BaseDeepSeek · Open weight
30.1%
8
DeepSeek V4 Flash (High)DeepSeek · Closed
28.9%
9
DeepSeek V4 FlashDeepSeek · Closed
23.1%

According to BenchLM.ai, DeepSeek V4 Pro 0813 leads the SimpleQA benchmark with a score of 57.9%, followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

9 models have been evaluated on SimpleQA. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, SimpleQA contributes 5% of the category score, so strong performance here directly affects a model's overall ranking.

About SimpleQA

Year

2024

Tasks

Factual questions

Format

Short-form Q&A

Difficulty

Factual accuracy focused

SimpleQA prioritizes two key properties: questions should have short, factual answers that can be easily verified, and questions should be diverse and challenging. It serves as a crucial test of factual knowledge and accuracy.

Freshness and provenance

Version

SimpleQA 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SimpleQA measure?

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Which model scores highest on SimpleQA?

DeepSeek V4 Pro 0813 by DeepSeek currently leads with a score of 57.9% on SimpleQA.

How many models are evaluated on SimpleQA?

9 AI models have been evaluated on SimpleQA on BenchLM.

Last updated: September 18, 2026 · BenchLM version SimpleQA 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.