Measuring Short-Form Factuality in Large Language Models (SimpleQA)
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
Top models on SimpleQA — September 18, 2026
As of September 18, 2026, DeepSeek V4 Pro 0813 leads the SimpleQA leaderboard with 57.9% , followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%).
DeepSeek V4 Pro 0813
DeepSeek
DeepSeek V4 Pro Base
DeepSeek
DeepSeek V4 Pro (High)
DeepSeek
9 modelsKnowledge5% of category scoreRefreshingUpdated September 18, 2026
Leaderboard (9 models)
ScoreAccording to BenchLM.ai, DeepSeek V4 Pro 0813 leads the SimpleQA benchmark with a score of 57.9%, followed by DeepSeek V4 Pro Base (55.2%) and DeepSeek V4 Pro (High) (46.2%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
9 models have been evaluated on SimpleQA. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, SimpleQA contributes 5% of the category score, so strong performance here directly affects a model's overall ranking.
About SimpleQA
Year
2024
Tasks
Factual questions
Format
Short-form Q&A
Difficulty
Factual accuracy focused
SimpleQA prioritizes two key properties: questions should have short, factual answers that can be easily verified, and questions should be diverse and challenging. It serves as a crucial test of factual knowledge and accuracy.
Freshness and provenance
Version
SimpleQA 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does SimpleQA measure?
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
Which model scores highest on SimpleQA?
DeepSeek V4 Pro 0813 by DeepSeek currently leads with a score of 57.9% on SimpleQA.
How many models are evaluated on SimpleQA?
9 AI models have been evaluated on SimpleQA on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.