DeepSearchQA
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
Top models on DeepSearchQA — September 22, 2026
As of September 22, 2026, Atria Dawn Preview leads the DeepSearchQA leaderboard with 96.0% , followed by Claude Opus 5 (95.0%) and Kimi K3 (95.0%).
Atria Dawn Preview
Shanghai Artificial Intelligence Laboratory
Claude Opus 5
Anthropic
Kimi K3
Moonshot AI
19 modelsAgentic2% of Agentic reference weightCurrentUpdated September 22, 2026
Leaderboard (19 models)
ScoreAccording to BenchLM.ai, Atria Dawn Preview leads the DeepSearchQA benchmark with a score of 96.0%, followed by Claude Opus 5 (95.0%) and Kimi K3 (95.0%). The top models are clustered within 1.0 points, suggesting this benchmark is nearing saturation for frontier models.
19 models have been evaluated on DeepSearchQA. The benchmark falls in the Agentic category. BenchAlign v5.6 gives DeepSearchQA 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.
About DeepSearchQA
Year
2026
Tasks
Agentic browsing and list-answer questions
Format
Search / open / find browser-agent evaluation
Difficulty
Agentic web research
Meta describes DeepSearchQA as a browser-tool evaluation graded with an F1-style semantic set match. BenchLM stores it as an agentic search benchmark.
Freshness and provenance
Version
DeepSearchQA 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does DeepSearchQA measure?
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
Which model scores highest on DeepSearchQA?
Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory currently leads with a score of 96.0% on DeepSearchQA.
How many models are evaluated on DeepSearchQA?
19 AI models have been evaluated on DeepSearchQA on BenchLM.
Compare top models on DeepSearchQA
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.