Skip to main content
BenchLM

DeepSearchQA

Data verified 40 confirmed releases in the last 30 daysFollow model changes

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

Top models on DeepSearchQA — September 22, 2026

As of September 22, 2026, Atria Dawn Preview leads the DeepSearchQA leaderboard with 96.0% , followed by Claude Opus 5 (95.0%) and Kimi K3 (95.0%).

19 modelsAgentic2% of Agentic reference weightCurrentUpdated September 22, 2026

Leaderboard (19 models)

Score
1
Atria Dawn PreviewShanghai Artificial Intelligence Laboratory · Open weight
96.0%
2
Claude Opus 5Anthropic · Closed
95.0%
3
Kimi K3Moonshot AI · Closed
95.0%
4
Claude Opus 4.8Anthropic · Closed
93.1%
5
Step 3.7 FlashStepFun · Open weight
92.8%
6
Kimi K2.6Moonshot AI · Open weight
92.5%
7
Apodex 1.1Apodex · Closed
92.4%
8
dots3-note PreviewDots Studio · Open weight
92.1%
9
Muse Spark 1.3Meta · Closed
89.4%
10
Muse Spark 1.1Meta · Closed
84.9%
11
Kimi K2.5Moonshot AI · Open weight
77.1%
12
Muse SparkMeta · Closed
74.8%
13
Muse Glimmer 30BMeta · Open weight
74.6%
14
Claude Opus 4.6Anthropic · Closed
73.7%
15
GPT-5.4OpenAI · Closed
73.6%
16
Gemini 3.1 ProGoogle · Closed
69.7%
17
Grok 4.20xAI · Closed
62.8%
18
Mercury 2.5Inception · Closed
34.0%
19
Mercury 2Inception · Closed
34.0%

According to BenchLM.ai, Atria Dawn Preview leads the DeepSearchQA benchmark with a score of 96.0%, followed by Claude Opus 5 (95.0%) and Kimi K3 (95.0%). The top models are clustered within 1.0 points, suggesting this benchmark is nearing saturation for frontier models.

19 models have been evaluated on DeepSearchQA. The benchmark falls in the Agentic category. BenchAlign v5.6 gives DeepSearchQA 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About DeepSearchQA

Year

2026

Tasks

Agentic browsing and list-answer questions

Format

Search / open / find browser-agent evaluation

Difficulty

Agentic web research

Meta describes DeepSearchQA as a browser-tool evaluation graded with an F1-style semantic set match. BenchLM stores it as an agentic search benchmark.

Freshness and provenance

Version

DeepSearchQA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does DeepSearchQA measure?

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

Which model scores highest on DeepSearchQA?

Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory currently leads with a score of 96.0% on DeepSearchQA.

How many models are evaluated on DeepSearchQA?

19 AI models have been evaluated on DeepSearchQA on BenchLM.

Last updated: September 22, 2026 · BenchLM version DeepSearchQA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.