# DeepSearchQA

> An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

Canonical page: https://benchlm.ai/benchmarks/deepsearchqa

- Category: [Agentic](/agentic)
- Last updated: September 22, 2026

## About DeepSearchQA

- Year: 2026
- Tasks: Agentic browsing and list-answer questions
- Format: Search / open / find browser-agent evaluation
- Difficulty: Agentic web research
- Paper: [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology)

Meta describes DeepSearchQA as a browser-tool evaluation graded with an F1-style semantic set match. BenchLM stores it as an agentic search benchmark.

BenchAlign v5.6 gives DeepSearchQA 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (18 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Atria Dawn Preview](/models/atria-dawn-preview) | Shanghai Artificial Intelligence Laboratory | 96.0% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 95.0% |
| 3 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 95.0% |
| 4 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 93.1% |
| 5 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 92.8% |
| 6 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 92.5% |
| 7 | [Apodex 1.1](/models/apodex-1-1) | Apodex | 92.4% |
| 8 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 92.1% |
| 9 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 89.4% |
| 10 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 84.9% |
| 11 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 77.1% |
| 12 | [Muse Spark](/models/muse-spark) | Meta | 74.8% |
| 13 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 74.6% |
| 14 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 73.7% |
| 15 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 73.6% |
| 16 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 69.7% |
| 17 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 62.8% |
| 18 | [Mercury 2.5](/models/mercury-2-5) | Inception | 34.0% |

## FAQ

### What does DeepSearchQA measure?

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

### Which model scores highest on DeepSearchQA?

Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory currently leads with a score of 96.0% on DeepSearchQA.

### How many models are evaluated on DeepSearchQA?

18 AI models have been evaluated on DeepSearchQA on BenchLM.

### Does DeepSearchQA affect BenchLM's overall score?

Yes. BenchAlign v5.6 gives DeepSearchQA 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on DeepSearchQA

- [Atria Dawn Preview vs Claude Opus 5](/compare/atria-dawn-preview-vs-claude-opus-5)
- [Claude Opus 5 vs Kimi K3](/compare/claude-opus-5-vs-kimi-k3)
- [Kimi K3 vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-kimi-k3)
- [Claude Opus 4.8 vs Step 3.7 Flash](/compare/claude-opus-4-8-vs-step-3-7-flash)
