# Parallel Search Capability Leaderboard (Parallel Search Capability)

> A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.

Canonical page: https://benchlm.ai/benchmarks/parallel-search-capability

- Category: [Agentic](/agentic)
- Last updated: September 22, 2026

## About Parallel Search Capability

- Year: 2026
- Tasks: 100 questions each from DeepSearchQA, HLE, and WISER
- Format: Equal-weight Search Intelligence Score (0–100)
- Difficulty: Multi-step web research and specialist questions

Parallel runs 100 questions from each of DeepSearchQA, Humanity's Last Exam, and WISER with and without Parallel Search Fast and Extract. The Search Intelligence Score equally weights DSQA F1, HLE accuracy, and WISER accuracy. We mirror the 25 published composite scores and preserve no-search scores, search lift, cost, and time in the source snapshot.

Parallel Search Capability is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | +32.9 search lift · $1,656 / 1K tasks | Anthropic | 75.5 |
| 2 | [Claude Opus 5](/models/claude-opus-5) | +30.8 search lift · $1,437 / 1K tasks | Anthropic | 71.2 |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | +27.9 search lift · $401 / 1K tasks | OpenAI | 70.8 |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | +24.7 search lift · $3,101 / 1K tasks | Anthropic | 69.3 |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | +26.3 search lift · $269 / 1K tasks | OpenAI | 67.7 |
| 6 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | +27.4 search lift · $130 / 1K tasks | Google | 66.8 |
| 7 | [Muse Spark 1.3](/models/muse-spark-1-3) | +35.8 search lift · $142 / 1K tasks | Meta | 66.1 |
| 8 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | +26.1 search lift · $177 / 1K tasks | Google | 65.1 |
| 9 | [Kimi K3](/models/kimi-k3) | +33.6 search lift · $266 / 1K tasks | Moonshot AI | 64.2 |
| 10 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | +35.7 search lift · $35.8 / 1K tasks | DeepSeek | 62.7 |
| 11 | [GLM 5.3](https://parallel.ai/leaderboard) | +40.8 search lift · $78.5 / 1K tasks | Z.ai | 62.7 |
| 12 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | +33.9 search lift · $36.2 / 1K tasks | OpenAI | 60.7 |
| 13 | [Claude Sonnet 5](/models/claude-sonnet-5) | +38.1 search lift · $1,000 / 1K tasks | Anthropic | 60.4 |
| 14 | [DeepSeek V4 Flash (0731)](https://parallel.ai/leaderboard) | +35.2 search lift · $13.6 / 1K tasks | DeepSeek | 59.2 |
| 15 | [Gemini 3 Flash](/models/gemini-3-flash) | +24.5 search lift · $114 / 1K tasks | Google | 58.6 |
| 16 | [DeepSeek V4 Pro](https://parallel.ai/leaderboard) | +27.9 search lift · $110 / 1K tasks | DeepSeek | 58.4 |
| 17 | [Hunyuan 3](https://parallel.ai/leaderboard) | +35.6 search lift · $28.4 / 1K tasks | Tencent | 58.2 |
| 18 | [GLM-5.2](/models/glm-5-2) | +35.9 search lift · $66.6 / 1K tasks | Z.AI | 54.3 |
| 19 | [DeepSeek V4 Pro (0813)](https://parallel.ai/leaderboard) | +20.4 search lift · $234 / 1K tasks | DeepSeek | 53.5 |
| 20 | [DeepSeek V4 Flash](https://parallel.ai/leaderboard) | +33.7 search lift · $24.4 / 1K tasks | DeepSeek | 53.1 |
| 21 | [Nemotron 3 Ultra 550B](https://parallel.ai/leaderboard) | +37.7 search lift · $113 / 1K tasks | NVIDIA | 53.0 |
| 22 | [MiniMax M3](/models/minimax-m3) | +29.5 search lift · $41.1 / 1K tasks | MiniMax | 52.6 |
| 23 | [MiMo v2.5](https://parallel.ai/leaderboard) | +37.9 search lift · $23.5 / 1K tasks | Xiaomi | 49.5 |
| 24 | [Laguna S 2.1](/models/laguna-s-2-1) | +27.2 search lift · $21.4 / 1K tasks | Poolside | 39.5 |
| 25 | [Nemotron 3.5 Lightning](https://parallel.ai/leaderboard) | +28.0 search lift · $12.1 / 1K tasks | NVIDIA | 37.9 |

## FAQ

### What does Parallel Search Capability measure?

A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.

### Which model leads the published Parallel Search Capability snapshot?

Claude Opus 5.5 currently leads the published Parallel Search Capability snapshot with a score of 75.5.

### How many models are evaluated on Parallel Search Capability?

The September 22, 2026 contains 25 AI models.

### Does Parallel Search Capability affect BenchLM's overall score?

Not directly. Parallel Search Capability is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
