Skip to main content
BenchLM

Parallel Search Capability Leaderboard (Parallel Search Capability)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.

Search Intelligence Score on Parallel Search Capability Leaderboard — September 22, 2026

We mirror the published search intelligence score view for Parallel Search Capability Leaderboard. Claude Opus 5.5 leads the public snapshot at 75.5, followed by Claude Opus 5 (71.2) and GPT-6 Astra (70.8). We do not use these results to rank models overall.

25 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026

Search Intelligence Score table (25 models)

Score
1
Claude Opus 5.5Anthropic · Closed+32.9 search lift · $1,656 / 1K tasks
75.5
2
Claude Opus 5Anthropic · Closed+30.8 search lift · $1,437 / 1K tasks
71.2
3
GPT-6 AstraOpenAI · Closed+27.9 search lift · $401 / 1K tasks
70.8
4
Claude Fable 5.1Anthropic · Closed+24.7 search lift · $3,101 / 1K tasks
69.3
5
GPT-5.6 SolOpenAI · Closed+26.3 search lift · $269 / 1K tasks
67.7
6
Gemini 3.7 FlashGoogle · Closed+27.4 search lift · $130 / 1K tasks
66.8
7
Muse Spark 1.3Meta · Closed+35.8 search lift · $142 / 1K tasks
66.1
8
Gemini 3.8 FlashGoogle · Closed+26.1 search lift · $177 / 1K tasks
65.1
9
Kimi K3Moonshot AI · Closed+33.6 search lift · $266 / 1K tasks
64.2
10
DeepSeek V4.1 FlashDeepSeek · Open weight+35.7 search lift · $35.8 / 1K tasks
62.7
11
GLM 5.3Z.ai+40.8 search lift · $78.5 / 1K tasks
62.7
12
GPT-5.6 LunaOpenAI · Closed+33.9 search lift · $36.2 / 1K tasks
60.7
13
Claude Sonnet 5Anthropic · Closed+38.1 search lift · $1,000 / 1K tasks
60.4
14
DeepSeek V4 Flash (0731)DeepSeek+35.2 search lift · $13.6 / 1K tasks
59.2
15
Gemini 3 FlashGoogle · Closed+24.5 search lift · $114 / 1K tasks
58.6
16
DeepSeek V4 ProDeepSeek+27.9 search lift · $110 / 1K tasks
58.4
17
Hunyuan 3Tencent+35.6 search lift · $28.4 / 1K tasks
58.2
18
GLM-5.2Z.AI · Open weight+35.9 search lift · $66.6 / 1K tasks
54.3
19
DeepSeek V4 Pro (0813)DeepSeek+20.4 search lift · $234 / 1K tasks
53.5
20
DeepSeek V4 FlashDeepSeek+33.7 search lift · $24.4 / 1K tasks
53.1
21
Nemotron 3 Ultra 550BNVIDIA+37.7 search lift · $113 / 1K tasks
53.0
22
MiniMax M3MiniMax · Open weight+29.5 search lift · $41.1 / 1K tasks
52.6
23
MiMo v2.5Xiaomi+37.9 search lift · $23.5 / 1K tasks
49.5
24
Laguna S 2.1Poolside · Open weight+27.2 search lift · $21.4 / 1K tasks
39.5
25
Nemotron 3.5 LightningNVIDIA+28.0 search lift · $12.1 / 1K tasks
37.9

How to read this leaderboard

Compare the models within Parallel's fixed tool setup. Search lift shows the difference between the search and no-search runs; cost is reported per 1,000 tasks and includes inference plus estimated Search and Extract usage.

Operator receipt: 25 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5.5 at 75.5.

Honest limit: The composite includes Parallel's proprietary WISER set and uses Parallel's search tools for every model. Reasoning settings and token limits can vary by model. Parallel says cost records may omit recovery attempts. The table remains display only and does not enter overall or category rankings.

How to read the search capability table

We captured 25 rows from Parallel's September 22, 2026 Search Capability Leaderboard. The score equally weights DeepSearchQA F1, Humanity's Last Exam accuracy, and WISER accuracy. Each model was tested on 100 questions from each set with Parallel Search Fast and Extract.

The score reflects a model using Parallel's search tools in this test setup. Parallel also reports a no-search score, lift, cost per 1,000 tasks, and time per task. The table is display only and does not enter overall or category rankings.

Snapshot

25 model rows3 benchmark components100 questions per componentDisplay only

The published Parallel Search Capability snapshot places Claude Opus 5.5 first at 75.5. The third row is 4.7 score units behind. The broader top-10 range is 12.8 score units, so the table still separates the published systems.

25 models have been evaluated on Parallel Search Capability. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Parallel Search Capability is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Parallel Search Capability

Year

2026

Tasks

100 questions each from DeepSearchQA, HLE, and WISER

Format

Equal-weight Search Intelligence Score (0–100)

Difficulty

Multi-step web research and specialist questions

Parallel runs 100 questions from each of DeepSearchQA, Humanity's Last Exam, and WISER with and without Parallel Search Fast and Extract. The Search Intelligence Score equally weights DSQA F1, HLE accuracy, and WISER accuracy. We mirror the 25 published composite scores and preserve no-search scores, search lift, cost, and time in the source snapshot.

Freshness and provenance

Version

Parallel Search Capability 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Parallel Search Capability measure?

A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.

Which model leads the published Parallel Search Capability snapshot?

Claude Opus 5.5 currently leads the published Parallel Search Capability snapshot with 75.5 search intelligence score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Parallel Search Capability?

The September 22, 2026 snapshot contains 25 AI models.

Last updated: September 22, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.