Parallel Search Capability Leaderboard (Parallel Search Capability)
We show this table for reference; we do not rank on it.
A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.
Search Intelligence Score on Parallel Search Capability Leaderboard — September 22, 2026
We mirror the published search intelligence score view for Parallel Search Capability Leaderboard. Claude Opus 5.5 leads the public snapshot at 75.5, followed by Claude Opus 5 (71.2) and GPT-6 Astra (70.8). We do not use these results to rank models overall.
Claude Opus 5.5
Anthropic
+32.9 search lift · $1,656 / 1K tasks
Claude Opus 5
Anthropic
+30.8 search lift · $1,437 / 1K tasks
GPT-6 Astra
OpenAI
+27.9 search lift · $401 / 1K tasks
25 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026
Search Intelligence Score table (25 models)
ScoreHow to read this leaderboard
Compare the models within Parallel's fixed tool setup. Search lift shows the difference between the search and no-search runs; cost is reported per 1,000 tasks and includes inference plus estimated Search and Extract usage.
Operator receipt: 25 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5.5 at 75.5.
Honest limit: The composite includes Parallel's proprietary WISER set and uses Parallel's search tools for every model. Reasoning settings and token limits can vary by model. Parallel says cost records may omit recovery attempts. The table remains display only and does not enter overall or category rankings.
How to read the search capability table
We captured 25 rows from Parallel's September 22, 2026 Search Capability Leaderboard. The score equally weights DeepSearchQA F1, Humanity's Last Exam accuracy, and WISER accuracy. Each model was tested on 100 questions from each set with Parallel Search Fast and Extract.
The score reflects a model using Parallel's search tools in this test setup. Parallel also reports a no-search score, lift, cost per 1,000 tasks, and time per task. The table is display only and does not enter overall or category rankings.
Snapshot
The published Parallel Search Capability snapshot places Claude Opus 5.5 first at 75.5. The third row is 4.7 score units behind. The broader top-10 range is 12.8 score units, so the table still separates the published systems.
25 models have been evaluated on Parallel Search Capability. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Parallel Search Capability is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Parallel Search Capability
Year
2026
Tasks
100 questions each from DeepSearchQA, HLE, and WISER
Format
Equal-weight Search Intelligence Score (0–100)
Difficulty
Multi-step web research and specialist questions
Parallel runs 100 questions from each of DeepSearchQA, Humanity's Last Exam, and WISER with and without Parallel Search Fast and Extract. The Search Intelligence Score equally weights DSQA F1, HLE accuracy, and WISER accuracy. We mirror the 25 published composite scores and preserve no-search scores, search lift, cost, and time in the source snapshot.
Freshness and provenance
Version
Parallel Search Capability 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Parallel Search Capability measure?
A 25-model comparison of how well models use Parallel Search Fast and Extract to answer web research questions.
Which model leads the published Parallel Search Capability snapshot?
Claude Opus 5.5 currently leads the published Parallel Search Capability snapshot with 75.5 search intelligence score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Parallel Search Capability?
The September 22, 2026 snapshot contains 25 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.