Vals Web Search Index (Web Search Index)
We show this table for reference; we do not rank on it.
A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.
Web Search Index score on Web Search Index — July 16, 2026
We mirror the published web search index score view for Web Search Index. Claude Fable 5 Exa leads the public snapshot at 48.45%, followed by Claude Fable 5 (46.94%) and GPT-5.6 Sol Exa (45.24%). We do not use these results to rank models overall.
Claude Fable 5 Exa
Anthropic
Exa
Exa
Claude Fable 5
Anthropic
Native
Native
GPT-5.6 Sol Exa
OpenAI
Exa · max reasoning
Exa
8 modelsAgenticCurrentDisplay onlyUpdated July 16, 2026
Web Search Index score table (8 models)
ScoreHow Web Search Index is shown here
BenchLM mirrors the public Vals AI Web Search Index leaderboard captured from https://www.vals.ai/benchmarks/web_search and updated by Vals on July 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Web Search Index is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Snapshot
The published Web Search Index snapshot places Claude Fable 5 Exa first at 48.45%. The third row is 3.21 points behind. The broader top-10 range is 12.92 points, so the table still separates the published systems.
8 models have been evaluated on Web Search Index. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Web Search Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Web Search Index
Year
2026
Tasks
Finance Agent Benchmark v2 and Legal Research Benchmark tasks
Format
Accuracy by model and search-tool combination
Difficulty
Professional web research with controlled search-tool variants
The index holds each model and the rest of the agent harness constant while swapping the search tool. We mirror the native and Exa rows, task scores, uncertainty, latency, and cost per task. The private dataset and tool-specific runs keep the table outside weighted model rankings.
Freshness and provenance
Version
Web Search Index 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Web Search Index measure?
A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.
Which model leads the published Web Search Index snapshot?
Claude Fable 5 Exa currently leads the published Web Search Index snapshot with 48.45% web search index score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Web Search Index?
The July 16, 2026 snapshot contains 8 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.