Benchmark profile
Vals Web Search Index (Web Search Index)
A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.
Data verifiedHow BenchLM shows Web Search Index
BenchLM mirrors the public Vals AI Web Search Index leaderboard captured from https://www.vals.ai/benchmarks/web_search and updated by Vals on July 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Web Search Index is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Web Search Index score on Web Search Index — July 16, 2026
BenchLM mirrors the published web search index score view for Web Search Index. Claude Fable 5 Exa leads the public snapshot at 48.45% , followed by Claude Fable 5 (46.94%) and GPT-5.6 Sol Exa (45.24%). BenchLM does not use these results to rank models overall.
Claude Fable 5 Exa
Anthropic
Exa
anthropic/claude-fable-5-exa
Claude Fable 5
Anthropic
Native
anthropic/claude-fable-5
GPT-5.6 Sol Exa
OpenAI
Exa
openai/gpt-5.6-sol-exa
Web Search Index score table (8 models)
ScoreThe published Web Search Index snapshot places Claude Fable 5 Exa first at 48.45%. The third row is 3.21 points behind. The broader top-10 range is 12.92 points, so the table still separates the published systems.
8 models have been evaluated on Web Search Index. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Web Search Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Web Search Index
Year
2026
Tasks
Finance Agent Benchmark v2 and Legal Research Benchmark tasks
Format
Accuracy by model and search-tool combination
Difficulty
Professional web research with controlled search-tool variants
The index holds each model and the rest of the agent harness constant while swapping the search tool. We mirror the native and Exa rows, task scores, uncertainty, latency, and cost per task. The private dataset and tool-specific runs keep the table outside weighted model rankings.
BenchLM freshness & provenance
Version
Web Search Index 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Web Search Index measure?
A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.
Which model leads the published Web Search Index snapshot?
Claude Fable 5 Exa currently leads the published Web Search Index snapshot with 48.45% web search index score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Web Search Index?
8 AI models are included in BenchLM's mirrored Web Search Index snapshot, based on the public leaderboard captured on July 16, 2026.