Skip to main content
BenchLM

Vals Web Search Index (Web Search Index)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.

Web Search Index score on Web Search Index — July 16, 2026

We mirror the published web search index score view for Web Search Index. Claude Fable 5 Exa leads the public snapshot at 48.45%, followed by Claude Fable 5 (46.94%) and GPT-5.6 Sol Exa (45.24%). We do not use these results to rank models overall.

8 modelsAgenticCurrentDisplay onlyUpdated July 16, 2026

Web Search Index score table (8 models)

Score
1
Claude Fable 5 ExaAnthropicExaExa
48.45%
2
Claude Fable 5Anthropic · ClosedNativeNative
46.94%
3
GPT-5.6 Sol ExaOpenAIExa · max reasoningExa
45.24%
4
GPT-5.6 SolOpenAI · ClosedNative · max reasoningNative
43.58%
5
Gemini 3.5 Flash ExaGoogleExa · high reasoningExa
41.36%
6
Grok 4.5 ExaxAIExa · high reasoningExa
38.75%
7
Gemini 3.5 FlashGoogle · ClosedNative · high reasoningNative
37.04%
8
Grok 4.5xAI · ClosedNative · high reasoningNative
35.54%

How Web Search Index is shown here

BenchLM mirrors the public Vals AI Web Search Index leaderboard captured from https://www.vals.ai/benchmarks/web_search and updated by Vals on July 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Web Search Index is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

8 Vals rows3 task viewsprivate datasetTasks: Overall, Finance Analysis, Legal ResearchDisplay only

The published Web Search Index snapshot places Claude Fable 5 Exa first at 48.45%. The third row is 3.21 points behind. The broader top-10 range is 12.92 points, so the table still separates the published systems.

8 models have been evaluated on Web Search Index. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Web Search Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Web Search Index

Year

2026

Tasks

Finance Agent Benchmark v2 and Legal Research Benchmark tasks

Format

Accuracy by model and search-tool combination

Difficulty

Professional web research with controlled search-tool variants

The index holds each model and the rest of the agent harness constant while swapping the search tool. We mirror the native and Exa rows, task scores, uncertainty, latency, and cost per task. The private dataset and tool-specific runs keep the table outside weighted model rankings.

Freshness and provenance

Version

Web Search Index 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Web Search Index measure?

A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.

Which model leads the published Web Search Index snapshot?

Claude Fable 5 Exa currently leads the published Web Search Index snapshot with 48.45% web search index score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Web Search Index?

The July 16, 2026 snapshot contains 8 AI models.

Last updated: July 16, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.