Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

BenchLM recommendation

Best AI Models for Web Research in 2026

Data verified

As of August 21, 2026, the top model in best ai models for web research on the BenchLM leaderboard is GPT-5.6 Sol with a score of 92.2.

Last verified: August 21, 2026

This reporting page isolates the web research slice of agentic performance. It prioritizes sourced benchmarks for browsing, evidence gathering, and multi-step web task completion rather than generic overall agent scores.

This page ranks models using only sourced web research benchmarks in the reporting family.

Bottom line: Web research agents need to browse, gather evidence, and synthesize findings. BrowseComp is the most predictive benchmark here.

GPT-5.6 Sol leads this ranking with a score of 92.2, followed by Kimi K3 (91.2) and Claude Opus 5 (90.8). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Ornith-1.5-397B (ranked #8 with a score of 86.6). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.

This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

How to choose

Full Rankings (38 models)

1
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

92.2

sourced avg

2
Kimi K3
Moonshot AI·Pending·1.05M

91.2

sourced avg

3
Claude Opus 5
Anthropic·Proprietary·

90.8

sourced avg

4
GPT-5.5 Pro
OpenAI·Proprietary·1M

90.1

sourced avg

5
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

89.3

sourced avg

6
Claude Mythos 5
Anthropic·Proprietary·1M+

88

sourced avg

7
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

87.5

sourced avg

8
Ornith-1.5-397B
Ornith AI·Open Weight·262K

86.6

sourced avg

9
Claude Sonnet 5
Anthropic·Proprietary·1M

84.7

sourced avg

10
GPT-5.5
OpenAI·Proprietary·1M

84.4

sourced avg

11
Claude Opus 4.8
Anthropic·Proprietary·1M

84.3

sourced avg

12
Claude Opus 4.6
Anthropic·Proprietary·1M

83.7

sourced avg

13
MiniMax M3
MiniMax·Open Weight·1M

83.5

sourced avg

14
DeepSeek V4 Pro 0813
DeepSeek·Proprietary·1M

83.4

sourced avg

15
dots3-note Preview
Dots Studio·Open Weight·512K

83.3

sourced avg

16
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

83.3

sourced avg

17
Kimi K2.6
Moonshot AI·Open Weight·256K

83.2

sourced avg

18
GPT-5.4
OpenAI·Proprietary·1.05M

82.7

sourced avg

19
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

79.3

sourced avg

20
Inkling-Small
Thinking Machines Lab·Open Weight·1M

77.4

sourced avg

21
Inkling
Thinking Machines Lab·Open Weight·1M

77.1

sourced avg

22
Step 3.7 Flash
StepFun·Open Weight·256K

75.8

sourced avg

23
Agents-A1
InternScience·Open Weight·262K

75.5

sourced avg

24
DeepSeek V4 Flash 0731
DeepSeek·Proprietary·1M

73.2

sourced avg

25
Ling 3.0 Flash
InclusionAI·Open Weight·262K

72.2

sourced avg

26
GLM-5.1
Z.AI·Open Weight·203K

68

sourced avg

27
Ornith-1.5-35B-A3B
Ornith AI·Open Weight·262K

67.6

sourced avg

28
GPT-5.2
OpenAI·Proprietary·400K

65.8

sourced avg

29
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

63.8

sourced avg

30
Qwen3.5 397B
Alibaba·Open Weight·128K

62

sourced avg

31
Qwen3.5-27B
Alibaba·Open Weight·262K

61

sourced avg

32
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

61

sourced avg

33
Kimi K2.5
Moonshot AI·Open Weight·256K

60.6

sourced avg

34
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

60.6

sourced avg

35
Ornith-1.5-9B
Ornith AI·Open Weight·262K

56.4

sourced avg

36
GLM-4.7
Z.AI·Open Weight·200K

52

sourced avg

37
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

44.4

sourced avg

38

36.8

sourced avg

Key Takeaways

The top model on this sourced reporting-family slice is GPT-5.6 Sol by OpenAI with an average of 92.2.

The best open-weight model is Ornith-1.5-397B at position #8.

38 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

This ranking averages sourced web-research benchmarks. It isolates the browsing and evidence-gathering slice of agentic performance.

Known limitations

Web research benchmarks test specific browsing patterns. Real-world web research also depends on access to search APIs, page rendering quality, and anti-bot measures that benchmarks do not capture.

Last updated: August 21, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.