Skip to main content
BenchLM
Data
Recommendation

Best AI Models for Web Research in 2026

As of October 5, 2026, the top model in best ai models for web research on the BenchLM leaderboard is Atria Dawn Preview with a score of 92.5.

Bottom line: Web research agents need to browse, gather evidence, and synthesize findings. BrowseComp is the most predictive benchmark here.

Ranking data as of

Full Rankings (47 models)

1
Atria Dawn Preview
Shanghai Artificial Intelligence Laboratory·Open Weight·256K

92.5

sourced avg

2
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

92.2

sourced avg

3
GPT-6 Astra
OpenAI·Proprietary·1.05M

91.5

sourced avg

4
Kimi K3
Moonshot AI·Pending·1.05M

91.2

sourced avg

5
Claude Opus 5
Anthropic·Proprietary·

90.8

sourced avg

6
GPT-5.5 Pro
OpenAI·Proprietary·1M

90.1

sourced avg

7
Fara1.5-27B
Microsoft·Open Weight·262K

89.3

sourced avg

8
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

89.3

sourced avg

9
Step 5 Preview
StepFun·Pending·1M

88.7

sourced avg

10
Claude Mythos 5
Anthropic·Proprietary·1M+

88

sourced avg

11
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

87.5

sourced avg

12
Ornith-1.5-397B
Ornith AI·Open Weight·262K

86.6

sourced avg

13
Claude Sonnet 5
Anthropic·Proprietary·1M

84.7

sourced avg

14
GPT-5.5
OpenAI·Proprietary·1M

84.4

sourced avg

15
Claude Opus 4.8
Anthropic·Proprietary·1M

84.3

sourced avg

16
Claude Opus 4.6
Anthropic·Proprietary·1M

83.7

sourced avg

17
MiniMax M3
MiniMax·Open Weight·1M

83.5

sourced avg

18
DeepSeek V4 Pro 0813
DeepSeek·Open Weight·1M

83.4

sourced avg

19
dots3-note Preview
Dots Studio·Open Weight·512K

83.3

sourced avg

20
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

83.3

sourced avg

21
Kimi K2.6
Moonshot AI·Open Weight·256K

83.2

sourced avg

22
GPT-5.4
OpenAI·Proprietary·1.05M

82.7

sourced avg

23
Fara1.5-4B
Microsoft·Open Weight·262K

80.8

sourced avg

24
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

79.3

sourced avg

25
Beam
Reflection AI·Pending·

77.4

sourced avg

26
Inkling-Small
Thinking Machines Lab·Open Weight·1M

77.4

sourced avg

27
Inkling
Thinking Machines Lab·Open Weight·1M

77.1

sourced avg

28
Step 3.7 Flash
StepFun·Open Weight·256K

75.8

sourced avg

29
Agents-A1
InternScience·Open Weight·262K

75.5

sourced avg

30
DeepSeek V4 Flash 0731
DeepSeek·Open Weight·1M

73.2

sourced avg

31
Ling 3.0 Flash
InclusionAI·Open Weight·262K

72.2

sourced avg

32
GLM-5.1
Z.AI·Open Weight·203K

68

sourced avg

33
Ornith-1.5-35B-A3B
Ornith AI·Open Weight·262K

67.6

sourced avg

34
Agents-A1-4B
InternScience·Open Weight·262K

66.8

sourced avg

35
GPT-5.2
OpenAI·Proprietary·400K

65.8

sourced avg

36
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

63.8

sourced avg

37
Qwen3.5 397B
Alibaba·Open Weight·128K

62

sourced avg

38
Qwen3.5-27B
Alibaba·Open Weight·262K

61

sourced avg

39
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

61

sourced avg

40
Kimi K2.5
Moonshot AI·Open Weight·256K

60.6

sourced avg

41
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

60.6

sourced avg

42
Ornith-1.5-9B
Ornith AI·Open Weight·262K

56.4

sourced avg

43
GLM-4.7
Z.AI·Open Weight·200K

52

sourced avg

44
Solar Pro 4
Upstage·Proprietary·512K

49.2

sourced avg

45
LongCat-Flash-Lite-Sparse
Meituan·Open Weight·1M

48.6

sourced avg

46
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

44.4

sourced avg

47

36.8

sourced avg

How to choose

Key Takeaways

The top model on this sourced reporting-family slice is Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory with an average of 92.5.

The best open-weight model is Atria Dawn Preview at position #1.

47 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

This ranking averages sourced web-research benchmarks. It isolates the browsing and evidence-gathering slice of agentic performance.

Known limitations

Web research benchmarks test specific browsing patterns. Real-world web research also depends on access to search APIs, page rendering quality, and anti-bot measures that benchmarks do not capture.

About this ranking

Ranking data as of October 5, 2026

This reporting page isolates the web research slice of agentic performance. It prioritizes sourced benchmarks for browsing, evidence gathering, and multi-step web task completion rather than generic overall agent scores.

This page ranks models using only sourced web research benchmarks in the reporting family.

Atria Dawn Preview leads this ranking with a score of 92.5, followed by GPT-5.6 Sol (92.2) and GPT-6 Astra (91.5). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Atria Dawn Preview (ranked #1 with a score of 92.5). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Explore More

Last updated: October 5, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.