Skip to main content
BenchLM
Recommendation

Best LLMs for Research in 2026

As of September 29, 2026, the top model in best llms for research on the BenchLM leaderboard is Atria Dawn Preview with a score of 93.9.

Bottom line: research is where frontier reasoning models earn their premium — the HLE and BrowseComp leaders below are the models that can both know and find.

Ranking data as of

Full Rankings (87 models)

1
Atria Dawn Preview
Shanghai Artificial Intelligence Laboratory·Open Weight·256K

93.9

sourced avg

2
GPT-6 Astra
OpenAI·Proprietary·1.05M

93.8

sourced avg

3
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

93.4

sourced avg

4
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

90.2

sourced avg

5
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

87.8

sourced avg

6
MAI-Thinking-1
Microsoft·Proprietary·256K

84.4

sourced avg

7
MiMo-V2-Flash
Xiaomi·Open Weight·256K

84

sourced avg

8
Kimi K3
Moonshot AI·Pending·1.05M

82.9

sourced avg

9
Step 3.7 Flash
StepFun·Open Weight·256K

82.6

sourced avg

10
Claude Mythos 5
Anthropic·Proprietary·1M+

82.2

sourced avg

11
Claude Opus 5
Anthropic·Proprietary·

82.1

sourced avg

12
Claude Opus 4.8
Anthropic·Proprietary·1M

81.2

sourced avg

13
GPT-5.2
OpenAI·Proprietary·400K

79.1

sourced avg

14
Qwen3 235B 2507
Alibaba·Open Weight·128K

78.9

sourced avg

15
Gemma 4 12B
Google·Open Weight·256K

78.4

sourced avg

16
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

76.8

sourced avg

17
GPT-5.5
OpenAI·Proprietary·1M

76.7

sourced avg

18
Claude Opus 4.6
Anthropic·Proprietary·1M

76.1

sourced avg

19
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

76.1

sourced avg

20
Kimi K2.5 (Reasoning)
Moonshot AI·Proprietary·128K

76

sourced avg

21
GPT-5.4
OpenAI·Proprietary·1.05M

75.5

sourced avg

22
Qwen3.5-27B
Alibaba·Open Weight·262K

75.1

sourced avg

23
Ornith-1.5-397B
Ornith AI·Open Weight·262K

74.7

sourced avg

24
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

74.4

sourced avg

25
dots3-note Preview
Dots Studio·Open Weight·512K

74

sourced avg

26
Hy4 preview
Tencent·Open Weight·1M

73.9

sourced avg

27
Kimi K2.6
Moonshot AI·Open Weight·256K

73.7

sourced avg

28
DeepSeek V4 Pro 0813
DeepSeek·Open Weight·1M

73.6

sourced avg

29
GPT-5.5 Pro
OpenAI·Proprietary·1M

73.6

sourced avg

30
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

73.5

sourced avg

31
GLM-5.2
Z.AI·Open Weight·1M

73

sourced avg

32
ZAYA1-8B
Zyphra·Open Weight·131K

71.8

sourced avg

33
Inkling-Small
Thinking Machines Lab·Open Weight·1M

71.6

sourced avg

34
Muse Spark 1.1
Meta·Proprietary·1M

71.2

sourced avg

35
Claude Sonnet 5
Anthropic·Proprietary·1M

71.1

sourced avg

36
Claude Sonnet 4.6
Anthropic·Proprietary·200K

70.8

sourced avg

37
GLM-5
Z.AI·Open Weight·200K

70.7

sourced avg

38
Inkling
Thinking Machines Lab·Open Weight·1M

70.3

sourced avg

39
Qwen3.7 Max
Alibaba·Proprietary·1M

70.1

sourced avg

40
Granite 4.2 30B
IBM·Open Weight·128K

69.2

sourced avg

41
Apodex 1.1
Apodex·Proprietary·N/A

68.5

sourced avg

42
Qwen3.8 Max
Alibaba·Open Weight·1M

68.1

sourced avg

43
Step 5 Preview
StepFun·Pending·1M

67.6

sourced avg

44
DeepSeek V4 Flash 0731
DeepSeek·Open Weight·1M

67.5

sourced avg

45
Granite 4.2 8B
IBM·Open Weight·128K

66.6

sourced avg

46
Gemini 3.5 Flash
Google·Proprietary·1M

66.2

sourced avg

47
Qwen3.7 Plus
Alibaba·Proprietary·1M

66.2

sourced avg

48
GPT-5.4 mini
OpenAI·Proprietary·400K

64.8

sourced avg

49
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

64.7

sourced avg

50
Kimi K2.5
Moonshot AI·Open Weight·256K

64.7

sourced avg

51
DeepSeek V4.1 Flash
DeepSeek·Open Weight·1M

63.9

sourced avg

52
Qwen3.8-Flash-Next
Alibaba·Open Weight·262K

63.8

sourced avg

53
Qwen3.8-Omni-Flash
Alibaba·Proprietary·1M

63.8

sourced avg

54
Qwen3.6 Plus
Alibaba·Proprietary·1M

63.7

sourced avg

55
Claude Opus 4.5
Anthropic·Proprietary·200K

63.3

sourced avg

56
DeepSeek V3
DeepSeek·Open Weight·128K

63.3

sourced avg

57
Grok 4.3
xAI·Proprietary·1M

62.5

sourced avg

58
Qwen3.5 397B
Alibaba·Open Weight·128K

62.5

sourced avg

59
Agents-A1
InternScience·Open Weight·262K

61.6

sourced avg

60
Gemma 4 E4B
Google·Open Weight·128K

61.3

sourced avg

61
Ornith-1.5-35B-A3B
Ornith AI·Open Weight·262K

60.8

sourced avg

62
Agents-A1-4B
InternScience·Open Weight·262K

60.4

sourced avg

63
GPT-5.4 nano
OpenAI·Proprietary·400K

60.3

sourced avg

64
GLM-5.1
Z.AI·Open Weight·203K

60.2

sourced avg

65
Muse Spark
Meta·Proprietary·262K

60.2

sourced avg

66
Qwen3.6-27B
Alibaba·Open Weight·262K

60.2

sourced avg

67
Ling 3.0 Flash
InclusionAI·Open Weight·262K

60

sourced avg

68
Qwen3.8-27B
Alibaba·Open Weight·262K

60

sourced avg

69
ZAYA1-74B-Preview
Zyphra·Open Weight·256K

60

sourced avg

70

59.8

sourced avg

71
Gemma 4 31B
Google·Open Weight·256K

59.7

sourced avg

72
Solar Pro 4
Upstage·Proprietary·512K

58.5

sourced avg

73
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

58.2

sourced avg

74
Granite 4.2 3B
IBM·Open Weight·128K

58.1

sourced avg

75
GLM-4.7
Z.AI·Open Weight·200K

57.2

sourced avg

76
Hy3 Preview
Tencent·Open Weight·256K

56.4

sourced avg

77
LongCat-Flash-Lite-Sparse
Meituan·Open Weight·1M

56.3

sourced avg

78
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

56.1

sourced avg

79
Ornith-1.5-9B
Ornith AI·Open Weight·262K

54.3

sourced avg

80
Gemini 2.5 Pro
Google·Proprietary·1M

50.9

sourced avg

81
Gemma 4 E2B
Google·Open Weight·128K

47.6

sourced avg

82
Soofi S 30B-A3B
Soofi Project·Open Weight·1M

45.4

sourced avg

83
K-EXAONE 2.0
LG AI Research·Open Weight·262K

34.6

sourced avg

84
Gemma 4 26B A4B
Google·Open Weight·256K

33.6

sourced avg

85
MiniCPM5-2B
OpenBMB·Open Weight·131K

24.4

sourced avg

86
LFM2.5-230M
LiquidAI·Open Weight·32K

24.1

sourced avg

87
LFM2.5-VL-450M
LiquidAI·Open Weight·128K

24.1

sourced avg

How to choose

Key Takeaways

The top model on this sourced reporting-family slice is Atria Dawn Preview by Shanghai Artificial Intelligence Laboratory with an average of 93.9.

The best open-weight model is Atria Dawn Preview at position #1.

87 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

This is a reporting-family ranking: a weighted average of sourced hard-knowledge and agentic-research benchmarks. It rewards models that combine deep knowledge with the ability to search, browse, and synthesize.

Known limitations

Research quality also depends on the harness (search tools, retrieval, citations UI), which benchmarks only partly capture. Models need sourced coverage on at least a quarter of the family to appear.

About this ranking

Ranking data as of September 29, 2026

Research work stresses two things at once: deep, reliable knowledge (GPQA, Humanity's Last Exam, frontier-science evaluations) and the ability to actually go find and synthesize sources (BrowseComp, DeepSearch-QA, GAIA). This reporting family blends both, weighted toward the hard-knowledge and browsing benchmarks that separate research-grade models from good chat models.

This page ranks models using only sourced benchmarks in the research reporting family — hard knowledge plus agentic web research — rather than the full provisional leaderboard.

Atria Dawn Preview leads this ranking with a score of 93.9, followed by GPT-6 Astra (93.8) and GPT-5.6 Sol (93.4). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Atria Dawn Preview (ranked #1 with a score of 93.9). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Questions

What is the best LLM for research?

The top rows of this table lead the sourced blend of hard-knowledge (GPQA, HLE) and agentic-research (BrowseComp, DeepSearch-QA) benchmarks — the two capabilities research work actually stresses. The ranking recomputes on every data refresh; check the answer box above for the current leader.

What is the best AI for deep research?

Deep-research products bundle a model with a browsing-and-synthesis harness, so pick from the BrowseComp and DeepSearch-QA leaders here, then compare the products built on them. A strong model in a weak harness will still miss sources; benchmark scores set the ceiling, the product sets how close you get.

Can I trust LLM citations in research?

Only after verification. Even the top HLE scorers fabricate citations at a nonzero rate, and browsing-enabled models can misread the sources they find. Use models from this table to draft and discover, and verify every load-bearing citation before it ships — the leaders lower the error rate, none eliminate it.

What is the best free LLM for research?

The strongest open-weight rows in this family can be self-hosted at no per-token cost — check which open models appear in the table, then see the open-source rankings and local LLM guide for hardware requirements. For occasional use, most frontier providers offer rate-limited free tiers of their chat products.

Explore More

Last updated: September 29, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.