# WideResearch

> A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

Canonical page: https://benchlm.ai/benchmarks/wideresearch

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026

## About WideResearch

- Year: 2026
- Tasks: Open-ended research tasks
- Format: Multi-source research evaluation
- Difficulty: Broad research-agent workflows
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

WideResearch evaluates whether a model can sustain a research process over multiple sources and branches rather than answering from shallow retrieval. BenchLM tracks it as a display-only browsing and synthesis benchmark.

WideResearch is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (15 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Hy4 preview](/models/hy4-preview) | Tencent | 83.9% |
| 2 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 81.9% |
| 3 | [Atria Dawn Preview](/models/atria-dawn-preview) | Shanghai Artificial Intelligence Laboratory | 81.9% |
| 4 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 80.8% |
| 5 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 80.8% |
| 6 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 78.9% |
| 7 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 76.4% |
| 8 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 74.3% |
| 9 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 74.0% |
| 10 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 73.6% |
| 11 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 72.7% |
| 12 | [GLM-5](/models/glm-5) | Z.AI | 69.8% |
| 13 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 67.8% |
| 14 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 60.1% |
| 15 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 59.5% |

## FAQ

### What does WideResearch measure?

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

### Which model scores highest on WideResearch?

Hy4 preview by Tencent currently leads with a score of 83.9% on WideResearch.

### How many models are evaluated on WideResearch?

15 AI models have been evaluated on WideResearch on BenchLM.

### Does WideResearch affect BenchLM's overall score?

Not directly. WideResearch is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on WideResearch

- [Hy4 preview vs Qwen3.8 Max](/compare/hy4-preview-vs-qwen3-8-max)
- [Qwen3.8 Max vs Atria Dawn Preview](/compare/atria-dawn-preview-vs-qwen3-8-max)
- [Atria Dawn Preview vs Kimi K2.6](/compare/atria-dawn-preview-vs-kimi-2-6)
- [Kimi K2.6 vs Ornith-1.5-397B](/compare/kimi-2-6-vs-ornith-1-5-397b)
