# BrowseComp

> A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Canonical page: https://benchlm.ai/benchmarks/browsecomp

- Category: [Agentic](/agentic)
- Last updated: September 10, 2026

## About BrowseComp

- Year: 2025
- Tasks: Research questions requiring browsing
- Format: Web search and evidence synthesis
- Difficulty: Hard web research
- Paper: [BrowseComp](https://openai.com/index/browsecomp/)

BrowseComp is designed to measure real web research behavior, not just latent world knowledge. It rewards models that can plan searches, inspect multiple pages, and avoid shallow answer synthesis.

BrowseComp is currently weighted in BenchLM's scoring formula. The Agentic category carries 22% of the overall score, and BrowseComp contributes 25% of that category score.

## Leaderboard (42 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 92.2% |
| 2 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 91.5% |
| 3 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 91.2% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 90.8% |
| 5 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | OpenAI | 90.1% |
| 6 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 89.3% |
| 7 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 88% |
| 8 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 87.5% |
| 9 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 86.6% |
| 10 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 84.7% |
| 11 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 84.4% |
| 12 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 84.3% |
| 13 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 83.7% |
| 14 | [MiniMax M3](/models/minimax-m3) | MiniMax | 83.5% |
| 15 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 83.4% |
| 16 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 83.3% |
| 17 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 83.3% |
| 18 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 83.2% |
| 19 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 82.7% |
| 20 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 79.3% |
| 21 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 77.4% |
| 22 | [Inkling](/models/inkling) | Thinking Machines Lab | 77.1% |
| 23 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 75.8% |
| 24 | [Agents-A1](/models/agents-a1) | InternScience | 75.5% |
| 25 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 73.2% |
| 26 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 72.2% |
| 27 | [GLM-5.1](/models/glm-5-1) | Z.AI | 68% |
| 28 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 67.6% |
| 29 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 66.8% |
| 30 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 65.8% |
| 31 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 63.8% |
| 32 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 62% |
| 33 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 61% |
| 34 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 61% |
| 35 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 60.6% |
| 36 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 60.6% |
| 37 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 56.4% |
| 38 | [GLM-4.7](/models/glm-4-7) | Z.AI | 52% |
| 39 | [Solar Pro 4](/models/solar-pro-4) | Upstage | 49.2% |
| 40 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 48.6% |
| 41 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 44.4% |
| 42 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 36.8% |

## FAQ

### What does BrowseComp measure?

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

### Which model scores highest on BrowseComp?

GPT-5.6 Sol by OpenAI currently leads with a score of 92.2% on BrowseComp.

### How many models are evaluated on BrowseComp?

42 AI models have been evaluated on BrowseComp on BenchLM.

### Does BrowseComp affect BenchLM's overall score?

Yes. BrowseComp is a weighted benchmark inside the Agentic category, which carries 22% of BenchLM's overall score. BrowseComp itself contributes 25% of that category score.

## Compare Top Models on BrowseComp

- [GPT-5.6 Sol vs GPT-6 Astra](/compare/gpt-5-6-sol-vs-gpt-6-astra)
- [GPT-6 Astra vs Kimi K3](/compare/gpt-6-astra-vs-kimi-k3)
- [Kimi K3 vs Claude Opus 5](/compare/claude-opus-5-vs-kimi-k3)
- [Claude Opus 5 vs GPT-5.5 Pro](/compare/claude-opus-5-vs-gpt-5-5-pro)

## Related Reading

- [BrowseComp benchmark explainer](/blog/posts/browsecomp-browsing-benchmark)
