# BrowseComp Extended — Mercor evaluation (BrowseComp Extended (Mercor))

> Finding difficult information through web research. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Canonical page: https://benchlm.ai/benchmarks/mercorbrowsecompextended

- Category: [Agentic](/agentic)
- Last updated: October 1, 2026 capture

Evaluation results published by [Mercor](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/). Published evaluation aggregates only; task contents and dataset license grants are not included.

We captured Mercor's BrowseComp Extended results on October 1, 2026. The table preserves 33 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

The headline uses the 100 Mercor-built tasks. The paired original-set values refer to the separately listed 130-task public benchmark. Missing baselines remain unreported. We keep the extension and original set separate from their upstream benchmark-owner and provider-run tables.

The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists llm-as-a-judge (lmaaj) and a judge of Gemini 3.5 Flash. Its run-count label is k = 4. Those notes describe the original public benchmark; extension-specific changes are not established by that reference.

- [Mercor BrowseComp Extended results](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/)
- [Mercor APEX benchmark index](https://www.mercor.com/apex/)
- [Mercor methodology](https://www.mercor.com/apex/methodology/)
- [Paired Mercor original-set results](https://www.mercor.com/apex/oss-benchmarks/oss-browsecomp-leaderboard/)
- [Source paper](https://arxiv.org/abs/2504.12516)
- [Upstream evaluation code](https://github.com/openai/simple-evals)
- [Upstream dataset](https://github.com/openai/simple-evals)

## About BrowseComp Extended (Mercor)

- Tasks: 100 Mercor tasks; 130 tasks in the separately reported original set
- Format: Pass@1
- Difficulty: Source-specific evaluation
- Paper: [BrowseComp Extended evaluation results](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/)

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

BrowseComp Extended (Mercor) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (33 configurations)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Sonnet 5.5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8.1 · Original set 86.6% ±5 · 338 reported samples | Anthropic | 73.00% |
| 2 | [Fable 5.1](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8.5 · Original set 85.2% ±5.3 · 377 reported samples | Anthropic | 63.30% |
| 3 | [GPT-6 Sol](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8.3 · Original set 89.4% ±4.5 · 400 reported samples | OpenAI | 63.20% |
| 4 | [Opus 5.5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8.4 · Original set 88.5% ±5 · 400 reported samples | Anthropic | 60.50% |
| 5 | [DeepSeek-V4.1-Flash](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±9.3 · Original set 85.8% ±9 · 186 reported samples | DeepSeek | 48.20% |
| 6 | [GPT-6 Luna](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8.9 · Original set 83.3% ±5.4 · 400 reported samples | OpenAI | 47.20% |
| 7 | [Fable 5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±8 · Original set 82.5% ±5.6 · 400 reported samples | Anthropic | 44.50% |
| 8 | [Opus 5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±8 · Original set 84.6% ±5.1 · 400 reported samples | Anthropic | 44.00% |
| 9 | [GPT-5.6 Sol](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · xhigh reasoning · Published margin ±7.6 · Original set 90.6% ±4.4 · 400 reported samples | OpenAI | 32.30% |
| 10 | [Opus 4.8](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6.9 · Original set 73.7% ±6.2 · 400 reported samples | Anthropic | 30.00% |
| 11 | [Kimi K3](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6.9 · Original set 89% ±4.4 · 400 reported samples | Kimi | 28.50% |
| 12 | [DeepSeek-V4-Pro-0813](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6.9 · Original set 73.5% ±6 · 400 reported samples | DeepSeek | 27.00% |
| 13 | [Gemini 3.7 Flash](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±6.9 · Original set 79.2% ±6 · 400 reported samples | Google | 25.50% |
| 14 | [GPT-5.6 Terra](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6.8 · Original set 85.8% ±5.2 · 400 reported samples | OpenAI | 25.30% |
| 15 | [DeepSeek-V4-Flash](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6.5 · Original set 77.1% ±6.1 · 400 reported samples | DeepSeek | 22.00% |
| 16 | [Qwen3.8-Max](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · xhigh reasoning · Published margin ±6 · Original set 67.9% ±6.5 · 400 reported samples | Alibaba | 21.50% |
| 17 | [Grok 4.6](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · xhigh reasoning · Published margin ±6.6 · Original set 84% ±5.1 · 400 reported samples | xAI | 21.50% |
| 18 | [GPT-5.5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · xhigh reasoning · Published margin ±6 · Original set 77.7% ±5.8 · 400 reported samples | OpenAI | 19.00% |
| 19 | [Sonnet 5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±6 · Original set 68.7% ±6.4 · 400 reported samples | Anthropic | 19.00% |
| 20 | [GPT-5.6 Luna](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±5.9 · Original set 83.8% ±5.4 · 400 reported samples | OpenAI | 16.50% |
| 21 | [Grok 4.5](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±5.9 · Original set 75.8% ±6.1 · 400 reported samples | xAI | 14.20% |
| 22 | [Opus 4.7](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±5 · Original set 60.6% ±7 · 400 reported samples | Anthropic | 14.20% |
| 23 | [Sonnet 4.6](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±5.3 · Original set 56.7% ±7.2 · 400 reported samples | Anthropic | 14.20% |
| 24 | [GPT-5.4](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · xhigh reasoning · Published margin ±5.4 · Original set 73.3% ±6.2 · 400 reported samples | OpenAI | 14.00% |
| 25 | [Gemini 3.6 Flash](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±5 · Original set 54.6% ±6.1 · 400 reported samples | Google | 13.80% |
| 26 | [Gemini 3.1 Pro](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±4.5 · Original set 75.6% ±6.3 · 400 reported samples | Google | 10.70% |
| 27 | [DeepSeek-V4-Pro](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · max reasoning · Published margin ±4 · Original set 62.3% ±6.4 · 400 reported samples | DeepSeek | 8.30% |
| 28 | [Kimi K2.7 Code](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±4 · Original set 43.5% ±6.2 · 400 reported samples | Kimi | 8.00% |
| 29 | [MiniMax-M3](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±3.7 · Original set 53.3% ±6.7 · 400 reported samples | MiniMax | 7.70% |
| 30 | [Inkling](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±2.7 · Original set 50.4% ±6.9 · 400 reported samples | Thinking Machines | 4.50% |
| 31 | [DeepSeek-V3.2](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · Published margin ±2.9 · Original set 38.1% ±7 · 400 reported samples | DeepSeek | 3.30% |
| 32 | [Gemini 3.5 Flash](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±0.9 · Original set 43.3% ±6.7 · 400 reported samples | Google | 0.70% |
| 33 | [MiniMax-M2.7](https://www.mercor.com/apex/mercor-extended/ots-held-out-browsecomp-leaderboard/) | Mercor extended-set run · high reasoning · Published margin ±0.7 · Original set 38.3% ±7 · 400 reported samples | MiniMax | 0.50% |

## FAQ

### What does BrowseComp Extended (Mercor) measure?

Finding difficult information through web research. This table shows Mercor-run configurations for reference and is excluded from model rankings.

### Which model leads the published BrowseComp Extended (Mercor) snapshot?

Sonnet 5.5 currently leads the published BrowseComp Extended (Mercor) snapshot with a score of 73.00%.

### How many models are evaluated on BrowseComp Extended (Mercor)?

The October 1, 2026 capture contains 33 source configurations.

### Does BrowseComp Extended (Mercor) affect BenchLM's overall score?

Not directly. BrowseComp Extended (Mercor) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
