Skip to main content
BenchLM

BrowseComp Extended — Mercor evaluation (BrowseComp Extended (Mercor))

We show this table for reference; we do not rank on it.

Finding difficult information through web research. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Evaluation results published by Mercor. Published evaluation aggregates only; task contents and dataset license grants are not included.

Pass@1 on BrowseComp Extended (Mercor) — October 1, 2026 capture

We mirror the published pass@1 view for BrowseComp Extended (Mercor). Sonnet 5.5 leads the public snapshot at 73.00%, followed by Fable 5.1 (63.30%) and GPT-6 Sol (63.20%). We do not use these results to rank models overall.

33 configurationsAgenticDated Mercor configurationsDisplay onlyUpdated October 1, 2026 capture

Pass@1 table (33 configurations)

Score
1
Sonnet 5.5AnthropicMercor extended-set run · max reasoningPublished margin ±8.1 · Original set 86.6% ±5 · 338 reported samples
73.00%
2
Fable 5.1AnthropicMercor extended-set run · max reasoningPublished margin ±8.5 · Original set 85.2% ±5.3 · 377 reported samples
63.30%
3
GPT-6 SolOpenAIMercor extended-set run · max reasoningPublished margin ±8.3 · Original set 89.4% ±4.5 · 400 reported samples
63.20%
4
Opus 5.5AnthropicMercor extended-set run · max reasoningPublished margin ±8.4 · Original set 88.5% ±5 · 400 reported samples
60.50%
5
DeepSeek-V4.1-FlashDeepSeekMercor extended-set run · max reasoningPublished margin ±9.3 · Original set 85.8% ±9 · 186 reported samples
48.20%
6
GPT-6 LunaOpenAIMercor extended-set run · max reasoningPublished margin ±8.9 · Original set 83.3% ±5.4 · 400 reported samples
47.20%
7
Fable 5AnthropicMercor extended-set run · high reasoningPublished margin ±8 · Original set 82.5% ±5.6 · 400 reported samples
44.50%
8
Opus 5AnthropicMercor extended-set run · max reasoningPublished margin ±8 · Original set 84.6% ±5.1 · 400 reported samples
44.00%
9
GPT-5.6 SolOpenAIMercor extended-set run · xhigh reasoningPublished margin ±7.6 · Original set 90.6% ±4.4 · 400 reported samples
32.30%
10
Opus 4.8AnthropicMercor extended-set run · max reasoningPublished margin ±6.9 · Original set 73.7% ±6.2 · 400 reported samples
30.00%
11
Kimi K3KimiMercor extended-set run · max reasoningPublished margin ±6.9 · Original set 89% ±4.4 · 400 reported samples
28.50%
12
DeepSeek-V4-Pro-0813DeepSeekMercor extended-set run · max reasoningPublished margin ±6.9 · Original set 73.5% ±6 · 400 reported samples
27.00%
13
Gemini 3.7 FlashGoogleMercor extended-set run · high reasoningPublished margin ±6.9 · Original set 79.2% ±6 · 400 reported samples
25.50%
14
GPT-5.6 TerraOpenAIMercor extended-set run · max reasoningPublished margin ±6.8 · Original set 85.8% ±5.2 · 400 reported samples
25.30%
15
DeepSeek-V4-FlashDeepSeekMercor extended-set run · max reasoningPublished margin ±6.5 · Original set 77.1% ±6.1 · 400 reported samples
22.00%
16
Qwen3.8-MaxAlibabaMercor extended-set run · xhigh reasoningPublished margin ±6 · Original set 67.9% ±6.5 · 400 reported samples
21.50%
17
Grok 4.6xAIMercor extended-set run · xhigh reasoningPublished margin ±6.6 · Original set 84% ±5.1 · 400 reported samples
21.50%
18
GPT-5.5OpenAIMercor extended-set run · xhigh reasoningPublished margin ±6 · Original set 77.7% ±5.8 · 400 reported samples
19.00%
19
Sonnet 5AnthropicMercor extended-set run · max reasoningPublished margin ±6 · Original set 68.7% ±6.4 · 400 reported samples
19.00%
20
GPT-5.6 LunaOpenAIMercor extended-set run · max reasoningPublished margin ±5.9 · Original set 83.8% ±5.4 · 400 reported samples
16.50%
21
Grok 4.5xAIMercor extended-set run · high reasoningPublished margin ±5.9 · Original set 75.8% ±6.1 · 400 reported samples
14.20%
22
Opus 4.7AnthropicMercor extended-set run · max reasoningPublished margin ±5 · Original set 60.6% ±7 · 400 reported samples
14.20%
23
Sonnet 4.6AnthropicMercor extended-set run · high reasoningPublished margin ±5.3 · Original set 56.7% ±7.2 · 400 reported samples
14.20%
24
GPT-5.4OpenAIMercor extended-set run · xhigh reasoningPublished margin ±5.4 · Original set 73.3% ±6.2 · 400 reported samples
14.00%
25
Gemini 3.6 FlashGoogleMercor extended-set run · high reasoningPublished margin ±5 · Original set 54.6% ±6.1 · 400 reported samples
13.80%
26
Gemini 3.1 ProGoogleMercor extended-set run · high reasoningPublished margin ±4.5 · Original set 75.6% ±6.3 · 400 reported samples
10.70%
27
DeepSeek-V4-ProDeepSeekMercor extended-set run · max reasoningPublished margin ±4 · Original set 62.3% ±6.4 · 400 reported samples
8.30%
28
Kimi K2.7 CodeKimiMercor extended-set run · high reasoningPublished margin ±4 · Original set 43.5% ±6.2 · 400 reported samples
8.00%
29
MiniMax-M3MiniMaxMercor extended-set run · high reasoningPublished margin ±3.7 · Original set 53.3% ±6.7 · 400 reported samples
7.70%
30
InklingThinking MachinesMercor extended-set run · high reasoningPublished margin ±2.7 · Original set 50.4% ±6.9 · 400 reported samples
4.50%
31
DeepSeek-V3.2DeepSeekMercor extended-set runPublished margin ±2.9 · Original set 38.1% ±7 · 400 reported samples
3.30%
32
Gemini 3.5 FlashGoogleMercor extended-set run · high reasoningPublished margin ±0.9 · Original set 43.3% ±6.7 · 400 reported samples
0.70%
33
MiniMax-M2.7MiniMaxMercor extended-set run · high reasoningPublished margin ±0.7 · Original set 38.3% ±7 · 400 reported samples
0.50%

How to read this Mercor evaluation

We captured Mercor's BrowseComp Extended results on October 1, 2026. The table preserves 33 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

The headline uses the 100 Mercor-built tasks. The paired original-set values refer to the separately listed 130-task public benchmark. Missing baselines remain unreported. We keep the extension and original set separate from their upstream benchmark-owner and provider-run tables.

The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists llm-as-a-judge (lmaaj) and a judge of Gemini 3.5 Flash. Its run-count label is k = 4. Those notes describe the original public benchmark; extension-specific changes are not established by that reference.

The published BrowseComp Extended (Mercor) snapshot places Sonnet 5.5 first at 73.00%. The third row is 9.80 points behind. The broader top-10 range is 43.00 points, so the table still separates the published systems.

33 configurations are shown for BrowseComp Extended (Mercor). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. BrowseComp Extended (Mercor) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About BrowseComp Extended (Mercor)

Tasks

100 Mercor tasks; 130 tasks in the separately reported original set

Format

Pass@1

Difficulty

Source-specific evaluation

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

Freshness and provenance

Version

BrowseComp Extended; Mercor Pass@1

Refresh cadence

Reviewed dated source capture

Staleness state

Dated Mercor configurations

Question availability

Upstream task access and terms; only aggregate results mirrored

Dated Mercor configurationsDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does BrowseComp Extended (Mercor) measure?

Finding difficult information through web research. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Which model leads the published BrowseComp Extended (Mercor) snapshot?

Sonnet 5.5 currently leads the published BrowseComp Extended (Mercor) snapshot with 73.00% pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on BrowseComp Extended (Mercor)?

The October 1, 2026 capture snapshot contains 33 source configurations.

Last updated: October 1, 2026 capture · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.