Skip to main content
BenchLM
Data

APEX-Agents 1.1 — Mercor evaluation (APEX-Agents 1.1 (Mercor))

We show this table for reference; we do not rank on it.

Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Evaluation results published by Mercor. Published evaluation aggregates only; task contents and dataset license grants are not included.

Pass@1 on APEX-Agents 1.1 (Mercor) — October 7, 2026 capture

We mirror the published pass@1 view for APEX-Agents 1.1 (Mercor). Gemini 4 Argon leads the public snapshot at 82.40%, followed by Sonnet 5.5 (75.50%) and Opus 5.5 (73.50%). We do not use these results to rank models overall.

55 configurationsAgenticDated Mercor configurationsDisplay onlyUpdated October 7, 2026 capture

Pass@1 table (55 configurations)

Pass@1 results for APEX-Agents 1.1 (Mercor)
RankModel / configurationScoreParameters (B)Open / closed
1Gemini 4 ArgonGoogleLoop (truncated tools) · high reasoningPublished margin ±4.3 · Mean score 87.4% ±3.3 · 954 reported samples
82.40%
Not reportedUnknown
2Sonnet 5.5AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.7 · Mean score 83.7% ±3.6 · 955 reported samples
75.50%
Not reportedUnknown
3Opus 5.5AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 81.3% ±4 · 950 reported samples
73.50%
Not reportedUnknown
4Fable 5.1AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 77.1% ±4.3 · 959 reported samples
68.60%
Not reportedUnknown
5Gemini 3.7 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.1 · Mean score 79.4% ±3.8 · 952 reported samples
67.80%
Not reportedUnknown
6Opus 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 77% ±4 · 956 reported samples
65.80%
Not reportedUnknown
7Grok 4.6xAILoop (truncated tools) · xhigh reasoningPublished margin ±5.2 · Mean score 77.3% ±4.1 · 960 reported samples
65.30%
Not reportedUnknown
8GPT 6 AstraOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 75.1% ±4.6 · 959 reported samples
64.70%
Not reportedUnknown
9Gemini 3.8 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.2 · Mean score 72.9% ±4.4 · 956 reported samples
64.30%
Not reportedUnknown
10Fable 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 74.5% ±4.3 · 959 reported samples
63.60%
Not reportedUnknown
11Qwen 3.8 MaxAlibabaLoop (truncated tools) · xhigh reasoningPublished margin ±4.9 · Mean score 74.3% ±4 · 951 reported samples
63.30%
Not reportedUnknown
12GPT 6.1 SolOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 73.2% ±4.6 · 954 reported samples
60.00%
Not reportedUnknown
13Fable 5.1AnthropicLoop (truncated tools) · high reasoningPublished margin ±5.6 · Mean score 71.8% ±4.6 · 954 reported samples
59.70%
Not reportedUnknown
14MiMo V2.6 ProXiaomiLoop (truncated tools) · auto reasoningPublished margin ±5.2 · Mean score 71.1% ±4.4 · 864 reported samples
59.50%
Not reportedUnknown
15Muse Spark 1.3MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.5 · Mean score 70.4% ±4.4 · 959 reported samples
58.60%
Not reportedUnknown
16GPT 5.6 TerraOpenAILoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 71.8% ±4.2 · 958 reported samples
58.20%
Not reportedUnknown
17MiMo V2.6 Flash RLXiaomiLoop (truncated tools) · auto reasoningPublished margin ±5.1 · Mean score 70.4% ±4 · 940 reported samples
57.40%
Not reportedUnknown
18GLM 5.3ZhipuLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 66.2% ±4.3 · 960 reported samples
56.60%
Not reportedUnknown
19Grok 4.5xAILoop (truncated tools) · high reasoningPublished margin ±5.4 · Mean score 70.5% ±4.3 · 960 reported samples
56.20%
Not reportedUnknown
20DeepSeek V4 FlashDeepSeekLoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 67.5% ±4.6 · 829 reported samples
55.30%
Not reportedUnknown
21GPT 5.5OpenAILoop (truncated tools) · xhigh reasoningPublished margin ±5.5 · Mean score 70.3% ±4.4 · 959 reported samples
55.10%
Not reportedUnknown
22Grok 4.7xAILoop (truncated tools) · xhigh reasoningPublished margin ±5.1 · Mean score 68.7% ±4.2 · 960 reported samples
54.60%
Not reportedUnknown
23Sonnet 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 67.4% ±4.5 · 958 reported samples
54.50%
Not reportedUnknown
24GPT 6 SolOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 68.7% ±4.7 · 951 reported samples
54.30%
Not reportedUnknown
25GLM 5.3 FlashZhipuLoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 66.8% ±4.4 · 957 reported samples
52.80%
Not reportedUnknown
26Opus 5.5AnthropicLoop (truncated tools) · medium reasoningPublished margin ±5.6 · Mean score 65.7% ±4.9 · 943 reported samples
52.50%
Not reportedUnknown
27GPT 5.4OpenAILoop (truncated tools) · xhigh reasoningPublished margin ±5.3 · Mean score 68.2% ±4.4 · 958 reported samples
52.40%
Not reportedUnknown
28GPT 5.6 Sol (Pro)OpenAILoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 65.2% ±4.8 · 958 reported samples
51.50%
Not reportedUnknown
29Kimi K3KimiLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 64% ±4.3 · 960 reported samples
50.60%
Not reportedUnknown
30Opus 4.7AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.4 · Mean score 63.1% ±4.8 · 960 reported samples
49.20%
Not reportedUnknown
31Opus 4.8AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 64.4% ±4.6 · 961 reported samples
48.90%
Not reportedUnknown
32Muse Spark 1.3MetaLoop (truncated tools) · max reasoningPublished margin ±5.7 · Mean score 61.3% ±5 · 960 reported samples
47.60%
Not reportedUnknown
33Qwen 3.8 27BAlibabaLoop (truncated tools) · xhigh reasoningPublished margin ±4.9 · Mean score 62.2% ±4.3 · 957 reported samples
47.50%
Not reportedUnknown
34DeepSeek V4 Pro 0813DeepSeekLoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 62.3% ±4.9 · 849 reported samples
47.30%
Not reportedUnknown
35Gemini 3.6 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.2 · Mean score 63.4% ±4.4 · 956 reported samples
46.90%
Not reportedUnknown
36Opus 4.6AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 61.8% ±4.8 · 955 reported samples
46.30%
Not reportedUnknown
37GLM 5.2ZhipuLoop (truncated tools) · max reasoningPublished margin ±5.7 · Mean score 61.1% ±4.9 · 846 reported samples
45.20%
Not reportedUnknown
38Sonnet 5.5AnthropicLoop (truncated tools) · medium reasoningPublished margin ±5.6 · Mean score 59.4% ±4.9 · 955 reported samples
44.60%
Not reportedUnknown
39GPT 6 LunaOpenAILoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 59% ±4.9 · 946 reported samples
44.30%
Not reportedUnknown
40Sonnet 4.6AnthropicLoop (truncated tools) · high reasoningPublished margin ±5.5 · Mean score 57.6% ±4.9 · 949 reported samples
43.00%
Not reportedUnknown
41GPT 5.6 LunaOpenAILoop (truncated tools) · max reasoningPublished margin ±5.4 · Mean score 57.8% ±4.9 · 945 reported samples
43.00%
Not reportedUnknown
42GLM 5.1ZhipuLoop (truncated tools)Published margin ±5.4 · Mean score 56.8% ±4.6 · 856 reported samples
40.90%
Not reportedUnknown
43DeepSeek V4.1 FlashDeepSeekLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 50.4% ±4.7 · 960 reported samples
39.50%
Not reportedUnknown
44MiniMax M3MiniMaxLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 53.5% ±4.5 · 951 reported samples
37.70%
Not reportedUnknown
45Kimi K2.7 CodeKimiLoop (truncated tools) · high reasoningPublished margin ±4.8 · Mean score 53.4% ±4.4 · 950 reported samples
37.60%
Not reportedUnknown
46Muse Spark 1.2MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.2 · Mean score 51.7% ±4.9 · 960 reported samples
36.40%
Not reportedUnknown
47Gemini 3.1 ProGoogleLoop (truncated tools) · high reasoningPublished margin ±5 · Mean score 52.5% ±4.6 · 947 reported samples
35.30%
Not reportedUnknown
48InklingThinking MachinesLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 48.8% ±4.8 · 946 reported samples
33.80%
Not reportedUnknown
49Muse Spark 1.1MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.1 · Mean score 46.2% ±5 · 819 reported samples
31.80%
Not reportedUnknown
50Gemini 3.5 Flash LiteGoogleLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 46.1% ±4.7 · 954 reported samples
29.30%
Not reportedUnknown
51Gemini 3.5 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5 · Mean score 44.3% ±4.9 · 959 reported samples
27.50%
Not reportedUnknown
52Qwen 3.5AlibabaLoop (truncated tools)Published margin ±4.7 · Mean score 40.3% ±4.6 · 960 reported samples
24.90%
Not reportedUnknown
53Nemotron 3 UltraNVIDIALoop (truncated tools) · high reasoningPublished margin ±4.5 · Mean score 38.3% ±4.6 · 887 reported samples
22.70%
Not reportedUnknown
54DeepSeek V3.2DeepSeekLoop (truncated tools)Published margin ±4.1 · Mean score 35.1% ±4.5 · 952 reported samples
21.30%
Not reportedUnknown
55GPT OSS 120BOpenAILoop (truncated tools) · high reasoningPublished margin ±2 · Mean score 8.3% ±2.7 · 954 reported samples
4.40%
Not reportedUnknown

How to read this Mercor evaluation

We captured Mercor's APEX-Agents 1.1 results on October 7, 2026. The table preserves 55 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.

This is the September 2026 APEX-Agents 1.1 revision: 240 tasks, updated task specifications, and a judge that penalizes multiple incompatible answers. Its release post and current selector rank by Pass@1; the general FAQ still calls mean score the primary metric. The general methodology lists three runs, while the revision discusses four. We preserve both metrics and reported sample counts without rescaling. Original APEX-Agents and Artificial Analysis APEX-Agents-AA remain separate.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists expert-authored rubrics over final deliverables and agent trajectories and a judge of LLM judge over expert rubrics. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.

The published APEX-Agents 1.1 (Mercor) snapshot places Gemini 4 Argon first at 82.40%. The third row is 8.90 points behind. The broader top-10 range is 18.80 points, so the table still separates the published systems.

55 configurations are shown for APEX-Agents 1.1 (Mercor). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. APEX-Agents 1.1 (Mercor) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About APEX-Agents 1.1 (Mercor)

Tasks

240 tasks in the version 1.1 revision

Format

Pass@1

Difficulty

Source-specific evaluation

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

Freshness and provenance

Version

APEX-Agents 1.1; Mercor Pass@1; loop_truncated_tools_agent

Refresh cadence

Reviewed dated source capture

Staleness state

Dated Mercor configurations

Question availability

Upstream task access and terms; only aggregate results mirrored

Dated Mercor configurationsDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does APEX-Agents 1.1 (Mercor) measure?

Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Which model leads the published APEX-Agents 1.1 (Mercor) snapshot?

Gemini 4 Argon currently leads the published APEX-Agents 1.1 (Mercor) snapshot with 82.40% pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on APEX-Agents 1.1 (Mercor)?

The October 7, 2026 capture snapshot contains 55 source configurations.

Last updated: October 7, 2026 capture · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.