APEX-Agents 1.1 — Mercor evaluation (APEX-Agents 1.1 (Mercor))
We show this table for reference; we do not rank on it.
Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.
Evaluation results published by Mercor. Published evaluation aggregates only; task contents and dataset license grants are not included.
Pass@1 on APEX-Agents 1.1 (Mercor) — October 7, 2026 capture
We mirror the published pass@1 view for APEX-Agents 1.1 (Mercor). Gemini 4 Argon leads the public snapshot at 82.40%, followed by Sonnet 5.5 (75.50%) and Opus 5.5 (73.50%). We do not use these results to rank models overall.
Gemini 4 Argon
Loop (truncated tools) · high reasoning
Published margin ±4.3 · Mean score 87.4% ±3.3 · 954 reported samples
Sonnet 5.5
Anthropic
Loop (truncated tools) · max reasoning
Published margin ±4.7 · Mean score 83.7% ±3.6 · 955 reported samples
Opus 5.5
Anthropic
Loop (truncated tools) · max reasoning
Published margin ±4.9 · Mean score 81.3% ±4 · 950 reported samples
55 configurationsAgenticDated Mercor configurationsDisplay onlyUpdated October 7, 2026 capture
Pass@1 table (55 configurations)
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Gemini 4 ArgonGoogleLoop (truncated tools) · high reasoningPublished margin ±4.3 · Mean score 87.4% ±3.3 · 954 reported samples | 82.40% | Not reported | Unknown |
| 2 | Sonnet 5.5AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.7 · Mean score 83.7% ±3.6 · 955 reported samples | 75.50% | Not reported | Unknown |
| 3 | Opus 5.5AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 81.3% ±4 · 950 reported samples | 73.50% | Not reported | Unknown |
| 4 | Fable 5.1AnthropicLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 77.1% ±4.3 · 959 reported samples | 68.60% | Not reported | Unknown |
| 5 | Gemini 3.7 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.1 · Mean score 79.4% ±3.8 · 952 reported samples | 67.80% | Not reported | Unknown |
| 6 | Opus 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 77% ±4 · 956 reported samples | 65.80% | Not reported | Unknown |
| 7 | Grok 4.6xAILoop (truncated tools) · xhigh reasoningPublished margin ±5.2 · Mean score 77.3% ±4.1 · 960 reported samples | 65.30% | Not reported | Unknown |
| 8 | GPT 6 AstraOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 75.1% ±4.6 · 959 reported samples | 64.70% | Not reported | Unknown |
| 9 | Gemini 3.8 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.2 · Mean score 72.9% ±4.4 · 956 reported samples | 64.30% | Not reported | Unknown |
| 10 | Fable 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 74.5% ±4.3 · 959 reported samples | 63.60% | Not reported | Unknown |
| 11 | Qwen 3.8 MaxAlibabaLoop (truncated tools) · xhigh reasoningPublished margin ±4.9 · Mean score 74.3% ±4 · 951 reported samples | 63.30% | Not reported | Unknown |
| 12 | GPT 6.1 SolOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 73.2% ±4.6 · 954 reported samples | 60.00% | Not reported | Unknown |
| 13 | Fable 5.1AnthropicLoop (truncated tools) · high reasoningPublished margin ±5.6 · Mean score 71.8% ±4.6 · 954 reported samples | 59.70% | Not reported | Unknown |
| 14 | MiMo V2.6 ProXiaomiLoop (truncated tools) · auto reasoningPublished margin ±5.2 · Mean score 71.1% ±4.4 · 864 reported samples | 59.50% | Not reported | Unknown |
| 15 | Muse Spark 1.3MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.5 · Mean score 70.4% ±4.4 · 959 reported samples | 58.60% | Not reported | Unknown |
| 16 | GPT 5.6 TerraOpenAILoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 71.8% ±4.2 · 958 reported samples | 58.20% | Not reported | Unknown |
| 17 | MiMo V2.6 Flash RLXiaomiLoop (truncated tools) · auto reasoningPublished margin ±5.1 · Mean score 70.4% ±4 · 940 reported samples | 57.40% | Not reported | Unknown |
| 18 | GLM 5.3ZhipuLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 66.2% ±4.3 · 960 reported samples | 56.60% | Not reported | Unknown |
| 19 | Grok 4.5xAILoop (truncated tools) · high reasoningPublished margin ±5.4 · Mean score 70.5% ±4.3 · 960 reported samples | 56.20% | Not reported | Unknown |
| 20 | DeepSeek V4 FlashDeepSeekLoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 67.5% ±4.6 · 829 reported samples | 55.30% | Not reported | Unknown |
| 21 | GPT 5.5OpenAILoop (truncated tools) · xhigh reasoningPublished margin ±5.5 · Mean score 70.3% ±4.4 · 959 reported samples | 55.10% | Not reported | Unknown |
| 22 | Grok 4.7xAILoop (truncated tools) · xhigh reasoningPublished margin ±5.1 · Mean score 68.7% ±4.2 · 960 reported samples | 54.60% | Not reported | Unknown |
| 23 | Sonnet 5AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 67.4% ±4.5 · 958 reported samples | 54.50% | Not reported | Unknown |
| 24 | GPT 6 SolOpenAILoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 68.7% ±4.7 · 951 reported samples | 54.30% | Not reported | Unknown |
| 25 | GLM 5.3 FlashZhipuLoop (truncated tools) · max reasoningPublished margin ±5.3 · Mean score 66.8% ±4.4 · 957 reported samples | 52.80% | Not reported | Unknown |
| 26 | Opus 5.5AnthropicLoop (truncated tools) · medium reasoningPublished margin ±5.6 · Mean score 65.7% ±4.9 · 943 reported samples | 52.50% | Not reported | Unknown |
| 27 | GPT 5.4OpenAILoop (truncated tools) · xhigh reasoningPublished margin ±5.3 · Mean score 68.2% ±4.4 · 958 reported samples | 52.40% | Not reported | Unknown |
| 28 | GPT 5.6 Sol (Pro)OpenAILoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 65.2% ±4.8 · 958 reported samples | 51.50% | Not reported | Unknown |
| 29 | Kimi K3KimiLoop (truncated tools) · max reasoningPublished margin ±5.1 · Mean score 64% ±4.3 · 960 reported samples | 50.60% | Not reported | Unknown |
| 30 | Opus 4.7AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.4 · Mean score 63.1% ±4.8 · 960 reported samples | 49.20% | Not reported | Unknown |
| 31 | Opus 4.8AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 64.4% ±4.6 · 961 reported samples | 48.90% | Not reported | Unknown |
| 32 | Muse Spark 1.3MetaLoop (truncated tools) · max reasoningPublished margin ±5.7 · Mean score 61.3% ±5 · 960 reported samples | 47.60% | Not reported | Unknown |
| 33 | Qwen 3.8 27BAlibabaLoop (truncated tools) · xhigh reasoningPublished margin ±4.9 · Mean score 62.2% ±4.3 · 957 reported samples | 47.50% | Not reported | Unknown |
| 34 | DeepSeek V4 Pro 0813DeepSeekLoop (truncated tools) · max reasoningPublished margin ±5.6 · Mean score 62.3% ±4.9 · 849 reported samples | 47.30% | Not reported | Unknown |
| 35 | Gemini 3.6 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5.2 · Mean score 63.4% ±4.4 · 956 reported samples | 46.90% | Not reported | Unknown |
| 36 | Opus 4.6AnthropicLoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 61.8% ±4.8 · 955 reported samples | 46.30% | Not reported | Unknown |
| 37 | GLM 5.2ZhipuLoop (truncated tools) · max reasoningPublished margin ±5.7 · Mean score 61.1% ±4.9 · 846 reported samples | 45.20% | Not reported | Unknown |
| 38 | Sonnet 5.5AnthropicLoop (truncated tools) · medium reasoningPublished margin ±5.6 · Mean score 59.4% ±4.9 · 955 reported samples | 44.60% | Not reported | Unknown |
| 39 | GPT 6 LunaOpenAILoop (truncated tools) · max reasoningPublished margin ±5.5 · Mean score 59% ±4.9 · 946 reported samples | 44.30% | Not reported | Unknown |
| 40 | Sonnet 4.6AnthropicLoop (truncated tools) · high reasoningPublished margin ±5.5 · Mean score 57.6% ±4.9 · 949 reported samples | 43.00% | Not reported | Unknown |
| 41 | GPT 5.6 LunaOpenAILoop (truncated tools) · max reasoningPublished margin ±5.4 · Mean score 57.8% ±4.9 · 945 reported samples | 43.00% | Not reported | Unknown |
| 42 | GLM 5.1ZhipuLoop (truncated tools)Published margin ±5.4 · Mean score 56.8% ±4.6 · 856 reported samples | 40.90% | Not reported | Unknown |
| 43 | DeepSeek V4.1 FlashDeepSeekLoop (truncated tools) · max reasoningPublished margin ±4.9 · Mean score 50.4% ±4.7 · 960 reported samples | 39.50% | Not reported | Unknown |
| 44 | MiniMax M3MiniMaxLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 53.5% ±4.5 · 951 reported samples | 37.70% | Not reported | Unknown |
| 45 | Kimi K2.7 CodeKimiLoop (truncated tools) · high reasoningPublished margin ±4.8 · Mean score 53.4% ±4.4 · 950 reported samples | 37.60% | Not reported | Unknown |
| 46 | Muse Spark 1.2MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.2 · Mean score 51.7% ±4.9 · 960 reported samples | 36.40% | Not reported | Unknown |
| 47 | Gemini 3.1 ProGoogleLoop (truncated tools) · high reasoningPublished margin ±5 · Mean score 52.5% ±4.6 · 947 reported samples | 35.30% | Not reported | Unknown |
| 48 | InklingThinking MachinesLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 48.8% ±4.8 · 946 reported samples | 33.80% | Not reported | Unknown |
| 49 | Muse Spark 1.1MetaLoop (truncated tools) · xhigh reasoningPublished margin ±5.1 · Mean score 46.2% ±5 · 819 reported samples | 31.80% | Not reported | Unknown |
| 50 | Gemini 3.5 Flash LiteGoogleLoop (truncated tools) · high reasoningPublished margin ±4.9 · Mean score 46.1% ±4.7 · 954 reported samples | 29.30% | Not reported | Unknown |
| 51 | Gemini 3.5 FlashGoogleLoop (truncated tools) · high reasoningPublished margin ±5 · Mean score 44.3% ±4.9 · 959 reported samples | 27.50% | Not reported | Unknown |
| 52 | Qwen 3.5AlibabaLoop (truncated tools)Published margin ±4.7 · Mean score 40.3% ±4.6 · 960 reported samples | 24.90% | Not reported | Unknown |
| 53 | Nemotron 3 UltraNVIDIALoop (truncated tools) · high reasoningPublished margin ±4.5 · Mean score 38.3% ±4.6 · 887 reported samples | 22.70% | Not reported | Unknown |
| 54 | DeepSeek V3.2DeepSeekLoop (truncated tools)Published margin ±4.1 · Mean score 35.1% ±4.5 · 952 reported samples | 21.30% | Not reported | Unknown |
| 55 | GPT OSS 120BOpenAILoop (truncated tools) · high reasoningPublished margin ±2 · Mean score 8.3% ±2.7 · 954 reported samples | 4.40% | Not reported | Unknown |
How to read this Mercor evaluation
We captured Mercor's APEX-Agents 1.1 results on October 7, 2026. The table preserves 55 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.
This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.
This is the September 2026 APEX-Agents 1.1 revision: 240 tasks, updated task specifications, and a judge that penalizes multiple incompatible answers. Its release post and current selector rank by Pass@1; the general FAQ still calls mean score the primary metric. The general methodology lists three runs, while the revision discusses four. We preserve both metrics and reported sample counts without rescaling. Original APEX-Agents and Artificial Analysis APEX-Agents-AA remain separate.
The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.
The general methodology lists expert-authored rubrics over final deliverables and agent trajectories and a judge of LLM judge over expert rubrics. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.
Snapshot
The published APEX-Agents 1.1 (Mercor) snapshot places Gemini 4 Argon first at 82.40%. The third row is 8.90 points behind. The broader top-10 range is 18.80 points, so the table still separates the published systems.
55 configurations are shown for APEX-Agents 1.1 (Mercor). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. APEX-Agents 1.1 (Mercor) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About APEX-Agents 1.1 (Mercor)
Tasks
240 tasks in the version 1.1 revision
Format
Pass@1
Difficulty
Source-specific evaluation
The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.
Freshness and provenance
Version
APEX-Agents 1.1; Mercor Pass@1; loop_truncated_tools_agent
Refresh cadence
Reviewed dated source capture
Staleness state
Dated Mercor configurations
Question availability
Upstream task access and terms; only aggregate results mirrored
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does APEX-Agents 1.1 (Mercor) measure?
Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.
Which model leads the published APEX-Agents 1.1 (Mercor) snapshot?
Gemini 4 Argon currently leads the published APEX-Agents 1.1 (Mercor) snapshot with 82.40% pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on APEX-Agents 1.1 (Mercor)?
The October 7, 2026 capture snapshot contains 55 source configurations.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.