# APEX-Agents 1.1 — Mercor evaluation (APEX-Agents 1.1 (Mercor))

> Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Canonical page: https://benchlm.ai/benchmarks/mercorapexagents11

- Category: [Agentic](/agentic)
- Last updated: October 7, 2026 capture

Evaluation results published by [Mercor](https://www.mercor.com/apex/apex-agents-leaderboard/). Published evaluation aggregates only; task contents and dataset license grants are not included.

We captured Mercor's APEX-Agents 1.1 results on October 7, 2026. The table preserves 55 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.

This is the September 2026 APEX-Agents 1.1 revision: 240 tasks, updated task specifications, and a judge that penalizes multiple incompatible answers. Its release post and current selector rank by Pass@1; the general FAQ still calls mean score the primary metric. The general methodology lists three runs, while the revision discusses four. We preserve both metrics and reported sample counts without rescaling. Original APEX-Agents and Artificial Analysis APEX-Agents-AA remain separate.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists expert-authored rubrics over final deliverables and agent trajectories and a judge of LLM judge over expert rubrics. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.

- [Mercor APEX-Agents 1.1 results](https://www.mercor.com/apex/apex-agents-leaderboard/)
- [Mercor APEX benchmark index](https://www.mercor.com/apex/)
- [Mercor methodology](https://www.mercor.com/apex/methodology/)
- [Source paper](https://arxiv.org/abs/2601.14242)
- [Upstream evaluation code](https://github.com/Mercor-Intelligence/apex_loop_truncated_tools_agent)
- [Upstream dataset](https://huggingface.co/datasets/mercor/apex-agents-v1.1)
- [Mercor evaluation notes](https://www.mercor.com/blog/introducing-apex-agents-1-1)

## About APEX-Agents 1.1 (Mercor)

- Tasks: 240 tasks in the version 1.1 revision
- Format: Pass@1
- Difficulty: Source-specific evaluation
- Paper: [APEX-Agents 1.1 evaluation results](https://www.mercor.com/apex/apex-agents-leaderboard/)

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

APEX-Agents 1.1 (Mercor) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (55 configurations)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Gemini 4 Argon](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.3 · Mean score 87.4% ±3.3 · 954 reported samples | Google | 82.40% |
| 2 | [Sonnet 5.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±4.7 · Mean score 83.7% ±3.6 · 955 reported samples | Anthropic | 75.50% |
| 3 | [Opus 5.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±4.9 · Mean score 81.3% ±4 · 950 reported samples | Anthropic | 73.50% |
| 4 | [Fable 5.1](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±4.9 · Mean score 77.1% ±4.3 · 959 reported samples | Anthropic | 68.60% |
| 5 | [Gemini 3.7 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.1 · Mean score 79.4% ±3.8 · 952 reported samples | Google | 67.80% |
| 6 | [Opus 5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.1 · Mean score 77% ±4 · 956 reported samples | Anthropic | 65.80% |
| 7 | [Grok 4.6](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.2 · Mean score 77.3% ±4.1 · 960 reported samples | xAI | 65.30% |
| 8 | [GPT 6 Astra](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.6 · Mean score 75.1% ±4.6 · 959 reported samples | OpenAI | 64.70% |
| 9 | [Gemini 3.8 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.2 · Mean score 72.9% ±4.4 · 956 reported samples | Google | 64.30% |
| 10 | [Fable 5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.3 · Mean score 74.5% ±4.3 · 959 reported samples | Anthropic | 63.60% |
| 11 | [Qwen 3.8 Max](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±4.9 · Mean score 74.3% ±4 · 951 reported samples | Alibaba | 63.30% |
| 12 | [GPT 6.1 Sol](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.6 · Mean score 73.2% ±4.6 · 954 reported samples | OpenAI | 60.00% |
| 13 | [Fable 5.1](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.6 · Mean score 71.8% ±4.6 · 954 reported samples | Anthropic | 59.70% |
| 14 | [MiMo V2.6 Pro](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · auto reasoning · Published margin ±5.2 · Mean score 71.1% ±4.4 · 864 reported samples | Xiaomi | 59.50% |
| 15 | [Muse Spark 1.3](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.5 · Mean score 70.4% ±4.4 · 959 reported samples | Meta | 58.60% |
| 16 | [GPT 5.6 Terra](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.3 · Mean score 71.8% ±4.2 · 958 reported samples | OpenAI | 58.20% |
| 17 | [MiMo V2.6 Flash RL](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · auto reasoning · Published margin ±5.1 · Mean score 70.4% ±4 · 940 reported samples | Xiaomi | 57.40% |
| 18 | [GLM 5.3](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.1 · Mean score 66.2% ±4.3 · 960 reported samples | Zhipu | 56.60% |
| 19 | [Grok 4.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.4 · Mean score 70.5% ±4.3 · 960 reported samples | xAI | 56.20% |
| 20 | [DeepSeek V4 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.6 · Mean score 67.5% ±4.6 · 829 reported samples | DeepSeek | 55.30% |
| 21 | [GPT 5.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.5 · Mean score 70.3% ±4.4 · 959 reported samples | OpenAI | 55.10% |
| 22 | [Grok 4.7](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.1 · Mean score 68.7% ±4.2 · 960 reported samples | xAI | 54.60% |
| 23 | [Sonnet 5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.5 · Mean score 67.4% ±4.5 · 958 reported samples | Anthropic | 54.50% |
| 24 | [GPT 6 Sol](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.6 · Mean score 68.7% ±4.7 · 951 reported samples | OpenAI | 54.30% |
| 25 | [GLM 5.3 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.3 · Mean score 66.8% ±4.4 · 957 reported samples | Zhipu | 52.80% |
| 26 | [Opus 5.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · medium reasoning · Published margin ±5.6 · Mean score 65.7% ±4.9 · 943 reported samples | Anthropic | 52.50% |
| 27 | [GPT 5.4](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.3 · Mean score 68.2% ±4.4 · 958 reported samples | OpenAI | 52.40% |
| 28 | [GPT 5.6 Sol (Pro)](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.5 · Mean score 65.2% ±4.8 · 958 reported samples | OpenAI | 51.50% |
| 29 | [Kimi K3](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.1 · Mean score 64% ±4.3 · 960 reported samples | Kimi | 50.60% |
| 30 | [Opus 4.7](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.4 · Mean score 63.1% ±4.8 · 960 reported samples | Anthropic | 49.20% |
| 31 | [Opus 4.8](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.5 · Mean score 64.4% ±4.6 · 961 reported samples | Anthropic | 48.90% |
| 32 | [Muse Spark 1.3](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.7 · Mean score 61.3% ±5 · 960 reported samples | Meta | 47.60% |
| 33 | [Qwen 3.8 27B](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±4.9 · Mean score 62.2% ±4.3 · 957 reported samples | Alibaba | 47.50% |
| 34 | [DeepSeek V4 Pro 0813](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.6 · Mean score 62.3% ±4.9 · 849 reported samples | DeepSeek | 47.30% |
| 35 | [Gemini 3.6 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.2 · Mean score 63.4% ±4.4 · 956 reported samples | Google | 46.90% |
| 36 | [Opus 4.6](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.5 · Mean score 61.8% ±4.8 · 955 reported samples | Anthropic | 46.30% |
| 37 | [GLM 5.2](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.7 · Mean score 61.1% ±4.9 · 846 reported samples | Zhipu | 45.20% |
| 38 | [Sonnet 5.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · medium reasoning · Published margin ±5.6 · Mean score 59.4% ±4.9 · 955 reported samples | Anthropic | 44.60% |
| 39 | [GPT 6 Luna](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.5 · Mean score 59% ±4.9 · 946 reported samples | OpenAI | 44.30% |
| 40 | [Sonnet 4.6](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5.5 · Mean score 57.6% ±4.9 · 949 reported samples | Anthropic | 43.00% |
| 41 | [GPT 5.6 Luna](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±5.4 · Mean score 57.8% ±4.9 · 945 reported samples | OpenAI | 43.00% |
| 42 | [GLM 5.1](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · Published margin ±5.4 · Mean score 56.8% ±4.6 · 856 reported samples | Zhipu | 40.90% |
| 43 | [DeepSeek V4.1 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · max reasoning · Published margin ±4.9 · Mean score 50.4% ±4.7 · 960 reported samples | DeepSeek | 39.50% |
| 44 | [MiniMax M3](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.9 · Mean score 53.5% ±4.5 · 951 reported samples | MiniMax | 37.70% |
| 45 | [Kimi K2.7 Code](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.8 · Mean score 53.4% ±4.4 · 950 reported samples | Kimi | 37.60% |
| 46 | [Muse Spark 1.2](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.2 · Mean score 51.7% ±4.9 · 960 reported samples | Meta | 36.40% |
| 47 | [Gemini 3.1 Pro](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5 · Mean score 52.5% ±4.6 · 947 reported samples | Google | 35.30% |
| 48 | [Inkling](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.9 · Mean score 48.8% ±4.8 · 946 reported samples | Thinking Machines | 33.80% |
| 49 | [Muse Spark 1.1](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · xhigh reasoning · Published margin ±5.1 · Mean score 46.2% ±5 · 819 reported samples | Meta | 31.80% |
| 50 | [Gemini 3.5 Flash Lite](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.9 · Mean score 46.1% ±4.7 · 954 reported samples | Google | 29.30% |
| 51 | [Gemini 3.5 Flash](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±5 · Mean score 44.3% ±4.9 · 959 reported samples | Google | 27.50% |
| 52 | [Qwen 3.5](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · Published margin ±4.7 · Mean score 40.3% ±4.6 · 960 reported samples | Alibaba | 24.90% |
| 53 | [Nemotron 3 Ultra](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±4.5 · Mean score 38.3% ±4.6 · 887 reported samples | NVIDIA | 22.70% |
| 54 | [DeepSeek V3.2](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · Published margin ±4.1 · Mean score 35.1% ±4.5 · 952 reported samples | DeepSeek | 21.30% |
| 55 | [GPT OSS 120B](https://www.mercor.com/apex/apex-agents-leaderboard/) | Loop (truncated tools) · high reasoning · Published margin ±2 · Mean score 8.3% ±2.7 · 954 reported samples | OpenAI | 4.40% |

## FAQ

### What does APEX-Agents 1.1 (Mercor) measure?

Professional-services tasks across investment banking, management consulting, and corporate law. This table shows Mercor-run configurations for reference and is excluded from model rankings.

### Which model leads the published APEX-Agents 1.1 (Mercor) snapshot?

Gemini 4 Argon currently leads the published APEX-Agents 1.1 (Mercor) snapshot with a score of 82.40%.

### How many models are evaluated on APEX-Agents 1.1 (Mercor)?

The October 7, 2026 capture contains 55 source configurations.

### Does APEX-Agents 1.1 (Mercor) affect BenchLM's overall score?

Not directly. APEX-Agents 1.1 (Mercor) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
