Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Benchmark profile

APEX-SWE

Mercor's software-engineering benchmark for repository work under a coding-agent harness.

Data verified 26 confirmed releases in the last 30 daysStart free brief

Benchmark score on APEX-SWE — August 12, 2026

BenchLM mirrors the published score view for APEX-SWE. Claude Fable 5 leads the public snapshot at 58.8% , followed by Grok 4.6 (56.4%) and Grok 4.5 (53.6%). We do not use these results to rank models overall.

8 modelsCodingCurrentDisplay onlyUpdated August 12, 2026

Benchmark score table (8 models)

Score
1
Claude Fable 5Anthropic · Closed
58.8%
2
Grok 4.6xAI · Closed
56.4%
3
Grok 4.5xAI · Closed
53.6%
4
Kimi K3Moonshot AI · Closed
48.0%
5
Claude Sonnet 5Anthropic · Closed
46.4%
6
GPT-5.6 SolOpenAI · Closed
45.8%
7
Claude Opus 4.8Anthropic · Closed
43.9%
8
GPT-5.5OpenAI · Closed
37.0%

The published APEX-SWE snapshot places Claude Fable 5 first at 58.8%. The third row is 5.2 points behind. The broader top-10 range is 21.8 points, so the table still separates the published systems.

8 models have been evaluated on APEX-SWE. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. APEX-SWE is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About APEX-SWE

Year

2026

Tasks

Repository-level software-engineering tasks

Format

Pass@1

Difficulty

Professional software engineering

BenchLM preserves exact public and model-card Pass@1 rows as display-only evidence because effort level and agent scaffold affect the result.

BenchLM freshness & provenance

Version

APEX-SWE 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does APEX-SWE measure?

Mercor's software-engineering benchmark for repository work under a coding-agent harness.

Which model scores highest on APEX-SWE?

Claude Fable 5 by Anthropic currently leads with a score of 58.8% on APEX-SWE.

How many models are evaluated on APEX-SWE?

8 AI models have been evaluated on APEX-SWE on BenchLM.

Last updated: August 12, 2026 · BenchLM version APEX-SWE 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.