Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

CursorBench

We mirror this table; we do not rank on it.

Data verified 35 confirmed releases in the last 30 daysFollow model changes

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Benchmark score on CursorBench — September 18, 2026

We mirror the published score view for CursorBench. Claude Fable 5.1 leads the public snapshot at 51.8%, followed by Claude Opus 5 (46.6%) and GPT-5.6 Sol (41.7%). We do not use these results to rank models overall.

10 modelsCodingCurrentDisplay onlyUpdated September 18, 2026

Benchmark score table (10 models)

Score
1
Claude Fable 5.1Anthropic · Closed
51.8%
2
Claude Opus 5Anthropic · Closed
46.6%
3
GPT-5.6 SolOpenAI · Closed
41.7%
4
Muse Spark 1.3Meta · Closed
41.6%
5
Grok 4.6xAI · Closed
41.4%
6
GPT-5.6 TerraOpenAI · Closed
41.3%
7
Gemini 3.8 FlashGoogle · Closed
39.6%
8
GPT-5.6 LunaOpenAI · Closed
35.9%
9
Claude Sonnet 5Anthropic · Closed
34.1%
10
Composer 2.5Cursor · Closed
27.7%

The published CursorBench snapshot places Claude Fable 5.1 first at 51.8%. The third row is 10.1 points behind. The broader top-10 range is 24.1 points, so the table still separates the published systems.

10 models have been evaluated on CursorBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. CursorBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CursorBench

Year

2026

Tasks

Long-horizon multi-file agentic coding tasks

Format

Cursor agent-loop evaluation

Difficulty

Professional agentic software engineering

Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.

Freshness and provenance

Version

CursorBench 4.0

Refresh cadence

Continuous

Staleness state

Current

Question availability

Tasks private; drawn from real Cursor sessions

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CursorBench measure?

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Which model scores highest on CursorBench?

Claude Fable 5.1 by Anthropic currently leads with a score of 51.8% on CursorBench.

How many models are evaluated on CursorBench?

10 AI models have been evaluated on CursorBench on BenchLM.

Last updated: September 18, 2026 · BenchLM version CursorBench 4.0

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.