Skip to main content
BenchLM
Data

CursorBench

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Benchmark score on CursorBench — October 2, 2026

We compile the CursorBench rows from benchmark-owner or independent runs. Claude Opus 5.5 leads the table at 57.8%, followed by Claude Sonnet 5.5 (55.5%) and Claude Fable 5.1 (51.8%). We do not use these results to rank models overall.

13 modelsCodingCurrentDisplay onlyUpdated October 2, 2026

Benchmark score table (13 models)

Score
1
Claude Opus 5.5Anthropic · Closed
57.8%
2
Claude Sonnet 5.5Anthropic · Closed
55.5%
3
Claude Fable 5.1Anthropic · Closed
51.8%
4
Claude Opus 5Anthropic · Closed
46.6%
5
Grok 4.7xAI · Closed
46.3%
6
GPT-5.6 SolOpenAI · Closed
41.7%
7
Muse Spark 1.3Meta · Closed
41.6%
8
Grok 4.6xAI · Closed
41.4%
9
GPT-5.6 TerraOpenAI · Closed
41.3%
10
Gemini 3.8 FlashGoogle · Closed
39.6%
11
GPT-5.6 LunaOpenAI · Closed
35.9%
12
Claude Sonnet 5Anthropic · Closed
34.1%
13
Composer 2.5Cursor · Closed
27.7%

Among the reported CursorBench rows, Claude Opus 5.5 is first at 57.8%. The third row is 6.0 points behind. The broader top-10 range is 18.2 points, so the table still separates the published systems.

13 models have been evaluated on CursorBench. The benchmark falls in the Coding category. CursorBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CursorBench

Year

2026

Tasks

Long-horizon multi-file agentic coding tasks

Format

Cursor agent-loop evaluation

Difficulty

Professional agentic software engineering

Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.

Freshness and provenance

Version

CursorBench 4.0

Refresh cadence

Continuous

Staleness state

Current

Question availability

Tasks private; drawn from real Cursor sessions

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CursorBench measure?

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Which model scores highest on CursorBench?

Claude Opus 5.5 by Anthropic currently leads with a score of 57.8% on CursorBench.

How many models are evaluated on CursorBench?

13 AI models have published results on CursorBench in the BenchLM catalog.

Last updated: October 2, 2026 · BenchLM version CursorBench 4.0

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.