Skip to main content
BenchLM
Data

CursorBench

We show this table for reference; we do not rank on it.

Data verified 40 confirmed releases in the last 30 daysFollow model changes

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Benchmark score on CursorBench — October 8, 2026

We compile the CursorBench rows from benchmark-owner or independent runs. Claude Opus 5.5 leads the table at 57.8%, followed by Claude Sonnet 5.5 (55.5%) and Claude Fable 5.1 (51.8%). We do not use these results to rank models overall.

13 modelsCodingCurrentDisplay onlyUpdated October 8, 2026

Benchmark score results for CursorBench
RankModel / configurationScoreParameters (B)Open / closed
1Claude Opus 5.5Anthropic
57.8%
Not reportedClosed
2Claude Sonnet 5.5Anthropic
55.5%
Not reportedClosed
3Claude Fable 5.1Anthropic
51.8%
Not reportedClosed
4Claude Opus 5Anthropic
46.6%
Not reportedClosed
5Grok 4.7xAI
46.3%
Not reportedClosed
6GPT-5.6 SolOpenAI
41.7%
Not reportedClosed
7Muse Spark 1.3Meta
41.6%
Not reportedClosed
8Grok 4.6xAI
41.4%
Not reportedClosed
9GPT-5.6 TerraOpenAI
41.3%
Not reportedClosed
10Gemini 3.8 FlashGoogle
39.6%
Not reportedClosed
11GPT-5.6 LunaOpenAI
35.9%
Not reportedClosed
12Claude Sonnet 5Anthropic
34.1%
Not reportedClosed
13Composer 2.5Cursor
27.7%
Not reportedClosed

Among the reported CursorBench rows, Claude Opus 5.5 is first at 57.8%. The third row is 6.0 points behind. The broader top-10 range is 18.2 points, so the table still separates the published systems.

13 models have been evaluated on CursorBench. The benchmark falls in the Coding category. CursorBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CursorBench

Year

2026

Tasks

Long-horizon multi-file agentic coding tasks

Format

Cursor agent-loop evaluation

Difficulty

Professional agentic software engineering

Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.

Freshness and provenance

Version

CursorBench 4.0

Refresh cadence

Continuous

Staleness state

Current

Question availability

Tasks private; drawn from real Cursor sessions

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CursorBench measure?

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Which model scores highest on CursorBench?

Claude Opus 5.5 by Anthropic currently leads with a score of 57.8% on CursorBench.

How many models are evaluated on CursorBench?

13 AI models have published results on CursorBench in the BenchLM catalog.

Last updated: October 8, 2026 · BenchLM version CursorBench 4.0

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.