CursorBench
We show this table for reference; we do not rank on it.
Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.
Benchmark score on CursorBench — October 2, 2026
We compile the CursorBench rows from benchmark-owner or independent runs. Claude Opus 5.5 leads the table at 57.8%, followed by Claude Sonnet 5.5 (55.5%) and Claude Fable 5.1 (51.8%). We do not use these results to rank models overall.
Claude Opus 5.5
Anthropic
Claude Sonnet 5.5
Anthropic
Claude Fable 5.1
Anthropic
13 modelsCodingCurrentDisplay onlyUpdated October 2, 2026
Benchmark score table (13 models)
ScoreAmong the reported CursorBench rows, Claude Opus 5.5 is first at 57.8%. The third row is 6.0 points behind. The broader top-10 range is 18.2 points, so the table still separates the published systems.
13 models have been evaluated on CursorBench. The benchmark falls in the Coding category. CursorBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About CursorBench
Year
2026
Tasks
Long-horizon multi-file agentic coding tasks
Format
Cursor agent-loop evaluation
Difficulty
Professional agentic software engineering
Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.
Freshness and provenance
Version
CursorBench 4.0
Refresh cadence
Continuous
Staleness state
Current
Question availability
Tasks private; drawn from real Cursor sessions
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CursorBench measure?
Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.
Which model scores highest on CursorBench?
Claude Opus 5.5 by Anthropic currently leads with a score of 57.8% on CursorBench.
How many models are evaluated on CursorBench?
13 AI models have published results on CursorBench in the BenchLM catalog.
Compare top models on CursorBench
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.