CursorBench
We mirror this table; we do not rank on it.
Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.
Benchmark score on CursorBench — September 18, 2026
We mirror the published score view for CursorBench. Claude Fable 5.1 leads the public snapshot at 51.8%, followed by Claude Opus 5 (46.6%) and GPT-5.6 Sol (41.7%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Opus 5
Anthropic
GPT-5.6 Sol
OpenAI
10 modelsCodingCurrentDisplay onlyUpdated September 18, 2026
Benchmark score table (10 models)
ScoreThe published CursorBench snapshot places Claude Fable 5.1 first at 51.8%. The third row is 10.1 points behind. The broader top-10 range is 24.1 points, so the table still separates the published systems.
10 models have been evaluated on CursorBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. CursorBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About CursorBench
Year
2026
Tasks
Long-horizon multi-file agentic coding tasks
Format
Cursor agent-loop evaluation
Difficulty
Professional agentic software engineering
Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.
Freshness and provenance
Version
CursorBench 4.0
Refresh cadence
Continuous
Staleness state
Current
Question availability
Tasks private; drawn from real Cursor sessions
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CursorBench measure?
Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.
Which model scores highest on CursorBench?
Claude Fable 5.1 by Anthropic currently leads with a score of 51.8% on CursorBench.
How many models are evaluated on CursorBench?
10 AI models have been evaluated on CursorBench on BenchLM.
Compare top models on CursorBench
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.