# CursorBench

> Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

Canonical page: https://benchlm.ai/benchmarks/cursorbench

- Category: [Coding](/coding)
- Last updated: September 18, 2026

## About CursorBench

- Year: 2026
- Tasks: Long-horizon multi-file agentic coding tasks
- Format: Cursor agent-loop evaluation
- Difficulty: Professional agentic software engineering
- Paper: [CursorBench 4.0](https://cursor.com/cursorbench)

Cursor published the CursorBench 4.0 task set on September 10, 2026, adding long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems; 4.0 scores are not comparable with the retired 3.2 set. BenchLM tracks 4.0 as display-only because it is a first-party benchmark and maps the highest published effort tier for each model. The frozen CursorBench 3.2 rows stay on model pages as the weighted coding-lane instrument.

CursorBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 51.8% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 46.6% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 41.7% |
| 4 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 41.6% |
| 5 | [Grok 4.6](/models/grok-4-6) | xAI | 41.4% |
| 6 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 41.3% |
| 7 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 39.6% |
| 8 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 35.9% |
| 9 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 34.1% |
| 10 | [Composer 2.5](/models/composer-2-5) | Cursor | 27.7% |

## FAQ

### What does CursorBench measure?

Cursor's first-party benchmark for ambiguous, multi-file coding-agent tasks drawn from real Cursor sessions, currently on the 4.0 task set.

### Which model scores highest on CursorBench?

Claude Fable 5.1 by Anthropic currently leads with a score of 51.8% on CursorBench.

### How many models are evaluated on CursorBench?

10 AI models have been evaluated on CursorBench on BenchLM.

### Does CursorBench affect BenchLM's overall score?

Not directly. CursorBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on CursorBench

- [Claude Fable 5.1 vs Claude Opus 5](/compare/claude-fable-5-1-vs-claude-opus-5)
- [Claude Opus 5 vs GPT-5.6 Sol](/compare/claude-opus-5-vs-gpt-5-6-sol)
- [GPT-5.6 Sol vs Muse Spark 1.3](/compare/gpt-5-6-sol-vs-muse-spark-1-3)
- [Muse Spark 1.3 vs Grok 4.6](/compare/grok-4-6-vs-muse-spark-1-3)
