# DeepSWE

> A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

Canonical page: https://benchlm.ai/benchmarks/deepswe

- Category: [external](/external)
- Last updated: September 1, 2026

## About DeepSWE

- Year: 2026
- Tasks: 113 software engineering tasks across 91 repositories and 5 languages
- Format: Pass@1 with confidence interval, cost, time, and token metadata
- Difficulty: Long-horizon software engineering
- Paper: [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog)

DeepSWE includes original tasks with isolated environments and program-based verifiers. BenchLM mirrors the public DeepSWE leaderboard JSON as display-only, using the best available mini-swe-agent configuration per model and preserving cost, time, token, and effort-level source metadata. Each row combines a model, agent harness, and reasoning-effort setting rather than a pure model-only benchmark score.

DeepSWE is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (28 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | mini-swe-agent · high reasoning | Google | 73.8% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | mini-swe-agent · max reasoning | Anthropic | 73.6% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | mini-swe-agent · max reasoning | OpenAI | 73.2% |
| 4 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | mini-swe-agent · max reasoning | OpenAI | 72.7% |
| 5 | [Claude Fable 5](/models/claude-fable) | mini-swe-agent · max reasoning | Anthropic | 69.7% |
| 6 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | mini-swe-agent · max reasoning | OpenAI | 69.6% |
| 7 | [GLM-5.3](/models/glm-5-3) | mini-swe-agent · max reasoning | Z.AI | 69.0% |
| 8 | [Kimi K3](/models/kimi-k3) | mini-swe-agent · max reasoning | Moonshot AI | 68.5% |
| 9 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | mini-swe-agent · max reasoning | OpenAI | 67.2% |
| 10 | [GPT-5.5](/models/gpt-5-5) | mini-swe-agent · xhigh reasoning | OpenAI | 67.0% |
| 11 | [Grok 4.6](/models/grok-4-6) | mini-swe-agent · xhigh reasoning | xAI | 66.7% |
| 12 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | mini-swe-agent · high reasoning | Google | 65.3% |
| 13 | [GLM-5.3-Flash](/models/glm-5-3-flash) | mini-swe-agent · max reasoning | Z.AI | 63.4% |
| 14 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | mini-swe-agent · max reasoning | DeepSeek | 62.8% |
| 15 | [Claude Opus 4.8](/models/claude-opus-4-8) | mini-swe-agent · max reasoning | Anthropic | 59.0% |
| 16 | [Qwen3.8 Max](/models/qwen3-8-max) | mini-swe-agent · xhigh reasoning | Alibaba | 57.5% |
| 17 | [Muse Spark 1.2](/models/muse-spark-1-2) | mini-swe-agent · xhigh reasoning | Meta | 54.9% |
| 18 | [Claude Sonnet 5](/models/claude-sonnet-5) | mini-swe-agent · max reasoning | Anthropic | 53.8% |
| 19 | [Grok 4.5](/models/grok-4-5) | mini-swe-agent · high reasoning | xAI | 53.8% |
| 20 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | mini-swe-agent · max reasoning | DeepSeek | 53.3% |
| 21 | [Muse Spark 1.1](/models/muse-spark-1-1) | mini-swe-agent · xhigh reasoning | Meta | 53.3% |
| 22 | [GPT-5.4](/models/gpt-5-4) | mini-swe-agent · xhigh reasoning | OpenAI | 51.8% |
| 23 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | mini-swe-agent · high reasoning | Google | 46.7% |
| 24 | [GLM-5.2](/models/glm-5-2) | mini-swe-agent · max reasoning | Z.AI | 43.8% |
| 25 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | mini-swe-agent · high reasoning | Google | 36.1% |
| 26 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | mini-swe-agent | Moonshot AI | 30.5% |
| 27 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | mini-swe-agent · high reasoning | Anthropic | 29.9% |
| 28 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | mini-swe-agent · high reasoning | Google | 11.7% |

## FAQ

### What does DeepSWE measure?

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

### Which model leads the published DeepSWE snapshot?

Gemini 3.8 Flash currently leads the published DeepSWE snapshot with a score of 73.8%.

### How many models are evaluated on DeepSWE?

The September 1, 2026 contains 28 AI models.

### Does DeepSWE affect BenchLM's overall score?

Not directly. DeepSWE is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
