# CWE-bench v1

> Defensive vulnerability patching on 120 held-out tasks spanning 73 weakness types.

Canonical page: https://benchlm.ai/benchmarks/cwe-bench-v1

- Category: [Agentic](/agentic)
- Last updated: September 30, 2026

## About CWE-bench v1

- Year: 2026
- Tasks: Audit and patch held-out C/C++, Go, Java, TypeScript/JavaScript, Python, and Rust repositories
- Format: Pass@1
- Difficulty: Frontier agent evaluation
- Paper: [CWE-bench v1 leaderboard and methodology](https://cwe-bench.com/)

The deterministic programmatic pass@1 score requires blocking the exploit while preserving existing tests. Models use high reasoning, four rollouts per task, and a one-hour limit per rollout. Version 1 uses 120 tasks and stays separate from v0, pass@4, and judge-panel scores.

CWE-bench v1 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (11 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 68.0% |
| 2 | [Gemini 4 Argon](/models/gemini-4-argon) | Google | 68.0% |
| 3 | [Grok 4.7](/models/grok-4-7) | xAI | 68.0% |
| 4 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 67.0% |
| 5 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 58.0% |
| 6 | [Grok 4.6](/models/grok-4-6) | xAI | 57.0% |
| 7 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 55.0% |
| 8 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 55.0% |
| 9 | [Hy4 preview](/models/hy4-preview) | Tencent | 53.0% |
| 10 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 52.0% |
| 11 | [Inkling](/models/inkling) | Thinking Machines Lab | 37.0% |

## FAQ

### What does CWE-bench v1 measure?

Defensive vulnerability patching on 120 held-out tasks spanning 73 weakness types.

### Which model scores highest on CWE-bench v1?

GPT-6 Astra by OpenAI currently leads with a score of 68.0% on CWE-bench v1.

### How many models are evaluated on CWE-bench v1?

11 AI models have been evaluated on CWE-bench v1 on BenchLM.

### Does CWE-bench v1 affect BenchLM's overall score?

Not directly. CWE-bench v1 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on CWE-bench v1

- [GPT-6 Astra vs Gemini 4 Argon](/compare/gemini-4-argon-vs-gpt-6-astra)
- [Gemini 4 Argon vs Grok 4.7](/compare/gemini-4-argon-vs-grok-4-7)
- [Grok 4.7 vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-grok-4-7)
- [Claude Opus 5.5 vs Claude Fable 5.1](/compare/claude-fable-5-1-vs-claude-opus-5-5)
