# VulcanBench v3

> An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories.

Canonical page: https://benchlm.ai/benchmarks/vulcanbench

- Category: [Coding](/coding)
- Last updated: September 18, 2026

## About VulcanBench v3

- Year: 2026
- Tasks: 23 post-cutoff repository tasks in the v3 report
- Format: Pass@1 with low, medium, and high effort
- Difficulty: Professional multi-file software engineering
- Paper: [VulcanBench](https://github.com/morganlinton/VulcanBench/tree/main)

VulcanBench v3 evaluates repository-level patches with deterministic hidden tests inside network-isolated Docker sandboxes. The August 24 board covers 23 tasks, 14 models, 39 model-by-effort columns, and 1,582 runs. We mirror each benchmark-owner best-per-model row as display-only evidence. Product-harness studies, extended-budget ablations, and other rows outside the raw uniform-harness board stay separate.

VulcanBench v3 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Grok 4.5](/models/grok-4-5) | xAI | 89.9% |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 89.5% |
| 3 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 88.4% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 87.0% |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 87.0% |
| 6 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 87.0% |
| 7 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 87.0% |
| 8 | [Grok 4.6](/models/grok-4-6) | xAI | 87.0% |
| 9 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 85.5% |
| 10 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 82.6% |
| 11 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 81.2% |
| 12 | [GLM-5.3](/models/glm-5-3) | Z.AI | 78.3% |
| 13 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 76.2% |
| 14 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 73.7% |

## FAQ

### What does VulcanBench v3 measure?

An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories.

### Which model scores highest on VulcanBench v3?

Grok 4.5 by xAI currently leads with a score of 89.9% on VulcanBench v3.

### How many models are evaluated on VulcanBench v3?

14 AI models have been evaluated on VulcanBench v3 on BenchLM.

### Does VulcanBench v3 affect BenchLM's overall score?

Not directly. VulcanBench v3 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on VulcanBench v3

- [Grok 4.5 vs Claude Fable 5](/compare/claude-fable-vs-grok-4-5)
- [Claude Fable 5 vs DeepSeek V4 Flash 0731](/compare/claude-fable-vs-deepseek-v4-flash-0731)
- [DeepSeek V4 Flash 0731 vs Claude Opus 5](/compare/claude-opus-5-vs-deepseek-v4-flash-0731)
- [Claude Opus 5 vs GPT-5.6 Sol](/compare/claude-opus-5-vs-gpt-5-6-sol)
