# Vals Vibe Code Bench 1-100 (Vibe Code Bench 1-100)

> Can a model extend one working web application across a long sequence of dependent requests?

Canonical page: https://benchlm.ai/benchmarks/vibe-code-bench-1-100

- Category: [external](/external)
- Last updated: September 16, 2026

## About Vibe Code Bench 1-100

- Year: 2026
- Tasks: Long chains of dependent web-application change requests
- Format: Accuracy score
- Difficulty: Long-horizon software delivery
- Paper: [Vals Vibe Code Bench 1-100](https://www.vals.ai/benchmarks/vcb-1-100)

Each run starts from a working web application and applies a long chain of dependent change requests, so an early mistake compounds through the rest of the sequence. That is a different question from the single-build Vibe Code Bench v1.1 board, which is why BenchLM keeps it on its own key. Private task set, external runner, display only.

Vibe Code Bench 1-100 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | OpenHands · OpenHands | Anthropic | 28.53% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | OpenHands · OpenHands | Anthropic | 28.00% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | OpenHands · OpenHands | OpenAI | 27.64% |
| 4 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenHands · OpenHands | OpenAI | 22.59% |
| 5 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | OpenHands · OpenHands | Meta | 20.46% |
| 6 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenHands · OpenHands | OpenAI | 20.01% |
| 7 | [GLM-5.3](/models/glm-5-3) | OpenHands · OpenHands | Z.AI | 19.99% |
| 8 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | OpenHands · OpenHands | Google | 18.77% |
| 9 | [Kimi K3](/models/kimi-k3) | OpenHands · OpenHands | Moonshot AI | 18.24% |
| 10 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | OpenHands · OpenHands | DeepSeek | 17.53% |
| 11 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | OpenHands · OpenHands | DeepSeek | 16.38% |
| 12 | [GLM-5.3-Flash](/models/glm-5-3-flash) | OpenHands · OpenHands | Z.AI | 16.03% |
| 13 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenHands · OpenHands | OpenAI | 14.82% |
| 14 | [Grok 4.6](/models/grok-4-6) | OpenHands · OpenHands | xAI | 14.75% |
| 15 | [Claude Sonnet 5](/models/claude-sonnet-5) | OpenHands · OpenHands | Anthropic | 13.82% |
| 16 | [Qwen3.8 Max](/models/qwen3-8-max) | OpenHands · OpenHands | Alibaba | 12.83% |
| 17 | [MiniMax M3](/models/minimax-m3) | OpenHands · OpenHands | MiniMax | 9.22% |
| 18 | [Inkling](/models/inkling) | OpenHands · OpenHands | Thinking Machines Lab | 7.33% |
| 19 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | OpenHands · OpenHands | Google | 6.69% |

## FAQ

### What does Vibe Code Bench 1-100 measure?

Can a model extend one working web application across a long sequence of dependent requests?

### Which model leads the published Vibe Code Bench 1-100 snapshot?

Claude Opus 5 currently leads the published Vibe Code Bench 1-100 snapshot with a score of 28.53%.

### How many models are evaluated on Vibe Code Bench 1-100?

The September 16, 2026 contains 19 AI models.

### Does Vibe Code Bench 1-100 affect BenchLM's overall score?

Not directly. Vibe Code Bench 1-100 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
