# Vibe Code Bench v1.1 (Vibe Code Bench)

> Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Canonical page: https://benchlm.ai/benchmarks/vibecodebench

- Category: [Coding](/coding)
- Last updated: September 10, 2026

## About Vibe Code Bench

- Year: 2026
- Tasks: End-to-end web application builds
- Format: Full-stack app implementation benchmark
- Difficulty: End-to-end software delivery
- Paper: [Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development](https://www.vals.ai/benchmarks/vibe-code)

Vibe Code Bench v1.1 asks models to build full web apps with services such as Supabase, Stripe test mode, email, browsing, and file editing available. The score is overall application pass accuracy across private end-to-end app tasks.

Vibe Code Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (94 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | OpenHands · OpenHands | Anthropic | 90.35% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | OpenHands · OpenHands | Anthropic | 90.26% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | OpenHands · OpenHands | OpenAI | 89.59% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | OpenHands · OpenHands | Anthropic | 88.40% |
| 5 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | OpenHands · OpenHands | Meta | 85.86% |
| 6 | [Kimi K3](/models/kimi-k3) | OpenHands · OpenHands | Moonshot AI | 84.96% |
| 7 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | OpenHands · OpenHands | DeepSeek | 84.74% |
| 8 | [Muse Spark 1.3](/models/muse-spark-1-3) | OpenHands · OpenHands | Meta | 82.86% |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | OpenHands · OpenHands | Anthropic | 82.72% |
| 10 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | OpenHands · OpenHands | DeepSeek | 82.30% |
| 11 | [Claude Sonnet 5](/models/claude-sonnet-5) | OpenHands · OpenHands | Anthropic | 81.33% |
| 12 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenHands · OpenHands | OpenAI | 80.50% |
| 13 | [Muse Spark 1.2](/models/muse-spark-1-2) | OpenHands · OpenHands | Meta | 79.10% |
| 14 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | OpenHands · OpenHands | Google | 78.65% |
| 15 | [GLM-5.3](/models/glm-5-3) | OpenHands · OpenHands | Z.AI | 78.13% |
| 16 | [Claude Opus 4.8 Claude Code](https://www.vals.ai/models/anthropic_claude-opus-4-8-claude-code) | Claude Code · Claude Code | Anthropic | 77.48% |
| 17 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenHands · OpenHands | OpenAI | 77.06% |
| 18 | [Grok 4.6](/models/grok-4-6) | OpenHands · OpenHands | xAI | 76.24% |
| 19 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | OpenHands · OpenHands | DeepSeek | 74.74% |
| 20 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenHands · OpenHands | OpenAI | 74.59% |
| 21 | [Muse Spark 1.1](/models/muse-spark-1-1) | OpenHands · OpenHands | Meta | 72.16% |
| 22 | [Claude Opus 4.7](/models/claude-opus-4-7) | OpenHands · OpenHands | Anthropic | 71.00% |
| 23 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | OpenHands · OpenHands | Google | 70.39% |
| 24 | [GPT-5.5](/models/gpt-5-5) | OpenHands · OpenHands | OpenAI | 69.85% |
| 25 | [Grok 4.5](/models/grok-4-5) | OpenHands · OpenHands | xAI | 69.00% |
| 26 | [GPT-5.4](/models/gpt-5-4) | OpenHands · OpenHands | OpenAI | 67.42% |
| 27 | [GPT-5.5 Factory](https://www.vals.ai/models/openai_gpt-5.5-factory) | Factory · Factory | OpenAI | 67.39% |
| 28 | [Qwen3.8-27B](/models/qwen3-8-27b) | openhands · openhands | Alibaba | 64.85% |
| 29 | [Qwen3.8 Max](/models/qwen3-8-max) | OpenHands · OpenHands | Alibaba | 64.70% |
| 30 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | OpenHands · OpenHands | Google | 64.00% |
| 31 | [GLM-5.2](/models/glm-5-2) | OpenHands · OpenHands | Z.AI | 63.96% |
| 32 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenHands · OpenHands | OpenAI | 61.77% |
| 33 | [GPT-5.5 Codex](https://www.vals.ai/models/openai_gpt-5.5-codex) | Codex · Codex | OpenAI | 58.19% |
| 34 | [Claude Opus 4.6](/models/claude-opus-4-6) | OpenHands · OpenHands | Anthropic | 57.57% |
| 35 | [Claude Sonnet 4.6 Claude Code](https://www.vals.ai/models/anthropic_claude-sonnet-4-6-claude-code) | Claude Code · Claude Code | Anthropic | 55.77% |
| 36 | [GPT-5.2](/models/gpt-5-2) | OpenHands · OpenHands | OpenAI | 53.50% |
| 37 | [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) | OpenHands · OpenHands | Anthropic | 53.50% |
| 38 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | OpenHands · OpenHands | Anthropic | 51.48% |
| 39 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | OpenHands · OpenHands | DeepSeek | 49.93% |
| 40 | [Composer 2.5](/models/composer-2-5) | Cursor CLI · Cursor CLI | Cursor | 49.61% |
| 41 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | OpenHands · OpenHands | Google | 48.68% |
| 42 | [GPT-5.4](/models/gpt-5-4) | Codex · Codex | OpenAI | 48.47% |
| 43 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenHands · OpenHands | OpenAI | 47.97% |
| 44 | [Qwen3.7 Max](/models/qwen3-7-max) | OpenHands · OpenHands | Alibaba | 47.67% |
| 45 | [MiniMax M3](/models/minimax-m3) | OpenHands · OpenHands | MiniMax | 47.57% |
| 46 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | OpenHands · OpenHands | Moonshot AI | 47.21% |
| 47 | [Qwen3.7 Plus](/models/qwen3-7-plus) | OpenHands · OpenHands | Alibaba | 46.39% |
| 48 | [MiMo-V2.5](/models/mimo-v2-5) | OpenHands · OpenHands | Xiaomi | 42.17% |
| 49 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenHands · OpenHands | OpenAI | 37.91% |
| 50 | [Kimi K2.6](/models/kimi-2-6) | OpenHands · OpenHands | Moonshot AI | 37.89% |
| 51 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | OpenHands · OpenHands | Google | 37.16% |
| 52 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | OpenHands · OpenHands | Xiaomi | 34.11% |
| 53 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | OpenHands · OpenHands | Google | 32.03% |
| 54 | [GLM-5.1](/models/glm-5-1) | OpenHands · OpenHands | Z.AI | 31.46% |
| 55 | [GLM-5.3-Flash](/models/glm-5-3-flash) | OpenHands · OpenHands | Z.AI | 30.76% |
| 56 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenHands · OpenHands | OpenAI | 26.10% |
| 57 | [Qwen3.6 Plus](/models/qwen3-6-plus) | OpenHands · OpenHands | Alibaba | 25.57% |
| 58 | [GPT-5.1](/models/gpt-5-1) | OpenHands · OpenHands | OpenAI | 24.61% |
| 59 | [GLM 5 Thinking](https://www.vals.ai/models/zai_glm-5-thinking) | OpenHands · OpenHands | zAI | 23.36% |
| 60 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | OpenHands · OpenHands | Anthropic | 22.62% |
| 61 | [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) | OpenHands · OpenHands | OpenAI | 22.17% |
| 62 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | OpenHands · OpenHands | Anthropic | 20.63% |
| 63 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | OpenHands · OpenHands | Google | 20.20% |
| 64 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | OpenHands · OpenHands | OpenAI | 20.09% |
| 65 | [Muse Spark](/models/muse-spark) | OpenHands · OpenHands | Meta | 19.67% |
| 66 | [Grok 4.3](/models/grok-4-3) | OpenHands · OpenHands | xAI | 19.40% |
| 67 | [Inkling](/models/inkling) | OpenHands · OpenHands | Thinking Machines Lab | 19.21% |
| 68 | [Inkling-Small](/models/inkling-small) | OpenHands · OpenHands | Thinking Machines Lab | 19.05% |
| 69 | [Kimi K2.5 Thinking](https://www.vals.ai/models/kimi_kimi-k2.5-thinking) | OpenHands · OpenHands | Moonshot AI | 17.54% |
| 70 | [Qwen3.5 Plus Thinking](https://www.vals.ai/models/alibaba_qwen3.5-plus-thinking) | OpenHands · OpenHands | Alibaba | 15.74% |
| 71 | [MiniMax M2.5](/models/minimax-m2-5) | OpenHands · OpenHands | MiniMax | 14.85% |
| 72 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | OpenHands · OpenHands | Google | 14.30% |
| 73 | [GPT-5 mini](/models/gpt-5-mini) | OpenHands · OpenHands | OpenAI | 14.17% |
| 74 | [Grok Build 0.1](/models/grok-build-0-1) | Grok Build · Grok Build | xAI | 13.35% |
| 75 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | OpenHands · OpenHands | OpenAI | 13.12% |
| 76 | [Qwen3.6-27B](/models/qwen3-6-27b) | OpenHands · OpenHands | Alibaba | 11.94% |
| 77 | [MiniMax M2.7](/models/minimax-m2-7) | OpenHands · OpenHands | MiniMax | 11.93% |
| 78 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | OpenHands · OpenHands | Anthropic | 11.39% |
| 79 | [Laguna M.1](/models/laguna-m-1) | OpenHands · OpenHands | Poolside | 11.04% |
| 80 | [Nemotron 3 Ultra 550b A55b](https://www.vals.ai/models/nvidia_nemotron-3-ultra-550b-a55b) | OpenHands · OpenHands | nvidia | 7.64% |
| 81 | [Swe 1.6 Fast](https://www.vals.ai/models/devin_swe-1-6-fast) | Devin CLI · Devin CLI | devin | 7.22% |
| 82 | [Laguna XS.2](/models/laguna-xs-2) | OpenHands · OpenHands | Poolside | 5.21% |
| 83 | [DeepSeek V3p2 Thinking](https://www.vals.ai/models/fireworks_deepseek-v3p2-thinking) | OpenHands · OpenHands | Fireworks | 5.11% |
| 84 | [Grok 4.20 0309 Reasoning](https://www.vals.ai/models/grok_grok-4.20-0309-reasoning) | OpenHands · OpenHands | xAI | 4.06% |
| 85 | [Qwen3 Max](/models/qwen3-max) | OpenHands · OpenHands | Alibaba | 3.51% |
| 86 | [GLM-4.6](/models/glm-4-6) | OpenHands · OpenHands | Z.AI | 3.09% |
| 87 | [Ling 3.0 Flash 2607](https://www.vals.ai/models/ant_ling-3.0-flash-2607) | OpenHands · OpenHands | ant | 2.91% |
| 88 | [Mistral Medium 3.5](https://www.vals.ai/models/mistralai_mistral-medium-3.5) | OpenHands · OpenHands | Mistral | 2.89% |
| 89 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | OpenHands · OpenHands | xAI | 1.20% |
| 90 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | OpenHands · OpenHands | Google | 0.40% |
| 91 | [Gemini 3.1 Flash Lite Preview](https://www.vals.ai/models/google_gemini-3.1-flash-lite-preview) | OpenHands · OpenHands | Google | 0.00% |
| 92 | [Nemotron Lightning 3p5 30b A3b](https://www.vals.ai/models/fireworks_nemotron-lightning-3p5-30b-a3b) | OpenHands · OpenHands | Fireworks | 0.00% |
| 93 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | OpenHands · OpenHands | xAI | 0.00% |
| 94 | [Mistral Small 2603](https://www.vals.ai/models/mistralai_mistral-small-2603) | OpenHands · OpenHands | Mistral | 0.00% |

## FAQ

### What does Vibe Code Bench measure?

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

### Which model leads the published Vibe Code Bench snapshot?

Claude Fable 5 currently leads the published Vibe Code Bench snapshot with a score of 90.35%.

### How many models are evaluated on Vibe Code Bench?

The September 10, 2026 contains 94 AI models.

### Does Vibe Code Bench affect BenchLM's overall score?

Not directly. Vibe Code Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
