Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Vibe Code Bench v1.1 (Vibe Code Bench)

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Vibe Code Bench v1.1

BenchLM mirrors the public Vals AI Vibe Code Bench v1.1 leaderboard captured from https://www.vals.ai/benchmarks/vibe-code and updated by Vals on September 10, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vibe Code Bench v1.1 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

94 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

Vibe Code score on Vibe Code Bench — September 10, 2026

We mirror the published vibe code score view for Vibe Code Bench. Claude Fable 5 leads the public snapshot at 90.35%, followed by Claude Fable 5.1 (90.26%) and GPT-6 Astra (89.59%). We do not use these results to rank models overall.

94 modelsCodingCurrentDisplay onlyUpdated September 10, 2026

Vibe Code score table (94 models)

Score
1
Claude Fable 5Anthropic · ClosedOpenHandsOpenHands
90.35%
2
Claude Fable 5.1Anthropic · ClosedOpenHandsOpenHands
90.26%
3
GPT-6 AstraOpenAI · ClosedOpenHandsOpenHands
89.59%
4
Claude Opus 5Anthropic · ClosedOpenHandsOpenHands
88.40%
5
Muse Spark 1.3 MaxMetaOpenHandsOpenHands
85.86%
6
Kimi K3Moonshot AI · ClosedOpenHandsOpenHands
84.96%
7
DeepSeek V4.1 FlashDeepSeek · Open weightOpenHandsOpenHands
84.74%
8
Muse Spark 1.3Meta · ClosedOpenHandsOpenHands
82.86%
9
Claude Opus 4.8Anthropic · ClosedOpenHandsOpenHands
82.72%
10
DeepSeek V4 Pro 0813DeepSeek · ClosedOpenHandsOpenHands
82.30%
11
Claude Sonnet 5Anthropic · ClosedOpenHandsOpenHands
81.33%
12
GPT-5.6 SolOpenAI · ClosedOpenHandsOpenHands
80.50%
13
Muse Spark 1.2Meta · ClosedOpenHandsOpenHands
79.10%
14
Gemini 3.8 FlashGoogle · ClosedOpenHandsOpenHands
78.65%
15
GLM-5.3Z.AI · Open weightOpenHandsOpenHands
78.13%
16
Claude Opus 4.8 Claude CodeAnthropicClaude CodeClaude Code
77.48%
17
GPT-5.6 LunaOpenAI · ClosedOpenHandsOpenHands
77.06%
18
Grok 4.6xAI · ClosedOpenHandsOpenHands
76.24%
19
DeepSeek V4 Flash 0731DeepSeek · ClosedOpenHandsOpenHands
74.74%
20
GPT-5.6 TerraOpenAI · ClosedOpenHandsOpenHands
74.59%
21
Muse Spark 1.1Meta · ClosedOpenHandsOpenHands
72.16%
22
Claude Opus 4.7Anthropic · ClosedOpenHandsOpenHands
71.00%
23
Gemini 3.7 FlashGoogle · ClosedOpenHandsOpenHands
70.39%
24
GPT-5.5OpenAI · ClosedOpenHandsOpenHands
69.85%
25
Grok 4.5xAI · ClosedOpenHandsOpenHands
69.00%
26
GPT-5.4OpenAI · ClosedOpenHandsOpenHands
67.42%
27
GPT-5.5 FactoryOpenAIFactoryFactory
67.39%
28
Qwen3.8-27BAlibaba · Open weightopenhandsopenhands
64.85%
29
Qwen3.8 MaxAlibaba · Open weightOpenHandsOpenHands
64.70%
30
Gemini 3.6 FlashGoogle · ClosedOpenHandsOpenHands
64.00%
31
GLM-5.2Z.AI · Open weightOpenHandsOpenHands
63.96%
32
GPT-5.3 CodexOpenAI · ClosedOpenHandsOpenHands
61.77%
33
GPT-5.5 CodexOpenAICodexCodex
58.19%
34
Claude Opus 4.6Anthropic · ClosedOpenHandsOpenHands
57.57%
35
Claude Sonnet 4.6 Claude CodeAnthropicClaude CodeClaude Code
55.77%
36
GPT-5.2OpenAI · ClosedOpenHandsOpenHands
53.50%
37
Claude Opus 4.6 (Adaptive)Anthropic · ClosedOpenHandsOpenHands
53.50%
38
Claude Sonnet 4.6Anthropic · ClosedOpenHandsOpenHands
51.48%
39
DeepSeek V4 Pro 0813DeepSeek · ClosedOpenHandsOpenHands
49.93%
40
Composer 2.5Cursor · ClosedCursor CLICursor CLI
49.61%
41
Gemini 3.5 FlashGoogle · ClosedOpenHandsOpenHands
48.68%
42
GPT-5.4OpenAI · ClosedCodexCodex
48.47%
43
GPT-5.4 miniOpenAI · ClosedOpenHandsOpenHands
47.97%
44
Qwen3.7 MaxAlibaba · ClosedOpenHandsOpenHands
47.67%
45
MiniMax M3MiniMax · Open weightOpenHandsOpenHands
47.57%
46
Kimi K2.7 CodeMoonshot AI · Open weightOpenHandsOpenHands
47.21%
47
Qwen3.7 PlusAlibaba · ClosedOpenHandsOpenHands
46.39%
48
MiMo-V2.5Xiaomi · ClosedOpenHandsOpenHands
42.17%
49
GPT-5.2-CodexOpenAI · ClosedOpenHandsOpenHands
37.91%
50
Kimi K2.6Moonshot AI · Open weightOpenHandsOpenHands
37.89%
51
Gemini 3.5 Flash-LiteGoogle · ClosedOpenHandsOpenHands
37.16%
52
MiMo-V2.5-ProXiaomi · ClosedOpenHandsOpenHands
34.11%
53
Gemini 3.1 Pro PreviewGoogleOpenHandsOpenHands
32.03%
54
GLM-5.1Z.AI · Open weightOpenHandsOpenHands
31.46%
55
GLM-5.3-FlashZ.AI · Open weightOpenHandsOpenHands
30.76%
56
GPT-5.4 nanoOpenAI · ClosedOpenHandsOpenHands
26.10%
57
Qwen3.6 PlusAlibaba · ClosedOpenHandsOpenHands
25.57%
58
GPT-5.1OpenAI · ClosedOpenHandsOpenHands
24.61%
59
GLM 5 ThinkingzAIOpenHandsOpenHands
23.36%
60
Claude Sonnet 4.5 ThinkingAnthropic · ClosedOpenHandsOpenHands
22.62%
61
GPT-5.1-Codex-MaxOpenAI · ClosedOpenHandsOpenHands
22.17%
62
Claude Opus 4.5 ThinkingAnthropic · ClosedOpenHandsOpenHands
20.63%
63
Gemini 3 Flash PreviewGoogleOpenHandsOpenHands
20.20%
64
GPT-5OpenAIOpenHandsOpenHands
20.09%
65
Muse SparkMeta · ClosedOpenHandsOpenHands
19.67%
66
Grok 4.3xAI · ClosedOpenHandsOpenHands
19.40%
67
InklingThinking Machines Lab · Open weightOpenHandsOpenHands
19.21%
68
Inkling-SmallThinking Machines Lab · Open weightOpenHandsOpenHands
19.05%
69
Kimi K2.5 ThinkingMoonshot AIOpenHandsOpenHands
17.54%
70
Qwen3.5 Plus ThinkingAlibabaOpenHandsOpenHands
15.74%
71
MiniMax M2.5MiniMax · ClosedOpenHandsOpenHands
14.85%
72
Gemini 3 Pro PreviewGoogleOpenHandsOpenHands
14.30%
73
GPT-5 miniOpenAI · ClosedOpenHandsOpenHands
14.17%
74
Grok Build 0.1xAI · ClosedGrok BuildGrok Build
13.35%
75
GPT-5.1-CodexOpenAI · ClosedOpenHandsOpenHands
13.12%
76
Qwen3.6-27BAlibaba · Open weightOpenHandsOpenHands
11.94%
77
MiniMax M2.7MiniMax · Open weightOpenHandsOpenHands
11.93%
78
Claude Haiku 4.5 ThinkingAnthropic · ClosedOpenHandsOpenHands
11.39%
79
Laguna M.1Poolside · ClosedOpenHandsOpenHands
11.04%
80
Nemotron 3 Ultra 550b A55bnvidiaOpenHandsOpenHands
7.64%
81
Swe 1.6 FastdevinDevin CLIDevin CLI
7.22%
82
Laguna XS.2Poolside · Open weightOpenHandsOpenHands
5.21%
83
DeepSeek V3p2 ThinkingFireworksOpenHandsOpenHands
5.11%
84
Grok 4.20 0309 ReasoningxAIOpenHandsOpenHands
4.06%
85
Qwen3 MaxAlibaba · ClosedOpenHandsOpenHands
3.51%
86
GLM-4.6Z.AI · Open weightOpenHandsOpenHands
3.09%
87
Ling 3.0 Flash 2607antOpenHandsOpenHands
2.91%
88
Mistral Medium 3.5MistralOpenHandsOpenHands
2.89%
89
Grok 4.1 Fast (Reasoning)xAI · ClosedOpenHandsOpenHands
1.20%
90
Gemini 2.5 ProGoogle · ClosedOpenHandsOpenHands
0.40%
91
Gemini 3.1 Flash Lite PreviewGoogleOpenHandsOpenHands
0.00%
92
Nemotron Lightning 3p5 30b A3bFireworksOpenHandsOpenHands
0.00%
93
Grok 4 Fast (Reasoning)xAI · ClosedOpenHandsOpenHands
0.00%
94
Mistral Small 2603MistralOpenHandsOpenHands
0.00%

The published Vibe Code Bench snapshot places Claude Fable 5 first at 90.35%. The third row is 0.76 points behind. The broader top-10 range is 8.06 points, so many of the published results sit in a relatively narrow band.

94 models have been evaluated on Vibe Code Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Vibe Code Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vibe Code Bench

Year

2026

Tasks

End-to-end web application builds

Format

Full-stack app implementation benchmark

Difficulty

End-to-end software delivery

Vibe Code Bench v1.1 asks models to build full web apps with services such as Supabase, Stripe test mode, email, browsing, and file editing available. The score is overall application pass accuracy across private end-to-end app tasks.

BenchLM freshness & provenance

Version

Vibe Code Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vibe Code Bench measure?

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Which model leads the published Vibe Code Bench snapshot?

Claude Fable 5 currently leads the published Vibe Code Bench snapshot with 90.35% vibe code score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vibe Code Bench?

The September 10, 2026 snapshot contains 94 AI models.

Last updated: September 10, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.