Skip to main content

Benchmark profile

Vibe Code Bench v1.1 (Vibe Code Bench)

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Data verified

How BenchLM shows Vibe Code Bench v1.1

BenchLM mirrors the public Vals AI Vibe Code Bench v1.1 leaderboard captured from https://www.vals.ai/benchmarks/vibe-code and updated by Vals on July 23, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vibe Code Bench v1.1 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

76 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

Vibe Code score on Vibe Code Bench — July 23, 2026

BenchLM mirrors the published vibe code score view for Vibe Code Bench. Claude Fable 5 leads the public snapshot at 90.35% , followed by Claude Opus 5 (88.40%) and Kimi K3 (84.96%). BenchLM does not use these results to rank models overall.

76 modelsCodingCurrentDisplay onlyUpdated July 23, 2026

Vibe Code score table (76 models)

Score
1
Claude Fable 5AnthropicOpenHands
90.35%
2
Claude Opus 5AnthropicOpenHands
88.40%
3
Kimi K3Moonshot AIOpenHands
84.96%
4
Claude Opus 4.8AnthropicOpenHands
82.72%
5
Claude Sonnet 5AnthropicOpenHands
81.33%
6
GPT-5.6 SolOpenAIOpenHands
80.50%
7
Claude Opus 4.8 Claude CodeAnthropicClaude Code
77.48%
8
GPT-5.6 LunaOpenAIOpenHands
77.06%
9
Muse Spark 1.1MetaOpenHands
72.16%
10
Claude Opus 4.7AnthropicOpenHands
71.00%
11
GPT-5.5OpenAIOpenHands
69.85%
12
Grok 4.5xAIOpenHands
69.00%
13
GPT-5.6 TerraOpenAIOpenHands
67.84%
14
GPT-5.4OpenAIOpenHands
67.42%
15
GPT-5.5 FactoryOpenAIFactory
67.39%
16
GLM 5.2Zhipu AIOpenHands
63.96%
17
GPT-5.3 CodexOpenAIOpenHands
61.77%
18
GPT-5.5 CodexOpenAICodex
58.19%
19
Claude Opus 4.6AnthropicOpenHands
57.57%
20
Gemini 3.6 FlashGoogleOpenHands
57.32%
21
Claude Sonnet 4.6 Claude CodeAnthropicClaude Code
55.77%
22
GPT-5.2OpenAIOpenHands
53.50%
23
Claude Opus 4.6 ThinkingAnthropicOpenHands
53.50%
24
Claude Sonnet 4.6AnthropicOpenHands
51.48%
25
DeepSeek V4 ProDeepSeekOpenHands
49.93%
26
Composer 2.5CursorCursor CLI
49.61%
27
Gemini 3.5 FlashGoogleOpenHands
48.68%
28
GPT-5.4OpenAICodex
48.47%
29
GPT-5.4 MiniOpenAIOpenHands
47.97%
30
Qwen3.7 MaxAlibabaOpenHands
47.67%
31
MiniMax M3MiniMaxOpenHands
47.57%
32
Kimi K2.7 CodeMoonshot AIOpenHands
47.21%
33
Qwen3.7 PlusAlibabaOpenHands
46.39%
34
Mimo V2.5XiaomiOpenHands
42.17%
35
GPT-5.2 CodexOpenAIOpenHands
37.91%
36
Kimi K2.6Moonshot AIOpenHands
37.89%
37
Gemini 3.5 Flash LiteGoogleOpenHands
37.16%
38
Mimo V2.5 ProXiaomiOpenHands
34.11%
39
Gemini 3.1 Pro PreviewGoogleOpenHands
32.03%
40
GLM 5.1Zhipu AIOpenHands
31.46%
41
GPT-5.4 NanoOpenAIOpenHands
26.10%
42
Qwen3.6 PlusAlibabaOpenHands
25.57%
43
GPT-5.1OpenAIOpenHands
24.61%
44
GLM 5 ThinkingZhipu AIOpenHands
23.36%
45
22.62%
46
GPT-5.1 Codex MaxOpenAIOpenHands
22.17%
47
20.63%
48
Gemini 3 Flash PreviewGoogleOpenHands
20.20%
49
GPT-5OpenAIOpenHands
20.09%
50
Muse SparkMetaOpenHands
19.67%
51
Grok 4.3xAIOpenHands
19.40%
52
InklingThinkingmachinesOpenHands
19.21%
53
Kimi K2.5 ThinkingMoonshot AIOpenHands
17.54%
54
Qwen3.5 Plus ThinkingAlibabaOpenHands
15.74%
55
MiniMax M2.5MiniMaxOpenHands
14.85%
56
Gemini 3 Pro PreviewGoogleOpenHands
14.30%
57
GPT-5 MiniOpenAIOpenHands
14.17%
58
Grok Build 0.1xAIGrok Build
13.35%
59
GPT-5.1 CodexOpenAIOpenHands
13.12%
60
Qwen3.6 27bAlibabaOpenHands
11.94%
61
MiniMax M2.7MiniMaxOpenHands
11.93%
62
11.39%
63
Laguna M.1PoolsideOpenHands
11.04%
64
7.64%
65
Swe 1.6 FastDevinDevin CLI
7.22%
66
Laguna Xs.2PoolsideOpenHands
5.21%
67
DeepSeek V3p2 ThinkingFireworks AIOpenHands
5.11%
68
4.06%
69
Qwen3 MaxAlibabaOpenHands
3.51%
70
GLM 4.6Zhipu AIOpenHands
3.09%
71
Mistral Medium 3.5Mistral AIOpenHands
2.89%
72
1.20%
73
Gemini 2.5 ProGoogleOpenHands
0.40%
74
0.00%
75
0.00%
76
Mistral Small 2603Mistral AIOpenHands
0.00%

The published Vibe Code Bench snapshot places Claude Fable 5 first at 90.35%. The third row is 5.39 points behind. The broader top-10 range is 19.35 points, so the table still separates the published systems.

76 models have been evaluated on Vibe Code Bench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Vibe Code Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vibe Code Bench

Year

2026

Tasks

End-to-end web application builds

Format

Full-stack app implementation benchmark

Difficulty

End-to-end software delivery

Vibe Code Bench v1.1 asks models to build full web apps with services such as Supabase, Stripe test mode, email, browsing, and file editing available. The score is overall application pass accuracy across private end-to-end app tasks.

BenchLM freshness & provenance

Version

Vibe Code Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vibe Code Bench measure?

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Which model leads the published Vibe Code Bench snapshot?

Claude Fable 5 currently leads the published Vibe Code Bench snapshot with 90.35% vibe code score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vibe Code Bench?

76 AI models are included in BenchLM's mirrored Vibe Code Bench snapshot, based on the public leaderboard captured on July 23, 2026.

Last updated: July 23, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.