Skip to main content
BenchLM

Design Arena Website Elo (Design Arena Website)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Benchmark score on Design Arena Website — September 27, 2026

We compile the Design Arena Website rows from secondary reports. Muse Spark 1.3 leads the table at 1365, followed by Kimi K3 (1351) and MiMo-V2.6-Pro (1331). We do not use these results to rank models overall.

97 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (97 models)

Score
1
Muse Spark 1.3Meta · Closed
1365
2
Kimi K3Moonshot AI · Closed
1351
3
MiMo-V2.6-ProXiaomi · Open weight
1331
4
Claude Fable 5.1Anthropic · Closed
1320
5
Claude Opus 5Anthropic · Closed
1319
6
Muse Spark 1.2Meta · Closed
1318
7
Gemini 3.7 FlashGoogle · Closed
1313
8
GLM-5.3Z.AI · Open weight
1312
9
Gemini 3.8 FlashGoogle · Closed
1311
10
Gemini 3.6 FlashGoogle · Closed
1310
11
Claude Fable 5Anthropic · Closed
1307
12
GLM-5.2Z.AI · Open weight
1302
13
Claude Opus 4.6Anthropic · Closed
1302
14
Grok 4.6xAI · Closed
1302
15
Claude Opus 4.6 (Adaptive)Anthropic · Closed
1302
16
Claude Sonnet 4.6Anthropic · Closed
1295
17
Grok 4.5xAI · Closed
1294
18
GLM-5.1Z.AI · Open weight
1290
19
Claude Sonnet 5Anthropic · Closed
1286
20
Qwen3.7 MaxAlibaba · Closed
1283
21
GLM-5.3-FlashZ.AI · Open weight
1282
22
Qwen3.7 PlusAlibaba · Closed
1281
23
Kimi K2.6Moonshot AI · Open weight
1281
24
MiMo-V2.5-ProXiaomi · Closed
1281
25
Muse Spark 1.1Meta · Closed
1280
26
Kimi K2.7 CodeMoonshot AI · Open weight
1277
27
MiMo-V2.5Xiaomi · Closed
1276
28
Gemini 3.5 FlashGoogle · Closed
1275
29
GLM-5-TurboZ.AI · Closed
1274
30
MiniMax M3MiniMax · Open weight
1269
31
GPT-5.5OpenAI · Closed
1267
32
Claude Opus 4.8Anthropic · Closed
1266
33
Gemini 3.1 ProGoogle · Closed
1264
34
Kimi K2.5Moonshot AI · Open weight
1260
35
Kimi K2.5 (Reasoning)Moonshot AI · Closed
1260
36
GLM-5Z.AI · Open weight
1259
37
GLM-5 (Reasoning)Z.AI · Open weight
1259
38
DeepSeek V4 Pro 0813DeepSeek · Open weight
1258
39
Claude Opus 4.5Anthropic · Closed
1258
40
Claude Opus 4.5 ThinkingAnthropic · Closed
1258
41
DeepSeek V4 Pro (High)DeepSeek · Open weight
1258
42
MiniMax M2.7MiniMax · Open weight
1254
43
Qwen3.6 PlusAlibaba · Closed
1251
44
DeepSeek V4 FlashDeepSeek · Open weight
1249
45
GLM-5V-TurboZ.AI · Closed
1241
46
Grok 4.20xAI · Closed
1240
47
GLM-4.7Z.AI · Open weight
1236
48
MiniMax M2.5MiniMax · Closed
1233
49
GPT-5.4OpenAI · Closed
1231
50
InklingThinking Machines Lab · Open weight
1229
51
DeepSeek V4 Flash 0731DeepSeek · Open weight
1219
52
DeepSeek V4 Flash (High)DeepSeek · Open weight
1219
53
Gemini 3 FlashGoogle · Closed
1207
54
GPT-5.2OpenAI · Closed
1206
55
Grok 4.3xAI · Closed
1206
56
Step 3.7 FlashStepFun · Open weight
1206
57
GLM-4.7-FlashZ.AI · Open weight
1205
58
Claude Sonnet 4.5Anthropic · Closed
1201
59
Claude Sonnet 4.5 ThinkingAnthropic · Closed
1201
60
GPT-5.1OpenAI · Closed
1198
61
GPT-5 (medium)OpenAI · Closed
1196
62
GPT-5 (high)OpenAI · Closed
1196
63
Hy3Tencent · Open weight
1193
64
Claude 4.1 OpusAnthropic · Closed
1188
65
Solar Pro 4Upstage · Closed
1187
66
DeepSeek V3.2DeepSeek · Open weight
1186
67
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
1186
68
GLM-4.5Z.AI · Closed
1181
69
Gemini 2.5 ProGoogle · Closed
1178
70
GPT-5.3 CodexOpenAI · Closed
1173
71
GPT-5.1-CodexOpenAI · Closed
1173
72
GLM-4.5-AirZ.AI · Closed
1158
73
Claude 4 SonnetAnthropic · Closed
1157
74
Nemotron 3 UltraNVIDIA · Open weight
1149
75
Trinity-Large-ThinkingArcee AI · Open weight
1148
76
Trinity-Large-PreviewArcee AI · Open weight
1148
77
GPT-5 miniOpenAI · Closed
1136
78
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
1134
79
DeepSeek V3.1DeepSeek · Open weight
1134
80
Claude Haiku 4.5Anthropic · Closed
1134
81
Claude Haiku 4.5 ThinkingAnthropic · Closed
1134
82
DeepSeek V3DeepSeek · Open weight
1131
83
Qwen3 MaxAlibaba · Closed
1130
84
Gemini 2.5 FlashGoogle · Closed
1126
85
GPT-5 nanoOpenAI · Closed
1113
86
Mistral Medium 3Mistral · Closed
1090
87
Kimi K2Moonshot AI · Closed
1062
88
GPT-4.1OpenAI · Closed
1050
89
o3OpenAI · Closed
1047
90
Mercury 2Inception · Closed
1038
91
GPT-4.1 miniOpenAI · Closed
1009
92
GPT-4.1 nanoOpenAI · Closed
984
93
GPT-OSS 120BOpenAI · Open weight
979
94
Llama 4 MaverickMeta · Open weight
882
95
GPT-OSS 20BOpenAI · Open weight
864
96
GPT-4oOpenAI · Closed
842
97
Llama 4 ScoutMeta · Open weight
761

Among the reported Design Arena Website rows, Muse Spark 1.3 is first at 1365. The third row is 34 score units behind. The broader top-10 range is 55 score units, so the table still separates the published systems.

97 models have been evaluated on Design Arena Website. The benchmark falls in the Multimodal & Grounded category. Design Arena Website is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Design Arena Website

Year

2026

Tasks

Website generation comparisons

Format

Elo

Difficulty

Design and website generation

OpenRouter's Grok 4.3 benchmark page reports Website at 1294 Elo, 56.5% win rate, 166.3s average generation time, and Top 13%. BenchLM stores the Elo as the benchmark value and keeps the supporting fields in the description.

Freshness and provenance

Version

Design Arena Website 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Design Arena Website measure?

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Which model scores highest on Design Arena Website?

Muse Spark 1.3 by Meta currently leads with a score of 1365 on Design Arena Website.

How many models are evaluated on Design Arena Website?

97 AI models have been evaluated on Design Arena Website on BenchLM.

Last updated: September 27, 2026 · BenchLM version Design Arena Website 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.