Skip to main content

Benchmark profile

VoxelBench Text-Prompt Leaderboard (VoxelBench Text)

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Data verified

How to read this leaderboard

A higher rating means a model's text-prompt builds won more of their pairwise comparisons after Glicko-2 adjusted for opponent strength. Read the rating together with its confidence interval and vote count.

Operator receipt: 49 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 (Max) at 2236.

Honest limit: The score belongs to the full VoxelBench setup, including its generation tool, instructions, reasoning setting, prompt mix, renderer, and voter pool. It is not a controlled base-model spatial-reasoning score.

How we show VoxelBench text-prompt ratings

We mirror the text-prompt table from VoxelBench's public leaderboard API. The source page shows models after at least 50 votes and ranks the eligible rows with Glicko-2; the snapshot keeps each rating, deviation, 95% confidence interval, vote count, win rate, and win/loss/tie record.

Voters compare two anonymous voxel structures produced for the same text prompt. The result mixes model behavior with VoxelBench's generation tools, instructions, settings, prompt mix, renderer, and voter pool, so we keep it display only and separate from weighted model rankings.

49 eligible model rows50+ votes per modelText-prompt buildsBlind pairwise votesGlicko-2Display only

Glicko-2 rating on VoxelBench text-prompt leaderboard — July 29, 2026

BenchLM mirrors the published glicko-2 rating view for VoxelBench text-prompt leaderboard. Claude Opus 5 (Max) leads the public snapshot at 2236 , followed by GPT-5.6 Sol (Max) (2235) and Claude Fable 5 (Max) (2191). BenchLM does not use these results to rank models overall.

49 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 29, 2026

Glicko-2 rating table (49 models)

Score
1
Claude Opus 5 (Max)Anthropic859 votes · 90.3% win · 95% CI 2156–2316
2236
2
GPT-5.6 Sol (Max)OpenAI2,516 votes · 93.0% win · 95% CI 2143–2327
2235
3
Claude Fable 5 (Max)Anthropic3,499 votes · 91.4% win · 95% CI 2097–2285
2191
4
Kimi K3 (Max)Moonshot AI1,701 votes · 84.5% win · 95% CI 1952–2094
2023
5
GPT-5.5 ProOpenAI4,047 votes · 89.2% win · 95% CI 1933–2113
2023
6
GPT-5.5 (xHigh)OpenAI4,785 votes · 87.6% win · 95% CI 1900–2080
1990
7
Grok 4.5xAI1,517 votes · 63.9% win · 95% CI 1725–1881
1803
8
Gemini 3.6 FlashGoogle727 votes · 65.1% win · 95% CI 1696–1844
1770
9
Claude Opus 4.8 (Max)Anthropic2,047 votes · 66.3% win · 95% CI 1619–1783
1701
10
Gemini 3.5 FlashGoogle2,074 votes · 66.7% win · 95% CI 1602–1770
1686
11
Gemini 3.1 Pro PreviewGoogle5,995 votes · 80.7% win · 95% CI 1583–1759
1671
12
Qwen 3.7 MaxAlibaba1,727 votes · 59.2% win · 95% CI 1546–1710
1628
13
Claude Sonnet 5 (xhigh)Anthropic1,128 votes · 55.0% win · 95% CI 1517–1689
1603
14
GLM-5.2 (Max)Z.AI1,773 votes · 58.3% win · 95% CI 1502–1666
1584
15
Muse Spark 1.1Meta622 votes · 43.7% win · 95% CI 1436–1592
1514
16
Claude Opus 4.5 (High)Anthropic3,164 votes · 60.4% win · 95% CI 1410–1594
1502
17
Gemini 2.5 Deep ThinkGoogle5,794 votes · 77.9% win · 95% CI 1406–1590
1498
18
GPT-5.4 (xHigh)OpenAI3,391 votes · 62.2% win · 95% CI 1409–1581
1495
19
Claude Opus 4.6 (Max)Anthropic3,433 votes · 61.7% win · 95% CI 1388–1568
1478
20
Gemini 3 Deep ThinkGoogle2,650 votes · 59.9% win · 95% CI 1367–1551
1459
21
GPT-5.2 (xHigh)OpenAI2,913 votes · 61.8% win · 95% CI 1352–1540
1446
22
Claude Opus 4.7 (Max)Anthropic2,622 votes · 50.6% win · 95% CI 1331–1507
1419
23
GLM-5.1Z.AI2,754 votes · 48.9% win · 95% CI 1334–1502
1418
24
MiniMax-M3MiniMax2,019 votes · 46.4% win · 95% CI 1325–1493
1409
25
Gemini 3 Pro PreviewGoogle3,062 votes · 57.2% win · 95% CI 1291–1487
1389
26
Gemini 3 Flash PreviewGoogle3,639 votes · 46.8% win · 95% CI 1265–1433
1349
27
Qwen 3.6 Plus PreviewAlibaba2,226 votes · 41.8% win · 95% CI 1234–1410
1322
28
GPT-5 (High)OpenAI5,775 votes · 63.6% win · 95% CI 1198–1382
1290
29
Gemini 2.5 ProGoogle5,790 votes · 56.3% win · 95% CI 1170–1366
1268
30
GLM-5Z.AI2,987 votes · 42.6% win · 95% CI 1176–1356
1266
31
DeepSeek V4 ProDeepSeek2,518 votes · 37.2% win · 95% CI 1169–1341
1255
32
GPT-5 Codex (High)OpenAI5,659 votes · 57.5% win · 95% CI 1157–1349
1253
33
GPT-5.4 Mini (xHigh)OpenAI2,708 votes · 34.9% win · 95% CI 1159–1327
1243
34
Kimi-K2.6Moonshot AI2,242 votes · 34.8% win · 95% CI 1151–1323
1237
35
Claude Opus 4.1 (64K Thinking)Anthropic5,771 votes · 54.6% win · 95% CI 1145–1325
1235
36
Claude Sonnet 4.5 (32K Thinking)Anthropic6,145 votes · 55.7% win · 95% CI 1137–1309
1223
37
GPT-5.1 (High)OpenAI3,214 votes · 44.4% win · 95% CI 1126–1318
1222
38
GPT-5 ProOpenAI5,205 votes · 61.9% win · 95% CI 1124–1308
1216
39
Claude Sonnet 4 (64k)Anthropic6,021 votes · 47.0% win · 95% CI 1065–1245
1155
40
DeepSeek V4 FlashDeepSeek1,940 votes · 23.6% win · 95% CI 1060–1232
1146
41
MiniMax M2.7MiniMax1,924 votes · 24.0% win · 95% CI 1027–1211
1119
42
DeepSeek V3.2 SpecialeDeepSeek3,257 votes · 32.3% win · 95% CI 1024–1212
1118
43
Kimi K2.5 (Thinking)Moonshot AI3,065 votes · 39.1% win · 95% CI 993–1181
1087
44
O3OpenAI5,584 votes · 38.6% win · 95% CI 938–1150
1044
45
Grok 4.20 Beta (03/26)xAI2,352 votes · 23.9% win · 95% CI 940–1124
1032
46
GPT-5.4 Nano (xHigh)OpenAI2,139 votes · 22.7% win · 95% CI 937–1125
1031
47
grok-4.3xAI105 votes · 17.1% win · 95% CI 835–1219
1027
48
Claude Haiku 4.5 (64K Thinking)Anthropic4,955 votes · 29.5% win · 95% CI 854–1062
958
49
Gemini 3.1 Flash Lite PreviewGoogle2,858 votes · 14.4% win · 95% CI 824–1028
926

The published VoxelBench Text snapshot places Claude Opus 5 (Max) first at 2236. The third row is 45 score units behind. The broader top-10 range is 550 score units, so the table still separates the published systems.

49 models have been evaluated on VoxelBench Text. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. VoxelBench Text is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About VoxelBench Text

Year

2025

Tasks

Live text prompts for 3D voxel construction

Format

Glicko-2 rating from blind pairwise votes

Difficulty

3D spatial construction and visual quality

We mirror the official text-prompt API rows that clear VoxelBench's 50-vote display gate. Glicko-2 ratings summarize blind pairwise preferences, while rating deviation and the 95% confidence interval show how uncertain each estimate remains.

BenchLM freshness & provenance

Version

Live VoxelBench Glicko-2

Refresh cadence

Rolling

Staleness state

Current

Question availability

Prompts browsable; full task set not versioned

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does VoxelBench Text measure?

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Which model leads the published VoxelBench Text snapshot?

Claude Opus 5 (Max) currently leads the published VoxelBench Text snapshot with 2236 glicko-2 rating. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on VoxelBench Text?

49 AI models are included in BenchLM's mirrored VoxelBench Text snapshot, based on the public leaderboard captured on July 29, 2026.

Last updated: July 29, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.