Skip to main content
BenchLM

VoxelBench Text-Prompt Leaderboard (VoxelBench Text)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Glicko-2 rating on VoxelBench text-prompt leaderboard — September 29, 2026

We mirror the published glicko-2 rating view for VoxelBench text-prompt leaderboard. GPT-6 Astra leads the public snapshot at 2692, followed by Claude Opus 5.5 (2566) and GPT-6 Sol (2455). We do not use these results to rank models overall.

53 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 29, 2026

Glicko-2 rating table (53 models)

Score
1
GPT-6 AstraOpenAI · Closed4,578 votes · 96.9% win · 95% CI 2457–2927
2692
2
Claude Opus 5.5Anthropic · Closed1,729 votes · 95.7% win · 95% CI 2486–2646
2566
3
GPT-6 SolOpenAI · Closed444 votes · 88.5% win · 95% CI 2347–2563
2455
4
Claude Opus 5Anthropic · Closed2,928 votes · 85.2% win · 95% CI 2113–2317
2215
5
GPT-5.6 SolOpenAI · Closed4,479 votes · 87.1% win · 95% CI 2058–2254
2156
6
Claude Fable 5Anthropic · Closed4,942 votes · 86.3% win · 95% CI 2022–2222
2122
7
Grok 4.6xAI · Closed2,353 votes · 73.9% win · 95% CI 1987–2179
2083
8
GPT-5.5 ProOpenAI · Closed5,590 votes · 83.2% win · 95% CI 1865–2061
1963
9
GPT-5.5OpenAI · Closed6,298 votes · 82.3% win · 95% CI 1819–2007
1913
10
GPT-5.6 TerraOpenAI · Closed2,181 votes · 63.7% win · 95% CI 1812–1996
1904
11
Kimi K3Moonshot AI · Closed3,719 votes · 74.5% win · 95% CI 1787–1983
1885
12
space-bunny-alphaUnknown133 votes · 45.1% win · 95% CI 1745–1991
1868
13
DeepSeek V4.1 FlashDeepSeek · Open weight2,316 votes · 59.4% win · 95% CI 1756–1936
1846
14
Qwen3.8 MaxAlibaba · Open weight2,489 votes · 58.6% win · 95% CI 1739–1923
1831
15
Gemini 3.7 FlashGoogle · Closed2,432 votes · 59.2% win · 95% CI 1736–1916
1826
16
GPT-5.6 LunaOpenAI · Closed2,315 votes · 57.2% win · 95% CI 1734–1914
1824
17
GPT-6 LunaOpenAI · Closed458 votes · 53.5% win · 95% CI 1735–1899
1817
18
Gemini 3.8 FlashGoogle · Closed2,222 votes · 49.4% win · 95% CI 1697–1885
1791
19
GLM-5.3Z.AI · Open weight2,473 votes · 52.5% win · 95% CI 1673–1853
1763
20
Gemini 3.6 FlashGoogle · Closed2,812 votes · 54.6% win · 95% CI 1669–1849
1759
21
Grok 4.5xAI · Closed3,388 votes · 56.8% win · 95% CI 1652–1840
1746
22
GLM-5.3-Flash (Max)Z.AI3,061 votes · 54.2% win · 95% CI 1637–1817
1727
23
Gemini 3.5 FlashGoogle · Closed3,824 votes · 57.0% win · 95% CI 1610–1790
1700
24
Gemini 3.1 Pro PreviewGoogle7,151 votes · 75.0% win · 95% CI 1573–1765
1669
25
Muse Spark 1.3Meta · Closed1,656 votes · 45.0% win · 95% CI 1563–1763
1663
26
Claude Opus 4.8 (Max)Anthropic3,732 votes · 55.3% win · 95% CI 1547–1731
1639
27
Qwen3.7 MaxAlibaba · Closed3,465 votes · 50.1% win · 95% CI 1547–1727
1637
28
DeepSeek V4 Pro 0813DeepSeek · Open weight2,134 votes · 42.2% win · 95% CI 1509–1697
1603
29
GLM-5.2Z.AI · Open weight3,536 votes · 48.3% win · 95% CI 1494–1682
1588
30
Gemini 2.5 Deep ThinkGoogle6,615 votes · 72.7% win · 95% CI 1427–1627
1527
31
Muse Spark 1.1Meta · Closed2,632 votes · 36.7% win · 95% CI 1420–1604
1512
32
Claude Sonnet 5 (xhigh)Anthropic2,766 votes · 44.5% win · 95% CI 1387–1575
1481
33
GPT-5.4OpenAI · Closed4,690 votes · 54.0% win · 95% CI 1380–1576
1478
34
DeepSeek V4 Flash 0731DeepSeek · Open weight2,452 votes · 30.1% win · 95% CI 1376–1560
1468
35
Claude Opus 4.6Anthropic · Closed4,555 votes · 52.8% win · 95% CI 1354–1558
1456
36
Claude Opus 4.5Anthropic · Closed4,113 votes · 52.8% win · 95% CI 1316–1524
1420
37
GLM-5.1Z.AI · Open weight4,317 votes · 40.9% win · 95% CI 1323–1511
1417
38
GPT-5.2 (xHigh)OpenAI3,795 votes · 54.4% win · 95% CI 1309–1521
1415
39
Gemini 2.5 ProGoogle · Closed6,640 votes · 51.7% win · 95% CI 1191–1637
1414
40
Gemini 3 Deep ThinkGoogle3,736 votes · 51.5% win · 95% CI 1313–1509
1411
41
Claude Opus 4.7Anthropic · Closed3,926 votes · 42.2% win · 95% CI 1311–1503
1407
42
MiniMax M3MiniMax · Open weight3,587 votes · 37.0% win · 95% CI 1276–1472
1374
43
Gemini 3 Pro PreviewGoogle3,967 votes · 49.1% win · 95% CI 1237–1465
1351
44
Gemini 3 Flash PreviewGoogle5,027 votes · 39.9% win · 95% CI 1220–1416
1318
45
Qwen 3.6 Plus PreviewAlibaba3,604 votes · 34.3% win · 95% CI 1216–1420
1318
46
DeepSeek V4 Pro 0813DeepSeek · Open weight4,068 votes · 30.6% win · 95% CI 1198–1398
1298
47
GPT-5 (high)OpenAI · Closed6,666 votes · 57.7% win · 95% CI 1189–1405
1297
48
GPT-5 ProOpenAI6,147 votes · 55.6% win · 95% CI 1170–1378
1274
49
GPT-5 Codex (High)OpenAI6,579 votes · 51.6% win · 95% CI 1126–1342
1234
50
GLM-5Z.AI · Open weight4,315 votes · 35.1% win · 95% CI 1125–1329
1227
51
GPT-5.1OpenAI · Closed4,234 votes · 37.7% win · 95% CI 1098–1318
1208
52
Claude Opus 4.1 (64K Thinking)Anthropic6,681 votes · 49.3% win · 95% CI 1100–1312
1206
53
Claude Sonnet 4.5 (32K Thinking)Anthropic7,238 votes · 49.8% win · 95% CI 1086–1290
1188

How to read this leaderboard

A higher rating means a model's text-prompt builds won more of their pairwise comparisons after Glicko-2 adjusted for opponent strength. Read the rating together with its confidence interval and vote count.

Operator receipt: 53 sourced rows are currently displayable on this page; the leading published row is GPT-6 Astra at 2692.

Honest limit: The score belongs to the full VoxelBench setup, including its generation tool, instructions, reasoning setting, prompt mix, renderer, and voter pool. It is not a controlled base-model spatial-reasoning score.

How we show VoxelBench text-prompt ratings

We mirror the text-prompt table from VoxelBench's public leaderboard API. The source page shows models after at least 50 votes and ranks the eligible rows with Glicko-2; the snapshot keeps each rating, deviation, 95% confidence interval, vote count, win rate, and win/loss/tie record.

Voters compare two anonymous voxel structures produced for the same text prompt. The result mixes model behavior with VoxelBench's generation tools, instructions, settings, prompt mix, renderer, and voter pool, so we keep it display only and separate from weighted model rankings.

Snapshot

53 eligible model rows50+ votes per modelText-prompt buildsBlind pairwise votesGlicko-2Display only

The published VoxelBench Text snapshot places GPT-6 Astra first at 2692. The third row is 237 score units behind. The broader top-10 range is 788 score units, so the table still separates the published systems.

53 models have been evaluated on VoxelBench Text. The benchmark falls in the Multimodal & Grounded category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. VoxelBench Text is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About VoxelBench Text

Year

2025

Tasks

Live text prompts for 3D voxel construction

Format

Glicko-2 rating from blind pairwise votes

Difficulty

3D spatial construction and visual quality

We mirror the official text-prompt API rows that clear VoxelBench's 50-vote display gate. Glicko-2 ratings summarize blind pairwise preferences, while rating deviation and the 95% confidence interval show how uncertain each estimate remains.

Freshness and provenance

Version

Live VoxelBench Glicko-2

Refresh cadence

Rolling

Staleness state

Current

Question availability

Prompts browsable; full task set not versioned

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does VoxelBench Text measure?

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Which model leads the published VoxelBench Text snapshot?

GPT-6 Astra currently leads the published VoxelBench Text snapshot with 2692 glicko-2 rating. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on VoxelBench Text?

The September 29, 2026 snapshot contains 53 AI models.

Last updated: September 29, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.