# VoxelBench Text-Prompt Leaderboard (VoxelBench Text)

> A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Canonical page: https://benchlm.ai/benchmarks/voxelbench

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 29, 2026

## About VoxelBench Text

- Year: 2025
- Tasks: Live text prompts for 3D voxel construction
- Format: Glicko-2 rating from blind pairwise votes
- Difficulty: 3D spatial construction and visual quality
- Paper: [VoxelBench leaderboard](https://voxelbench.ai/leaderboard)

We mirror the official text-prompt API rows that clear VoxelBench's 50-vote display gate. Glicko-2 ratings summarize blind pairwise preferences, while rating deviation and the 95% confidence interval show how uncertain each estimate remains.

VoxelBench Text is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (53 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | 4,578 votes · 96.9% win · 95% CI 2457–2927 | OpenAI | 2692 |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | 1,729 votes · 95.7% win · 95% CI 2486–2646 | Anthropic | 2566 |
| 3 | [GPT-6 Sol](/models/gpt-6-sol) | 444 votes · 88.5% win · 95% CI 2347–2563 | OpenAI | 2455 |
| 4 | [Claude Opus 5](/models/claude-opus-5) | 2,928 votes · 85.2% win · 95% CI 2113–2317 | Anthropic | 2215 |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | 4,479 votes · 87.1% win · 95% CI 2058–2254 | OpenAI | 2156 |
| 6 | [Claude Fable 5](/models/claude-fable) | 4,942 votes · 86.3% win · 95% CI 2022–2222 | Anthropic | 2122 |
| 7 | [Grok 4.6](/models/grok-4-6) | 2,353 votes · 73.9% win · 95% CI 1987–2179 | xAI | 2083 |
| 8 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | 5,590 votes · 83.2% win · 95% CI 1865–2061 | OpenAI | 1963 |
| 9 | [GPT-5.5](/models/gpt-5-5) | 6,298 votes · 82.3% win · 95% CI 1819–2007 | OpenAI | 1913 |
| 10 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | 2,181 votes · 63.7% win · 95% CI 1812–1996 | OpenAI | 1904 |
| 11 | [Kimi K3](/models/kimi-k3) | 3,719 votes · 74.5% win · 95% CI 1787–1983 | Moonshot AI | 1885 |
| 12 | [space-bunny-alpha](https://voxelbench.ai/leaderboard) | 133 votes · 45.1% win · 95% CI 1745–1991 | Unknown | 1868 |
| 13 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | 2,316 votes · 59.4% win · 95% CI 1756–1936 | DeepSeek | 1846 |
| 14 | [Qwen3.8 Max](/models/qwen3-8-max) | 2,489 votes · 58.6% win · 95% CI 1739–1923 | Alibaba | 1831 |
| 15 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | 2,432 votes · 59.2% win · 95% CI 1736–1916 | Google | 1826 |
| 16 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | 2,315 votes · 57.2% win · 95% CI 1734–1914 | OpenAI | 1824 |
| 17 | [GPT-6 Luna](/models/gpt-6-luna) | 458 votes · 53.5% win · 95% CI 1735–1899 | OpenAI | 1817 |
| 18 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | 2,222 votes · 49.4% win · 95% CI 1697–1885 | Google | 1791 |
| 19 | [GLM-5.3](/models/glm-5-3) | 2,473 votes · 52.5% win · 95% CI 1673–1853 | Z.AI | 1763 |
| 20 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | 2,812 votes · 54.6% win · 95% CI 1669–1849 | Google | 1759 |
| 21 | [Grok 4.5](/models/grok-4-5) | 3,388 votes · 56.8% win · 95% CI 1652–1840 | xAI | 1746 |
| 22 | [GLM-5.3-Flash (Max)](https://voxelbench.ai/leaderboard) | 3,061 votes · 54.2% win · 95% CI 1637–1817 | Z.AI | 1727 |
| 23 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | 3,824 votes · 57.0% win · 95% CI 1610–1790 | Google | 1700 |
| 24 | [Gemini 3.1 Pro Preview](https://voxelbench.ai/leaderboard) | 7,151 votes · 75.0% win · 95% CI 1573–1765 | Google | 1669 |
| 25 | [Muse Spark 1.3](/models/muse-spark-1-3) | 1,656 votes · 45.0% win · 95% CI 1563–1763 | Meta | 1663 |
| 26 | [Claude Opus 4.8 (Max)](https://voxelbench.ai/leaderboard) | 3,732 votes · 55.3% win · 95% CI 1547–1731 | Anthropic | 1639 |
| 27 | [Qwen3.7 Max](/models/qwen3-7-max) | 3,465 votes · 50.1% win · 95% CI 1547–1727 | Alibaba | 1637 |
| 28 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | 2,134 votes · 42.2% win · 95% CI 1509–1697 | DeepSeek | 1603 |
| 29 | [GLM-5.2](/models/glm-5-2) | 3,536 votes · 48.3% win · 95% CI 1494–1682 | Z.AI | 1588 |
| 30 | [Gemini 2.5 Deep Think](https://voxelbench.ai/leaderboard) | 6,615 votes · 72.7% win · 95% CI 1427–1627 | Google | 1527 |
| 31 | [Muse Spark 1.1](/models/muse-spark-1-1) | 2,632 votes · 36.7% win · 95% CI 1420–1604 | Meta | 1512 |
| 32 | [Claude Sonnet 5 (xhigh)](https://voxelbench.ai/leaderboard) | 2,766 votes · 44.5% win · 95% CI 1387–1575 | Anthropic | 1481 |
| 33 | [GPT-5.4](/models/gpt-5-4) | 4,690 votes · 54.0% win · 95% CI 1380–1576 | OpenAI | 1478 |
| 34 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | 2,452 votes · 30.1% win · 95% CI 1376–1560 | DeepSeek | 1468 |
| 35 | [Claude Opus 4.6](/models/claude-opus-4-6) | 4,555 votes · 52.8% win · 95% CI 1354–1558 | Anthropic | 1456 |
| 36 | [Claude Opus 4.5](/models/claude-opus-4-5) | 4,113 votes · 52.8% win · 95% CI 1316–1524 | Anthropic | 1420 |
| 37 | [GLM-5.1](/models/glm-5-1) | 4,317 votes · 40.9% win · 95% CI 1323–1511 | Z.AI | 1417 |
| 38 | [GPT-5.2 (xHigh)](https://voxelbench.ai/leaderboard) | 3,795 votes · 54.4% win · 95% CI 1309–1521 | OpenAI | 1415 |
| 39 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | 6,640 votes · 51.7% win · 95% CI 1191–1637 | Google | 1414 |
| 40 | [Gemini 3 Deep Think](https://voxelbench.ai/leaderboard) | 3,736 votes · 51.5% win · 95% CI 1313–1509 | Google | 1411 |
| 41 | [Claude Opus 4.7](/models/claude-opus-4-7) | 3,926 votes · 42.2% win · 95% CI 1311–1503 | Anthropic | 1407 |
| 42 | [MiniMax M3](/models/minimax-m3) | 3,587 votes · 37.0% win · 95% CI 1276–1472 | MiniMax | 1374 |
| 43 | [Gemini 3 Pro Preview](https://voxelbench.ai/leaderboard) | 3,967 votes · 49.1% win · 95% CI 1237–1465 | Google | 1351 |
| 44 | [Gemini 3 Flash Preview](https://voxelbench.ai/leaderboard) | 5,027 votes · 39.9% win · 95% CI 1220–1416 | Google | 1318 |
| 45 | [Qwen 3.6 Plus Preview](https://voxelbench.ai/leaderboard) | 3,604 votes · 34.3% win · 95% CI 1216–1420 | Alibaba | 1318 |
| 46 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | 4,068 votes · 30.6% win · 95% CI 1198–1398 | DeepSeek | 1298 |
| 47 | [GPT-5 (high)](/models/gpt-5-high) | 6,666 votes · 57.7% win · 95% CI 1189–1405 | OpenAI | 1297 |
| 48 | [GPT-5 Pro](https://voxelbench.ai/leaderboard) | 6,147 votes · 55.6% win · 95% CI 1170–1378 | OpenAI | 1274 |
| 49 | [GPT-5 Codex (High)](https://voxelbench.ai/leaderboard) | 6,579 votes · 51.6% win · 95% CI 1126–1342 | OpenAI | 1234 |
| 50 | [GLM-5](/models/glm-5) | 4,315 votes · 35.1% win · 95% CI 1125–1329 | Z.AI | 1227 |
| 51 | [GPT-5.1](/models/gpt-5-1) | 4,234 votes · 37.7% win · 95% CI 1098–1318 | OpenAI | 1208 |
| 52 | [Claude Opus 4.1 (64K Thinking)](https://voxelbench.ai/leaderboard) | 6,681 votes · 49.3% win · 95% CI 1100–1312 | Anthropic | 1206 |
| 53 | [Claude Sonnet 4.5 (32K Thinking)](https://voxelbench.ai/leaderboard) | 7,238 votes · 49.8% win · 95% CI 1086–1290 | Anthropic | 1188 |

## FAQ

### What does VoxelBench Text measure?

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

### Which model leads the published VoxelBench Text snapshot?

GPT-6 Astra currently leads the published VoxelBench Text snapshot with a score of 2692.

### How many models are evaluated on VoxelBench Text?

The September 29, 2026 contains 53 AI models.

### Does VoxelBench Text affect BenchLM's overall score?

Not directly. VoxelBench Text is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
