# VoxelBench Image-Prompt Leaderboard (VoxelBench Image)

> A live human-preference benchmark where multimodal models build voxel structures from image references and voters compare anonymous results produced from the same prompt.

Canonical page: https://benchlm.ai/benchmarks/voxelbench-image

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 29, 2026

## About VoxelBench Image

- Year: 2025
- Tasks: Live image-reference prompts for 3D voxel construction
- Format: Glicko-2 rating from blind pairwise votes
- Difficulty: Visual grounding, 3D construction, and aesthetic quality
- Paper: [VoxelBench leaderboard](https://voxelbench.ai/leaderboard)

The mirrored image-prompt view keeps only models with at least 50 votes, matching the public page. Each row preserves the Glicko-2 rating, deviation, confidence interval, vote count, win rate, and win/loss/tie record.

VoxelBench Image is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (12 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Gemini 3.1 Pro Preview](https://voxelbench.ai/leaderboard) | 2,230 votes · 81.1% win · 95% CI 1538–1746 | Google | 1642 |
| 2 | [GPT-5.4](/models/gpt-5-4) | 2,938 votes · 64.9% win · 95% CI 1404–1576 | OpenAI | 1490 |
| 3 | [Gemini 3 Pro Preview](https://voxelbench.ai/leaderboard) | 5,167 votes · 66.6% win · 95% CI 1388–1544 | Google | 1466 |
| 4 | [Claude Opus 4.6](/models/claude-opus-4-6) | 4,160 votes · 65.6% win · 95% CI 1357–1503 | Anthropic | 1430 |
| 5 | [GPT-5.2 (xHigh)](https://voxelbench.ai/leaderboard) | 5,038 votes · 64.1% win · 95% CI 1330–1478 | OpenAI | 1404 |
| 6 | [Claude Opus 4.5](/models/claude-opus-4-5) | 5,075 votes · 55.5% win · 95% CI 1202–1350 | Anthropic | 1276 |
| 7 | [Claude Opus 4.7](/models/claude-opus-4-7) | 4,441 votes · 49.0% win · 95% CI 1167–1313 | Anthropic | 1240 |
| 8 | [GPT-5 (high)](/models/gpt-5-high) | 4,000 votes · 46.2% win · 95% CI 1143–1303 | OpenAI | 1223 |
| 9 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | 3,804 votes · 43.9% win · 95% CI 1139–1299 | Google | 1219 |
| 10 | [GPT-5 Codex (High)](https://voxelbench.ai/leaderboard) | 4,010 votes · 43.4% win · 95% CI 1122–1278 | OpenAI | 1200 |
| 11 | [Claude Opus 4.1 (64K Thinking)](https://voxelbench.ai/leaderboard) | 4,025 votes · 38.0% win · 95% CI 1069–1225 | Anthropic | 1147 |
| 12 | [GPT-5.1](/models/gpt-5-1) | 5,159 votes · 39.0% win · 95% CI 1071–1219 | OpenAI | 1145 |

## FAQ

### What does VoxelBench Image measure?

A live human-preference benchmark where multimodal models build voxel structures from image references and voters compare anonymous results produced from the same prompt.

### Which model leads the published VoxelBench Image snapshot?

Gemini 3.1 Pro Preview currently leads the published VoxelBench Image snapshot with a score of 1642.

### How many models are evaluated on VoxelBench Image?

The September 29, 2026 contains 12 AI models.

### Does VoxelBench Image affect BenchLM's overall score?

Not directly. VoxelBench Image is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
