# Blueprint-Bench 2

> An agentic spatial reasoning benchmark reported as a normalized score.

Canonical page: https://benchlm.ai/benchmarks/blueprintbench2

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 23, 2026

## About Blueprint-Bench 2

- Year: 2026
- Tasks: Spatial reasoning from blueprints
- Format: Normalized score
- Difficulty: Agentic spatial reasoning
- Paper: [Gemini 3.5 Flash launch screenshots](https://x.com/GoogleDeepMind)

Google reported Blueprint-Bench 2 in the Gemini 3.5 Flash launch comparison table. BenchLM stores it as a display-only multimodal and spatial-reasoning benchmark until Google publishes the full methodology page.

Blueprint-Bench 2 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (26 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 49.7% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 41.9% |
| 3 | [Claude Fable 5](/models/claude-fable) | Anthropic | 38.6% |
| 4 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 38.6% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 36.2% |
| 6 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 33.6% |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 33.6% |
| 8 | [Grok 4.6](/models/grok-4-6) | xAI | 33.2% |
| 9 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 31.2% |
| 10 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 30.8% |
| 11 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 30.4% |
| 12 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 29.5% |
| 13 | [Grok 4.5](/models/grok-4-5) | xAI | 27.3% |
| 14 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 27.1% |
| 15 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 26.5% |
| 16 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 24.9% |
| 17 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 24.5% |
| 18 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 22.6% |
| 19 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 14.5% |
| 20 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 6.7% |
| 21 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 3.9% |
| 22 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 0.0% |
| 23 | [Grok 4.3](/models/grok-4-3) | xAI | 0.0% |
| 24 | [Gemini Robotics-ER 1.6](https://andonlabs.com/evals/blueprint-bench-2) | Google | 0.0% |
| 25 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 0.0% |
| 26 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 0.0% |

## FAQ

### What does Blueprint-Bench 2 measure?

An agentic spatial reasoning benchmark reported as a normalized score.

### Which model leads the published Blueprint-Bench 2 snapshot?

GPT-6 Astra currently leads the published Blueprint-Bench 2 snapshot with a score of 49.7%.

### How many models are evaluated on Blueprint-Bench 2?

The September 23, 2026 contains 26 AI models.

### Does Blueprint-Bench 2 affect BenchLM's overall score?

Not directly. Blueprint-Bench 2 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
