Skip to main content
BenchLM

Blueprint-Bench 2

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An agentic spatial reasoning benchmark reported as a normalized score.

Normalized connectivity score on Blueprint-Bench 2 — September 23, 2026

We mirror the published normalized connectivity score view for Blueprint-Bench 2. GPT-6 Astra leads the public snapshot at 49.7%, followed by Claude Fable 5.1 (41.9%) and Claude Fable 5 (38.6%). We do not use these results to rank models overall.

26 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 23, 2026

Normalized connectivity score table (26 models)

Score
1
GPT-6 AstraOpenAI · Closed
49.7%
2
Claude Fable 5.1Anthropic · Closed
41.9%
3
Claude Fable 5Anthropic · Closed
38.6%
4
Gemini 3.8 FlashGoogle · Closed
38.6%
5
GPT-5.5OpenAI · Closed
36.2%
6
Gemini 3.5 FlashGoogle · Closed
33.6%
7
GPT-5.6 SolOpenAI · Closed
33.6%
8
Grok 4.6xAI · Closed
33.2%
9
Gemini 3.6 FlashGoogle · Closed
31.2%
10
GPT-5.6 TerraOpenAI · Closed
30.8%
11
Claude Opus 5Anthropic · Closed
30.4%
12
Kimi K3Moonshot AI · Closed
29.5%
13
Grok 4.5xAI · Closed
27.3%
14
GPT-5.4OpenAI · Closed
27.1%
15
Gemini 3.1 ProGoogle · Closed
26.5%
16
Claude Sonnet 5Anthropic · Closed
24.9%
17
Claude Opus 4.7Anthropic · Closed
24.5%
18
GPT-5.6 LunaOpenAI · Closed
22.6%
19
Claude Opus 4.8Anthropic · Closed
14.5%
20
Claude Sonnet 4.6Anthropic · Closed
6.7%
21
Kimi K2.6Moonshot AI · Open weight
3.9%
22
Gemini 3 FlashGoogle · Closed
0.0%
23
Grok 4.3xAI · Closed
0.0%
25
Claude Haiku 4.5Anthropic · Closed
0.0%
26
Grok 4.20xAI · Closed
0.0%

How we show Blueprint-Bench 2

We mirror Andon Labs' Blueprint-Bench 2 model table as a display-only spatial-reasoning benchmark. Each agent processes 50 apartments sequentially, examines roughly 20 interior photographs per apartment, and draws a structured floor plan.

The composite measures room connectivity, graph degree and density, room and door counts, and orientation. Scores are normalized so the random baseline is 0 and a perfect result is 100. The human result covers 12 apartments, so it stays outside the model ranking.

Snapshot

26 model rows50 apartmentsHuman subset: 58.6%Display only

The published Blueprint-Bench 2 snapshot places GPT-6 Astra first at 49.7%. The third row is 11.1 points behind. The broader top-10 range is 18.9 points, so the table still separates the published systems.

26 models have been evaluated on Blueprint-Bench 2. The benchmark falls in the Multimodal & Grounded category. Blueprint-Bench 2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Blueprint-Bench 2

Year

2026

Tasks

Spatial reasoning from blueprints

Format

Normalized score

Difficulty

Agentic spatial reasoning

Google reported Blueprint-Bench 2 in the Gemini 3.5 Flash launch comparison table. BenchLM stores it as a display-only multimodal and spatial-reasoning benchmark until Google publishes the full methodology page.

Freshness and provenance

Version

Blueprint-Bench 2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Blueprint-Bench 2 measure?

An agentic spatial reasoning benchmark reported as a normalized score.

Which model leads the published Blueprint-Bench 2 snapshot?

GPT-6 Astra currently leads the published Blueprint-Bench 2 snapshot with 49.7% normalized connectivity score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Blueprint-Bench 2?

The September 23, 2026 snapshot contains 26 AI models.

Last updated: September 23, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.