Skip to main content
BenchLM

Pencil Puzzle Bench

We show this table for reference; we do not rank on it.

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

Best solve rate on Pencil Puzzle Bench — September 23, 2026 snapshot

We mirror the published best solve rate view for Pencil Puzzle Bench. Claude Fable 5 leads the public snapshot at 97.6%, followed by GPT-5.5 (83.3%) and GPT-5.4 (70.2%). We do not use these results to rank models overall.

73 modelsReasoningCurrentDisplay onlyUpdated September 23, 2026 snapshot

Best solve rate table (73 models)

Score
1
Claude Fable 5Anthropic · Closed
97.6%
2
GPT-5.5OpenAI · Closed
83.3%
3
GPT-5.4OpenAI · Closed
70.2%
4
GPT-5.2OpenAI · Closed
56.0%
5
Claude Opus 4.7Anthropic · Closed
50.0%
6
Gemini 3.5 FlashGoogle · Closed
41.9%
7
Qwen3.7 MaxAlibaba · Closed
40.0%
8
GPT-5.2OpenAI · Closed
36.7%
9
36.7%
10
Claude Opus 4.6 (Adaptive)Anthropic · Closed
33.3%
11
Gemini 3.1 ProGoogle · Closed
33.3%
12
Claude Opus 4.6Anthropic · Closed
30.0%
13
GLM-5.2Z.AI · Open weight
26.7%
14
Claude Sonnet 4.6Anthropic · Closed
26.7%
15
GPT-5.2 ProOpenAI · Closed
26.7%
16
GPT-5.2OpenAI · Closed
23.3%
17
Claude Opus 4.6Anthropic · Closed
23.3%
18
23.3%
19
Kimi K2.6Moonshot AI · Open weight
20.0%
20
Kimi K2.7 CodeMoonshot AI · Open weight
16.7%
21
Qwen3.7 PlusAlibaba · Closed
16.7%
22
Gemini 3 ProGoogle · Closed
16.7%
23
Claude Sonnet 4.6Anthropic · Closed
16.7%
24
Gemini 3 ProGoogle · Closed
13.3%
25
Gemini 3 ProGoogle · Closed
10.0%
26
GPT-5.2OpenAI · Closed
10.0%
27
Qwen3.6 PlusAlibaba · Closed
10.0%
28
GPT-5.1OpenAI · Closed
7.7%
29
MiniMax M3MiniMax · Open weight
7.1%
30
Claude Opus 4.5 ThinkingAnthropic · Closed
6.7%
31
Gemini 3 FlashGoogle · Closed
6.7%
32
Gemini 3 FlashGoogle · Closed
6.7%
33
Grok 4.20xAI · Closed
6.7%
34
GPT-5 (high)OpenAI · Closed
6.0%
35
Kimi K2.5Moonshot AI · Open weight
6.0%
36
Grok 4.1 FastxAI · Closed
5.7%
37
5.3%
38
DeepSeek V4 Pro 0813DeepSeek · Open weight
4.0%
39
Grok 4.3xAI · Closed
3.3%
40
o3OpenAI · Closed
3.3%
42
MiniMax M2.5MiniMax · Closed
3.3%
43
Claude Opus 4.5Anthropic · Closed
3.3%
44
Claude Sonnet 4.5Anthropic · Closed
3.3%
46
Claude Sonnet 4.5 ThinkingAnthropic · Closed
2.3%
47
DeepSeek V3.2DeepSeek · Open weight
2.0%
48
Grok 4.3xAI · Closed
2.0%
49
Kimi K2Moonshot AI · Closed
1.3%
50
MiMo-V2-ProXiaomi · Closed
1.0%
51
o1OpenAI · Closed
0.7%
52
MiniMax M2.7MiniMax · Open weight
0.7%
53
0.7%
54
GLM-5Z.AI · Open weight
0.7%
55
Gemini 2.5 ProGoogle · Closed
0.3%
56
GPT-5.2OpenAI · Closed
0.3%
57
0.3%
58
0.3%
59
GPT-OSS 120BOpenAI · Open weight
0.3%
63
MiMo-V2-FlashXiaomi · Open weight
0.3%
64
GLM-4.7Z.AI · Open weight
0.3%
65
Grok Code Fast 1xAI · Closed
0.3%
66
Gemini 3.5 FlashGoogle · Closed
0.0%
67
0.0%
68
GPT-4.1OpenAI · Closed
0.0%
69
GPT-4oOpenAI · Closed
0.0%
70
0.0%
71
0.0%
72
0.0%
73
0.0%

How Pencil Puzzle Bench is shown here

BenchLM mirrors the public Pencil Puzzle Bench leaderboard from September 23, 2026 snapshot. The source benchmark evaluates 51 frontier models on 300 curated puzzles spanning 20 puzzle types, with direct-ask and agentic solve rates reported separately.

Pencil Puzzle Bench is display only on BenchLM. It is a useful multi-step reasoning reference, but the public table mixes direct prompting and agentic runs and exposes variant-specific reasoning settings, so BenchLM keeps it out of weighted model rankings for now.

Snapshot

73 model variants300 evaluation puzzles20 puzzle types17,000 eval runsDisplay only

The published Pencil Puzzle Bench snapshot places Claude Fable 5 first at 97.6%. The third row is 27.4 points behind. The broader top-10 range is 64.3 points, so the table still separates the published systems.

73 models have been evaluated on Pencil Puzzle Bench. The benchmark falls in the Reasoning category. Pencil Puzzle Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Pencil Puzzle Bench

Year

2026

Tasks

300 evaluation puzzles

Format

Direct and agentic puzzle solve rate

Difficulty

Multi-step verifiable reasoning

BenchLM mirrors the public Pencil Puzzle Bench leaderboard as a display-only reasoning benchmark. The public site reports direct-ask and agentic solve rates across a 300-puzzle evaluation selection from the 62,231-puzzle dataset.

Freshness and provenance

Version

Pencil Puzzle Bench 2026

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Pencil Puzzle Bench measure?

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

Which model leads the published Pencil Puzzle Bench snapshot?

Claude Fable 5 currently leads the published Pencil Puzzle Bench snapshot with 97.6% best solve rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Pencil Puzzle Bench?

The September 23, 2026 snapshot snapshot contains 73 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.