Skip to main content
BenchLM

Vals ProofBench (ProofBench)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI automated theorem-proving benchmark.

ProofBench score on ProofBench — September 28, 2026

We mirror the published proofbench score view for ProofBench. Claude Sonnet 5.5 leads the public snapshot at 100.00%, followed by Claude Opus 5.5 (100.00%) and Claude Fable 5.1 (100.00%). We do not use these results to rank models overall.

44 modelsMathematicsCurrentDisplay onlyUpdated September 28, 2026

ProofBench score table (44 models)

Score
1
Claude Sonnet 5.5Anthropic · Closed
100.00%
2
Claude Opus 5.5Anthropic · Closed
100.00%
3
Claude Fable 5.1Anthropic · Closed
100.00%
4
Alephproverlogicalintelligence
100.00%
5
GPT-6 AstraOpenAI · Closed
99.00%
6
Claude Opus 5Anthropic · Closed
99.00%
7
Claude Fable 5Anthropic · Closed
95.00%
8
Kimi K3Moonshot AI · Closed
87.00%
9
AristotleAristotle
86.00%
10
GPT-6 SolOpenAI · Closed
83.00%
11
GPT-5.6 SolOpenAI · Closed
83.00%
12
Claude Sonnet 5Anthropic · Closed
77.00%
13
Hy4 previewTencent · Open weight
75.00%
14
GPT-5.6 TerraOpenAI · Closed
74.00%
15
MiMo-V2.6-ProXiaomi · Open weight
70.00%
16
GPT-6 LunaOpenAI · Closed
64.00%
17
MiMo-V2.6-FlashXiaomi · Open weight
63.00%
18
GPT-5.6 LunaOpenAI · Closed
60.00%
19
Gemini 3.7 FlashGoogle · Closed
58.00%
20
58.00%
21
Qwen3.8 MaxAlibaba · Open weight
58.00%
22
DeepSeek V4 Flash 0731DeepSeek · Open weight
56.00%
23
Muse Spark 1.3Meta · Closed
55.00%
24
DeepSeek V4.1 FlashDeepSeek · Open weight
54.00%
25
Grok 4.6xAI · Closed
51.00%
26
DeepSeek V4 Pro 0813DeepSeek · Open weight
50.00%
27
GLM-5.3Z.AI · Open weight
49.00%
28
Gemini 3.8 FlashGoogle · Closed
48.00%
29
Muse Spark 1.2Meta · Closed
43.00%
30
Gemini 3.5 FlashGoogle · Closed
31.00%
31
Grok 4.5xAI · Closed
31.00%
32
Grok 4.7xAI · Closed
26.00%
34
MiMo-V2.5-ProXiaomi · Closed
22.00%
35
GLM-5.3-FlashZ.AI · Open weight
21.00%
36
MiniMax M3MiniMax · Open weight
18.00%
37
DeepSeek V4 Pro 0813DeepSeek · Open weight
16.00%
38
MiMo-V2.5Xiaomi · Closed
16.00%
39
Qwen3.8-27BAlibaba · Open weight
16.00%
40
9.00%
41
Inkling-SmallThinking Machines Lab · Open weight
6.00%
42
Mercury 2.5Inception · Closed
3.00%
43
InklingThinking Machines Lab · Open weight
0.00%

How ProofBench is shown here

BenchLM mirrors the public Vals AI ProofBench leaderboard captured from https://www.vals.ai/benchmarks/proof_bench and updated by Vals on September 28, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

ProofBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

44 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

The published ProofBench snapshot places Claude Sonnet 5.5 first at 100.00%. The third row is 0.00 points behind. The broader top-10 range is 17.00 points, so the table still separates the published systems.

44 models have been evaluated on ProofBench. The benchmark falls in the Mathematics category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. ProofBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ProofBench

Year

2026

Tasks

Automated theorem proving

Format

Accuracy score

Difficulty

Formal proof reasoning

BenchLM mirrors Vals ProofBench as a display-only math and proof benchmark.

Freshness and provenance

Version

ProofBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ProofBench measure?

Vals AI automated theorem-proving benchmark.

Which model leads the published ProofBench snapshot?

Claude Sonnet 5.5 currently leads the published ProofBench snapshot with 100.00% proofbench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ProofBench?

The September 28, 2026 snapshot contains 44 AI models.

Last updated: September 28, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.