Skip to main content

Benchmark profile

Graduate-Level Google-Proof Q&A (GPQA)

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

Data verified

Sakana Fugu-Ultra leads the GPQA leaderboard on BenchLM's July 2026 update with 95.5%, ahead of Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%), across 71 tracked models.

Top models on GPQA — July 20, 2026

As of July 20, 2026, Sakana Fugu-Ultra leads the GPQA leaderboard with 95.5% , followed by Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%).

71 modelsKnowledge7% of category scoreRefreshingUpdated July 20, 2026

Leaderboard (71 models)

Score
1
Sakana Fugu-UltraSakana AI · Closed
95.5%
2
Sakana FuguSakana AI · Closed
95.5%
3
GPT-5.6 SolOpenAI · Closed
94.6%
4
Claude Opus 4.7 (Adaptive)Anthropic · Closed
94.2%
5
Claude Mythos 5Anthropic · Closed
94.1%
6
Claude Opus 4.8Anthropic · Closed
93.6%
7
GPT-5.5OpenAI · Closed
93.6%
8
Kimi K3Moonshot AI · Closed
93.5%
9
GPT-5.6 TerraOpenAI · Closed
92.9%
10
GPT-5.4OpenAI · Closed
92.8%
11
Qwen3.7 MaxAlibaba · Closed
92.4%
12
GPT-5.2OpenAI · Closed
92.4%
13
GPT-5.6 LunaOpenAI · Closed
92.3%
14
Gemini 3.5 FlashGoogle · Closed
92.2%
15
Claude Opus 4.6Anthropic · Closed
91.3%
16
GLM-5.2Z.AI · Open weight
91.2%
17
Kimi K2.6Moonshot AI · Open weight
90.5%
18
Qwen3.6 PlusAlibaba · Closed
90.4%
19
Qwen3.7 PlusAlibaba · Closed
90.3%
20
DeepSeek V4 Pro (Max)DeepSeek · Open weight
90.1%
21
Grok 4.3xAI · Closed
90.1%
22
Claude Sonnet 4.6Anthropic · Closed
89.9%
23
Interfaze BetaInterfaze · Closed
89.9%
24
DeepSeek V4 Pro (High)DeepSeek · Open weight
89.1%
25
Qwen3.5 397BAlibaba · Open weight
88.4%
26
DeepSeek V4 Flash (Max)DeepSeek · Open weight
88.1%
27
GPT-5.4 miniOpenAI · Closed
88%
28
InklingThinking Machines Lab · Open weight
87.9%
29
Qwen3.6-27BAlibaba · Open weight
87.8%
30
Kimi K2.5Moonshot AI · Open weight
87.6%
31
Kimi K2.5 (Reasoning)Moonshot AI · Closed
87.6%
32
DeepSeek V4 Flash (High)DeepSeek · Open weight
87.4%
33
Hy3 PreviewTencent · Open weight
87.2%
34
Claude Opus 4.5Anthropic · Closed
87%
35
Nemotron 3 UltraNVIDIA · Open weight
87%
36
Qwen3.5-122B-A10BAlibaba · Open weight
86.6%
37
GLM-5Z.AI · Open weight
86%
38
Qwen3.6-35B-A3BAlibaba · Open weight
86%
39
GLM-4.7Z.AI · Open weight
85.7%
40
Qwen3.5-27BAlibaba · Open weight
85.5%
41
Gemma 4 31BGoogle · Open weight
84.3%
42
MAI-Thinking-1Microsoft · Closed
84.2%
43
Qwen3.5-35B-A3BAlibaba · Open weight
84.2%
44
MiMo-V2-FlashXiaomi · Open weight
83.7%
45
Claude Sonnet 4.5Anthropic · Closed
83.4%
46
Gemini 2.5 ProGoogle · Closed
83%
47
GPT-5.4 nanoOpenAI · Closed
82.8%
48
o1-proOpenAI · Closed
79%
49
Gemma 4 12BGoogle · Open weight
78.8%
50
Qwen3 235B 2507Alibaba · Open weight
77.5%
51
o3-miniOpenAI · Closed
77.2%
52
o1OpenAI · Closed
75.7%
53
DeepSeek V4 ProDeepSeek · Open weight
72.9%
54
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
72.2%
55
DeepSeek V4 FlashDeepSeek · Open weight
71.2%
56
GPT-5 nanoOpenAI · Closed
71.2%
57
ZAYA1-8BZyphra · Open weight
71%
58
GPT-4.1OpenAI · Closed
66.3%
59
GPT-4.1 miniOpenAI · Closed
64.2%
60
Claude 3.5 SonnetAnthropic · Closed
59.4%
61
DeepSeek V3DeepSeek · Open weight
59.1%
62
Ling 2.6 FlashInclusionAI · Open weight
59%
63
Gemma 4 E4BGoogle · Open weight
58.6%
64
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weight
57.6%
65
ZAYA1-74B-PreviewZyphra · Open weight
57.3%
66
GPT-4.1 nanoOpenAI · Closed
50.3%
67
Soofi S 30B-A3BSoofi Project · Open weight
43.4%
68
Gemma 4 E2BGoogle · Open weight
43.4%
69
Mellum2-12B-A2.5B-InstructJetBrains · Open weight
40.9%
70
LFM2.5-VL-450MLiquidAI · Open weight
25.7%
71
LFM2.5-230MLiquidAI · Open weight
25.4%

According to BenchLM.ai, Sakana Fugu-Ultra leads the GPQA benchmark with a score of 95.5%, followed by Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%). The top models are clustered within 0.9 points, suggesting this benchmark is nearing saturation for frontier models.

71 models have been evaluated on GPQA. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, GPQA contributes 7% of the category score, so strong performance here directly affects a model's overall ranking.

About GPQA

Year

2023

Tasks

448 questions

Format

Multiple choice questions

Difficulty

Graduate level

GPQA questions are crafted by PhD-level domain experts and validated to be answerable by experts but challenging for non-experts even with internet access. This makes it an excellent test of deep scientific knowledge and reasoning.

BenchLM freshness & provenance

Version

GPQA Diamond

Refresh cadence

Static

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GPQA measure?

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

Which model scores highest on GPQA?

Sakana Fugu-Ultra by Sakana AI currently leads with a score of 95.5% on GPQA.

How many models are evaluated on GPQA?

71 AI models have been evaluated on GPQA on BenchLM.

Last updated: July 20, 2026 · BenchLM version GPQA Diamond

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.