Benchmark profile
Graduate-Level Google-Proof Q&A (GPQA)
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
Data verifiedSakana Fugu-Ultra leads the GPQA leaderboard on BenchLM's July 2026 update with 95.5%, ahead of Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%), across 71 tracked models.
Top models on GPQA — July 20, 2026
As of July 20, 2026, Sakana Fugu-Ultra leads the GPQA leaderboard with 95.5% , followed by Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%).
Sakana Fugu-Ultra
Sakana AI
sakana-fugu-ultra
Sakana Fugu
Sakana AI
sakana-fugu
GPT-5.6 Sol
OpenAI
gpt-5-6-sol
Leaderboard (71 models)
ScoreAccording to BenchLM.ai, Sakana Fugu-Ultra leads the GPQA benchmark with a score of 95.5%, followed by Sakana Fugu (95.5%) and GPT-5.6 Sol (94.6%). The top models are clustered within 0.9 points, suggesting this benchmark is nearing saturation for frontier models.
71 models have been evaluated on GPQA. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, GPQA contributes 7% of the category score, so strong performance here directly affects a model's overall ranking.
About GPQA
Year
2023
Tasks
448 questions
Format
Multiple choice questions
Difficulty
Graduate level
GPQA questions are crafted by PhD-level domain experts and validated to be answerable by experts but challenging for non-experts even with internet access. This makes it an excellent test of deep scientific knowledge and reasoning.
BenchLM freshness & provenance
Version
GPQA Diamond
Refresh cadence
Static
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does GPQA measure?
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
Which model scores highest on GPQA?
Sakana Fugu-Ultra by Sakana AI currently leads with a score of 95.5% on GPQA.
How many models are evaluated on GPQA?
71 AI models have been evaluated on GPQA on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.