ProteinGym Hard
We show this table for reference; we do not rank on it.
Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.
Benchmark score on ProteinGym Hard — September 27, 2026
We compile the ProteinGym Hard rows from provider self-reports. Claude Opus 5 leads the table at 47.7%. We do not use these results to rank models overall.
1 modelKnowledgeCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (1 model)
ScoreAbout ProteinGym Hard
Year
2026
Tasks
Hard protein mutation-effect ranking tasks
Format
Rank correlation
Difficulty
Computational protein science
Section 8.17.3 reports rank-correlation performance on the hard ProteinGym slice with bash and file-editing tools but no package manager.
Freshness and provenance
Version
ProteinGym Hard 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ProteinGym Hard measure?
Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.
Which model scores highest on ProteinGym Hard?
Claude Opus 5 by Anthropic currently leads with a score of 47.7% on ProteinGym Hard.
How many models are evaluated on ProteinGym Hard?
1 AI models have been evaluated on ProteinGym Hard on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.