Benchmark profile
ProteinGym Hard
Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.
Data verifiedBenchmark score on ProteinGym Hard — July 24, 2026
BenchLM mirrors the published score view for ProteinGym Hard. Claude Opus 5 leads the public snapshot at 47.7%. BenchLM does not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout ProteinGym Hard
Year
2026
Tasks
Hard protein mutation-effect ranking tasks
Format
Rank correlation
Difficulty
Computational protein science
Section 8.17.3 reports rank-correlation performance on the hard ProteinGym slice with bash and file-editing tools but no package manager.
BenchLM freshness & provenance
Version
ProteinGym Hard 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does ProteinGym Hard measure?
Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.
Which model scores highest on ProteinGym Hard?
Claude Opus 5 by Anthropic currently leads with a score of 47.7% on ProteinGym Hard.
How many models are evaluated on ProteinGym Hard?
1 AI models have been evaluated on ProteinGym Hard on BenchLM.