Skip to main content

Benchmark profile

ProteinGym Hard

Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.

Data verified

Benchmark score on ProteinGym Hard — July 24, 2026

BenchLM mirrors the published score view for ProteinGym Hard. Claude Opus 5 leads the public snapshot at 47.7%. BenchLM does not use these results to rank models overall.

1 modelKnowledgeCurrentDisplay onlyUpdated July 24, 2026

Benchmark score table (1 model)

Score
1
Claude Opus 5Anthropic · Closed
47.7%

About ProteinGym Hard

Year

2026

Tasks

Hard protein mutation-effect ranking tasks

Format

Rank correlation

Difficulty

Computational protein science

Section 8.17.3 reports rank-correlation performance on the hard ProteinGym slice with bash and file-editing tools but no package manager.

BenchLM freshness & provenance

Version

ProteinGym Hard 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ProteinGym Hard measure?

Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.

Which model scores highest on ProteinGym Hard?

Claude Opus 5 by Anthropic currently leads with a score of 47.7% on ProteinGym Hard.

How many models are evaluated on ProteinGym Hard?

1 AI models have been evaluated on ProteinGym Hard on BenchLM.

Last updated: July 24, 2026 · BenchLM version ProteinGym Hard 2026