# ProteinGym Hard

> Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.

Canonical page: https://benchlm.ai/benchmarks/proteingymhard

- Category: [Knowledge](/knowledge)
- Last updated: September 27, 2026

## About ProteinGym Hard

- Year: 2026
- Tasks: Hard protein mutation-effect ranking tasks
- Format: Rank correlation
- Difficulty: Computational protein science
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.17.3 reports rank-correlation performance on the hard ProteinGym slice with bash and file-editing tools but no package manager.

ProteinGym Hard is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 47.7% |

## FAQ

### What does ProteinGym Hard measure?

Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.

### Which model scores highest on ProteinGym Hard?

Claude Opus 5 by Anthropic currently leads with a score of 47.7% on ProteinGym Hard.

### How many models are evaluated on ProteinGym Hard?

1 AI models have been evaluated on ProteinGym Hard on BenchLM.

### Does ProteinGym Hard affect BenchLM's overall score?

Not directly. ProteinGym Hard is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
