# Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set (LAB criterion-pass (Harvey held-out))

> Harvey AI's mean criterion-level score on its held-out legal-agent evaluation.

Canonical page: https://benchlm.ai/benchmarks/legalagentbenchheldoutcriterionpass

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About LAB criterion-pass (Harvey held-out)

- Year: 2026
- Tasks: Harvey-held-out legal-agent tasks
- Format: Mean criterion-pass rate
- Difficulty: Professional legal work
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.13.3 separately reports Harvey's held-out criterion-pass result. It is not merged with Anthropic's internal-harness result.

LAB criterion-pass (Harvey held-out) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 94.1% |
| 2 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 91.2% |

## FAQ

### What does LAB criterion-pass (Harvey held-out) measure?

Harvey AI's mean criterion-level score on its held-out legal-agent evaluation.

### Which model scores highest on LAB criterion-pass (Harvey held-out)?

Claude Opus 5 by Anthropic currently leads with a score of 94.1% on LAB criterion-pass (Harvey held-out).

### How many models are evaluated on LAB criterion-pass (Harvey held-out)?

2 AI models have been evaluated on LAB criterion-pass (Harvey held-out) on BenchLM.

### Does LAB criterion-pass (Harvey held-out) affect BenchLM's overall score?

Not directly. LAB criterion-pass (Harvey held-out) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on LAB criterion-pass (Harvey held-out)

- [Claude Opus 5 vs Claude Opus 5.5](/compare/claude-opus-5-vs-claude-opus-5-5)
