# Legal Agent Benchmark all-pass rate — Anthropic harness (LAB all-pass (Anthropic harness))

> Strict task success requiring every expert-written legal-work rubric criterion to pass.

Canonical page: https://benchlm.ai/benchmarks/legalagentbenchallpass

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About LAB all-pass (Anthropic harness)

- Year: 2026
- Tasks: 1,235 legal-agent tasks
- Format: All-criteria pass rate
- Difficulty: Professional legal work
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.13.3 reports five-run performance on 1,235 tasks using Anthropic's reduced-tool reimplementation, production safeguards, and fallback to Opus 4.8 on classifier refusal.

LAB all-pass (Anthropic harness) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 23.58% |

## FAQ

### What does LAB all-pass (Anthropic harness) measure?

Strict task success requiring every expert-written legal-work rubric criterion to pass.

### Which model scores highest on LAB all-pass (Anthropic harness)?

Claude Opus 5 by Anthropic currently leads with a score of 23.58% on LAB all-pass (Anthropic harness).

### How many models are evaluated on LAB all-pass (Anthropic harness)?

1 AI models have been evaluated on LAB all-pass (Anthropic harness) on BenchLM.

### Does LAB all-pass (Anthropic harness) affect BenchLM's overall score?

Not directly. LAB all-pass (Anthropic harness) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
