# Claude Opus 5.5 Benchmark Scores & Performance

> Claude Opus 5.5 by Anthropic scores 81.03/100 overall, ranking #4 out of 505 AI models.

Canonical page: https://benchlm.ai/models/claude-opus-5-5

Last updated: September 22, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Anthropic |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [Claude Opus 5.5 system card](https://www.anthropic.com/claude-opus-5-5-system-card) |
| Overall Score | 81.03/100 |
| Overall Rank | #4 of 505 |

## Family & Coverage

- Family: Claude Opus 5.5
- Variant: base
- Benchmarks covered: 63 of 454
- Related earlier model: [Claude Opus 5](/models/claude-opus-5)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 4.0](/benchmarks/terminal-bench-4) | 66.40% |
| [Terminal-Bench-Science 0.1](/benchmarks/terminal-bench-science) | 58.7% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1846 |
| [AutomationBench](/benchmarks/automationbench) | 40.0% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 67.7% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 67.3% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1822 |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 69.5% |
| [AA Harvey LAB](/benchmarks/aaharveylab) | 91.2% |
| [GDP.pdf](/benchmarks/aagdppdf) | 26.2% |
| [OSWorld 2.0](/benchmarks/osworld2) | 48.7% |
| [LAB all-pass (Harvey held-out)](/benchmarks/legalagentbenchheldoutallpass) | 8.3% |
| [LAB criterion-pass (Harvey held-out)](/benchmarks/legalagentbenchheldoutcriterionpass) | 91.2% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 77.8% |
| [Toolathlon Verified Pass@3](/benchmarks/toolathlonverifiedpass3) | 82.4% |
| [Toolathlon Verified Pass³](/benchmarks/toolathlonverifiedpass3all) | 72.2% |
| [Toolathlon Verified avg. turns](/benchmarks/toolathlonverifiedavgturns) | 26.9 turns |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [FrontierCode 1.1 Main](/benchmarks/frontiercode) | 54.4% |
| [AA-SciCode](/benchmarks/aascicode) | 66.9% |
| [CursorBench 4.0](/benchmarks/cursorbench40) | 57.8% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 89.9% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 93.9% |
| [SWE Multimodal](/benchmarks/swe-bench-multimodal) | 61.4% |
| [DeepSWE](/benchmarks/deepswe) | 74.2% |
| [FrontierCode 1.1 Extended](/benchmarks/frontiercode11extended) | 63.6% |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 62.3% |
| [ProgramBench](/benchmarks/programbench) | 91.2% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Chartography (tools)](/benchmarks/chartographywithtools) | 89.0% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 87.7% |
| [Chartography (no tools)](/benchmarks/chartography) | 64.4% |
| [BenchCAD Vision2Code (no tools)](/benchmarks/benchcadvision2code) | 0.730 |
| [BenchCAD Vision2Code (tools)](/benchmarks/benchcadvision2codewithtools) | 0.962 |
| [Biomedical image analysis](/benchmarks/biomedicalimageanalysis) | 71.4% |
| [OfficeQA](/benchmarks/officeqa) | 78.9% |
| [OfficeQA Pro](/benchmarks/officeqapro) | 67.7% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 84.7% |
| [CritPt](/benchmarks/critpt) | 31.7% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [HLE](/benchmarks/hle) | 67.7% |
| [HLE w/o tools](/benchmarks/hlenotools) | 64.4% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 57.6% |
| [AA-HLE](/benchmarks/aahle) | 61.4% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 46.4% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 66.2% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 58.6% |
| [HealthBench (raw)](/benchmarks/healthbench) | 68.1% |
| [HealthBench (length-adjusted)](/benchmarks/healthbenchlengthadjusted) | 60.6% |
| [HealthBench Professional](/benchmarks/healthbenchprofessional) | 65.6% |
| [HealthBench Professional (raw)](/benchmarks/healthbenchprofessionalraw) | 77.1% |
| [BioMysteryBench (human-solvable)](/benchmarks/biomysterybenchhumansolvable) | 89.3% |
| [BioMysteryBench (human-difficult)](/benchmarks/biomysterybenchhumandifficult) | 50.0% |
| [SpatialBench Verified](/benchmarks/spatialbenchverified) | 72.0% |
| [SingleCellBench](/benchmarks/singlecellbench) | 61.2% |
| [Morphology-to-molecule matching](/benchmarks/morphologytomolecule) | 34.0% |
| [Medicinal chemistry](/benchmarks/medicinalchemistry) | 63.5% |
| [Protein Design](/benchmarks/proteindesign) | 60.2% |
| [Protein Design library ranking](/benchmarks/proteindesignlibraryranking) | 56.0% |
| [De novo protein-binder design](/benchmarks/proteinbinderdesign) | 82.6% |
| [Protocols (troubleshooting)](/benchmarks/protocolstroubleshooting) | 73.7% |
| [Protocols (understanding)](/benchmarks/protocolsunderstanding) | 69.0% |

## Multilingual Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GMMLU](/benchmarks/gmmlu) | 94.3% |
| [MILU](/benchmarks/milu) | 93.1% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [ArXivMath Aug. 2026 (no tools)](/benchmarks/arxivmathaugust2026) | 91.2% |
| [ArXivMath Aug. 2026 (tools)](/benchmarks/arxivmathaugust2026withtools) | 96.9% |

## Other Anthropic Models

- [Claude Fable 5.1](/models/claude-fable-5-1) - Score: 83.65
- [Claude Fable 5](/models/claude-fable) - Score: 81.06
- [Claude Opus 5](/models/claude-opus-5) - Score: 80.93
- [Claude Opus 4.8](/models/claude-opus-4-8) - Score: 72.09
- [Claude Opus 4.7](/models/claude-opus-4-7) - Score: 71.17
- [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) - Score: 69.63
- [Claude Opus 4.6](/models/claude-opus-4-6) - Score: 69.53
- [Claude Sonnet 5](/models/claude-sonnet-5) - Score: 69.52
- [Claude Sonnet 4.6](/models/claude-sonnet-4-6) - Score: 63.17
- [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) - Score: 62.7
- [Claude Opus 4.5](/models/claude-opus-4-5) - Score: 58.72
- [Claude Sonnet 4.5](/models/claude-sonnet-4-5) - Score: 54.69
- [Claude Haiku 4.5](/models/claude-haiku-4-5) - Score: 53.49
- [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) - Score: 52.99
- [Claude 4.1 Opus](/models/claude-4-1-opus) - Score: 42.01
- [Claude 4 Sonnet](/models/claude-4-sonnet) - Score: 39.58
- [Claude 3 Opus](/models/claude-3-opus) - Score: 34.85
- [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) - Score: 32.59
- [Claude 4.1 Opus Thinking](/models/claude-4-1-opus-thinking) - Score: 31.34
- [Claude 3 Haiku](/models/claude-3-haiku) - Score: 18.55
- [Claude Mythos 5](/models/claude-mythos-5) - Score: not computed
- [Claude Mythos 5.1](/models/claude-mythos-5-1) - Score: not computed
- [Claude Mythos Preview](/models/claude-mythos-preview) - Score: not computed
- [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) - Score: not computed
- [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) - Score: not computed
