Skip to main content

Model profile

Claude Opus 5

AnthropicCurrentReleased Jul 24, 2026
Data verified
Overall Score
85.88Public #1 of 215
Arena Elo
Not listed
Eligible category ranks
4of 8
Price (1M tokens)
$5 in / $25 out
API pricing
Speed
Not listed
Context
Not available

Evidence coverage

79 of 369 tracked benchmarks are published. 65 are verified and 14 provisional. 7 of 8 categories are measured.

Updated July 24, 2026Methodology
Published / tracked
79 / 369
Verified
65
Provisional
14
Categories with evidence
7 / 8

Evidence by category

  • Agentic24 benchmarks
    Mixed evidence
  • Coding12 benchmarks
    Mixed evidence
  • Reasoning5 benchmarks
    Mixed evidence
  • Knowledge21 benchmarks
    Mixed evidence
  • Math5 benchmarks
    Verified
  • Multilingual3 benchmarks
    Verified
  • Multimodal9 benchmarks
    Mixed evidence
  • Inst. Following0 benchmarks
    Not measured
ProprietaryReasoningClaude Opus 5 System Card
Confidence:
High
base

Claude Opus 5 ranks #1 out of 215 models on the public leaderboard with an overall score of 85.88/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. This places it among the top tier of AI models available in 2026, competing directly with the strongest models from all major AI labs.

Claude Opus 5 is a proprietary model. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.

Available on all Claude platforms and through the Claude API as claude-opus-5. It is the default model on Claude Max and the strongest model offered on Claude Pro. Anthropic lists Opus 5 at $5 input / $25 output per million tokens, with Fast mode running around 2.5× the default speed at twice the base price.

BenchLM links it directly to Claude Opus 4.8 as the earlier related model in that lineage. This profile currently has 79 of 369 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Knowledge (#1), while its weakest is Coding (#9). This performance profile makes it particularly effective for knowledge-intensive tasks like research, analysis, and factual Q&A.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 76.685.88

  1. Claude Opus 5Current model
    Anthropic
    #185.88
    Claude Opus 5 is #1 with a score of 85.88.
  2. Claude Mythos 5
    Anthropic
    #283.01
    Claude Mythos 5 is #2 with a score of 83.01.
    Compare
  3. Claude Fable 5
    Anthropic
    #382.76
    Claude Fable 5 is #3 with a score of 82.76.
    Compare
  4. GPT-5.6 Sol
    OpenAI
    #481.46
    GPT-5.6 Sol is #4 with a score of 81.46.
    Compare
  5. Kimi K3
    Moonshot AI
    #579.98
    Kimi K3 is #5 with a score of 79.98.
    Compare
  6. Claude Opus 4.8
    Anthropic
    #677.44
    Claude Opus 4.8 is #6 with a score of 77.44.
    Compare
  7. Muse Spark 1.1
    Meta
    #776.6
    Muse Spark 1.1 is #7 with a score of 76.6.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Knowledge100%
    Eligible cohort rank #1 of 53Category score 93.5
  2. Multimodal94%
    Eligible cohort rank #3 of 32Category score 89.1
  3. Agentic98%
    Eligible cohort rank #3 of 129Category score 69.4
  4. Coding94%
    Eligible cohort rank #9 of 130Category score 68.8

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #3 of 129Percentile 98thWeight 22%24 benchmarksMixed sources69.4
CodingRank #9 of 130Percentile 94thWeight 20%12 benchmarksMixed sources68.8
ReasoningRank Not rankedWeight 17%5 benchmarksMixed sources92.5
KnowledgeRank #1 of 53Percentile 100thWeight 12%21 benchmarksMixed sources93.5
MathWeight 5%5 benchmarksVerifiedScore pending
MultilingualWeight 7%3 benchmarksVerifiedScore pending
MultimodalRank #3 of 32Percentile 94thWeight 12%9 benchmarksMixed sources89.1
Inst. FollowingWeight 5%0 benchmarksNot measuredNot measured

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic24 benchmarks
BrowseCompProvider exact
90.8%Weighted 28%
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 90.8% on BrowseComp.
FrontierBench v0.1Provider exact
43.3%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 43.3% on Anthropic's internal FrontierBench v0.1 run, using mini-SWE-agent, a GKE backend, and five attempts per task.
HLE w/ toolsProvider exact

Humanity's Last Exam with tools

64.7%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 64.7% on Humanity's Last Exam with tools.
DeepSearchQAProvider exact
95.0%Display only
Source: Anthropic: Introducing Claude Opus 5Provenance: Anthropic's launch chart reports the max-effort DeepSearchQA endpoint at 95.0 mean F1 under a 980K-token budget.
DRACOProvider exact

Data Research and Analysis with Complex Operations

88.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.10.4.A reports an 88.6% max-effort normalized DRACO score at a 980K-token budget using Opus 4.6 as judge.
BrowseComp (10-agent, prerelease)Provider exact

Multi-Agent BrowseComp — 10-agent team prerelease configuration

93.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Sections 8.11.1 and 8.11.4 report 93.6% for a 10-agent BrowseComp team on a pre-release Opus 5 configuration without safeguards; the card limits this result to relative comparison.
OSWorld 2.0Provider exact
70.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 70.6% on OSWorld 2.0.
MCP AtlasProvider exact
85.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.2 reports an 85.8% MCP-Atlas pass rate and notes that production effort settings may have changed slightly since evaluation.
MCP-Atlas claim coverageProvider exact

MCP-Atlas mean claim coverage

89.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.2 reports 89.1% mean claim coverage on MCP-Atlas.
LAB all-pass (Anthropic harness)Provider exact

Legal Agent Benchmark all-pass rate — Anthropic harness

23.58%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.3 reports a 23.58% all-pass rate on 1,235 LAB tasks using Anthropic's reduced-tool internal reimplementation.
LAB criterion-pass (Anthropic harness)Provider exact

Legal Agent Benchmark mean criterion-pass rate — Anthropic harness

93.74%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.3 reports a 93.74% mean criterion-pass rate on Anthropic's LAB reimplementation.
LAB all-pass (Harvey held-out)Provider exact

Legal Agent Benchmark all-pass rate — Harvey held-out set

11.7%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.3 reports Harvey AI's 11.7% all-pass rate for Opus 5 on its held-out LAB set.
LAB criterion-pass (Harvey held-out)Provider exact

Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set

94.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.3 reports Harvey AI's 94.1% mean criterion-pass rate for Opus 5 on its held-out LAB set.
GDPval-AAProvider exact
1861Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.4 reports Claude Opus 5 at 1861 Elo on GDPval-AA v2 at max effort, independently evaluated by Artificial Analysis.
Toolathlon-VerifiedProvider exact
80.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.13.6.A reports 80.6% Pass@1 on Toolathlon Verified using Anthropic's pinned internal harness at max effort.
Toolathlon Verified Pass@3Provider exact
87.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.13.6.A reports 87.0% Pass@3 on Toolathlon Verified.
Toolathlon Verified Pass³Provider exact

Toolathlon Verified Pass cubed

73.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.13.6.A reports 73.1% Pass³, meaning all three trials succeeded, on Toolathlon Verified.
Toolathlon Verified avg. turnsProvider exact

Toolathlon Verified average assistant turns

23.5 turnsDisplay only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.13.6.A reports 23.5 average assistant turns per Toolathlon Verified trajectory.
AutomationBenchProvider exact
26.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.7 reports Claude Opus 5 at 26.0% on AutomationBench at max effort.
AA Agentic IndexReported

Artificial Analysis Agentic Index

55.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
GDPval-AAReported

GDPval-AA normalized

68.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA Tau3 BankingReported

Artificial Analysis Tau3-Banking

30.3%Display only
Source: Artificial Analysis: tau3-banking leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA BriefcaseProvider exact

Artificial Analysis Briefcase

1720Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 1720 Elo on AA-Briefcase.
aaTerminalBench21Reported
89.1%Display only
Source: Artificial Analysis: terminalbench-v2-1 leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Coding12 benchmarks
SWE-bench VerifiedProvider exact

Software Engineering Benchmark Verified

96%Weighted 16%
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.2 reports Claude Opus 5 at 96.0% on SWE-bench Verified, averaged over five trials.
SWE-bench ProProvider exact
79.2%Weighted 10%
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 79.2% on SWE-bench Pro.
SWE MultilingualProvider exact
89.5%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 89.5% on SWE-bench Multilingual.
SWE MultimodalProvider exact

SWE-bench Multimodal

59.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 59.4% on SWE-bench Multimodal.
deepSweProvider exact
68.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 68.8% on DeepSWE v1.1.
FrontierCode 1.1 MainProvider exact
53.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 53.4% on FrontierCode 1.1 Main.
FrontierCode 1.1 ExtendedProvider exact
63.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.4 reports Claude Opus 5 at 63.6% on FrontierCode 1.1 Extended at its best effort setting.
ProgramBench (episode 1)Provider exact

ProgramBench hidden-test pass rate after episode 1

83.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.9.1 reports an 83% hidden-test pass rate after the first episode on 166 golden ProgramBench tasks.
ProgramBenchProvider exact

ProgramBench: Can Language Models Rebuild Programs From Scratch?

93.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.9.1 reports a 93% hidden-test pass rate after the fifth sequential episode on 166 golden ProgramBench tasks.
cursorBench32Benchmark exact
70.0%Display only
Source: Cursor evals: CursorBench 3.2Provenance: Cursor reports Opus 5 Max at this exact CursorBench 3.2 score on its public evals page. BenchLM stores it on the Claude Opus 5 row as a display-only coding-agent benchmark.
AA Coding IndexProvider exact

Artificial Analysis Coding Index

78.0%Display only
Source: Anthropic: Introducing Claude Opus 5Provenance: Anthropic's launch chart reports the Artificial Analysis Coding Agent Index effort ladder. BenchLM stores the 64.8 max-effort endpoint; Artificial Analysis did not yet expose a live Opus 5 model row when checked on July 24, 2026.
AA-SciCodeReported

Artificial Analysis SciCode

55.7%Display only
Source: Artificial Analysis: scicode leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Reasoning5 benchmarks
ARC-AGI-2Provider exact

Abstraction and Reasoning Corpus for AGI v2

90.4%Weighted 31%
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 90.4% on ARC-AGI-2 at max effort.
ARC-AGI-1Provider exact

ARC-AGI-1 Semi-Private Evaluation

97.50%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.14.1 reports the ARC Prize Foundation's verified 97.50% ARC-AGI-1 semi-private score at max effort.
ARC-AGI-3Provider exact

Abstraction and Reasoning Corpus for AGI v3

30.2%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.14.2 reports a verified 30.16% RHAE score on the ARC-AGI-3 semi-private evaluation at high effort; max-effort results were not available at launch.
AA-LCRReported

Artificial Analysis Long Context Reasoning

70.0%Display only
Source: Artificial Analysis: artificial-analysis-long-context-reasoning leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
CritPtReported

Critical Physics Tasks

29.1%Display only
Source: Artificial Analysis: critpt leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Knowledge21 benchmarks
HLEProvider exact

Humanity's Last Exam

64.7%Weighted 45%
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 64.7% on Humanity's Last Exam with tools.
HLE w/o toolsProvider exact

Humanity's Last Exam without tools

56.3%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Table 8.1.A reports Claude Opus 5 at 56.3% on Humanity's Last Exam without tools.
HealthBench (raw)Provider exact

HealthBench raw score

67.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.15.1 reports a 67.1% raw HealthBench score at max effort, averaged over five trials without tools.
HealthBench (length-adjusted)Provider exact

HealthBench length-adjusted score

57.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.15.1 reports a 57.8% length-adjusted HealthBench score.
HealthBench ProfessionalProvider exact
59.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.15.2 reports a 59.8% length-adjusted HealthBench Professional score.
HealthBench Professional (raw)Provider exact

HealthBench Professional raw score

73.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.15.2 reports a 73.4% raw HealthBench Professional score.
BioMysteryBench (human-solvable)Provider exact

BioMysteryBench Human Solvable

90.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.1 reports 90.1% on the revised BioMysteryBench Human Solvable subset.
BioMysteryBench (human-difficult)Provider exact

BioMysteryBench Human Difficult

49.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.1 reports 49.4% on the revised BioMysteryBench Human Difficult subset.
SpatialBench VerifiedProvider exact

LatchBio SpatialBench Verified

72.5%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.2 reports 72.5% on LatchBio SpatialBench Verified.
SingleCellBenchProvider exact

LatchBio SingleCellBench

60.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.2 reports 60.6% on LatchBio SingleCellBench.
ProteinGym HardProvider exact
47.7%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.3 reports 47.7% on ProteinGym Hard.
Protein DesignProvider exact

Anthropic Protein Design evaluation

42.5%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.4 reports 42.5% on Anthropic's Protein Design evaluation.
Organic chemistry V2Provider exact

Anthropic Organic Chemistry V2 evaluation

61.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.5 reports 61.6% on Organic Chemistry V2; ten timed-out attempts were excluded from aggregation.
Protocols (troubleshooting)Provider exact

Molecular Biology Protocols Troubleshooting

61.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.6 reports 61.1% on the Protocols Troubleshooting variant.
Protocols (understanding)Provider exact

Benchling Molecular Biology Protocols Understanding

78.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.17.6 reports 78.4% on the Benchling-created Protocols Understanding variant.
Artificial Analysis Intelligence IndexReported
60.7%Display only
Source: Artificial Analysis: artificial-analysis-intelligence-index leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

93.2%Display only
Source: Artificial Analysis: gpqa-diamond leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

52.6%Display only
Source: Artificial Analysis: humanitys-last-exam leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

31.3%Display only
Source: Artificial Analysis: omniscience leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-Omniscience AccuracyReported

Artificial Analysis Omniscience Accuracy

54.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience Hallucination RateReported

Artificial Analysis Omniscience Hallucination Rate

50.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Math5 benchmarks
IMO 2026Provider exact

International Mathematical Olympiad 2026

42/42Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.6 reports 42/42 on IMO 2026: all 24 generated proofs were accepted by a three-model panel and the six pre-specified proofs received full credit from human experts.
RiemannBench (no tools)Provider exact

RiemannBench without tools

60.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.7.A reports 60% on the corrected private 25-problem RiemannBench setup at max effort without tools.
RiemannBench (tools)Provider exact

RiemannBench with tools

79.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.7.A reports 79% on the corrected private 25-problem RiemannBench setup at max effort with tools.
ArXivMath Jun. 2026 (no tools)Provider exact

ArXivMath June 2026 without tools

90.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.8 reports 90.8% on the 49-problem June 2026 ArXivMath release at max effort without tools, averaged over four runs.
ArXivMath Jun. 2026 (tools)Provider exact

ArXivMath June 2026 with tools

91.3%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.8 reports 91.3% on the 49-problem June 2026 ArXivMath release at max effort with tools, averaged over four runs.
Multilingual3 benchmarks
GMMLUProvider exact

Global MMLU

92.5%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.16.1.A reports 92.5% GMMLU average accuracy from one max-effort trial without tools.
MILUProvider exact

Multi-task Indic Language Understanding Benchmark

92.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.16.2.A reports 92.1% MILU average accuracy over five max-effort trials without tools.
INCLUDEProvider exact
89.8%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Figure 8.16.3.A reports 89.8% INCLUDE average accuracy over five max-effort trials without tools.
Multimodal9 benchmarks
OfficeQA ProProvider exact
66.9%Weighted 30%
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.1 reports 66.9% on the harder 133-question OfficeQA Pro subset.
Chartography (no tools)Provider exact

Chartography without tools

29.6%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.1 reports 29.6% on Chartography at max effort without tools, averaged over five runs.
Chartography (tools)Provider exact

Chartography with image and code tools

83.0%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.1 reports 83.0% on Chartography at max effort with a container and image-cropping tool, averaged over five runs.
BenchCAD Vision2Code (no tools)Provider exact

BenchCAD Vision2Code voxel IoU without tools

0.366Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.2 reports 0.366 voxel IoU on a random 1,000-file BenchCAD Vision2Code subset at max effort without tools.
BenchCAD Vision2Code (tools)Provider exact

BenchCAD Vision2Code voxel IoU with tools

0.821Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.2 reports 0.821 voxel IoU on a random 1,000-file BenchCAD Vision2Code subset at max effort with tools.
GDP.pdf (no tools)Provider exact

GDP.pdf mean criteria pass rate without tools

83.4%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.4 reports an 83.4% mean criteria pass rate on GDP.pdf without tools using Anthropic's corrected internal harness.
GDP.pdf (tools)Provider exact

GDP.pdf mean criteria pass rate with tools

85.5%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.12.4 reports an 85.5% mean criteria pass rate on GDP.pdf with tools using Anthropic's corrected internal harness.
OfficeQAProvider exact
78.1%Display only
Source: Anthropic: Claude Opus 5 system cardProvenance: Section 8.13.1 reports 78.1% on OfficeQA using extracted documents and code execution through the public Messages API with production safeguards.
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

84.7%Display only
Source: Artificial Analysis: mmmu-pro leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.

Claude Opus 5 Family

Base entry

Related Earlier Model

Claude Opus 4.8

Frequently Asked Questions

How does Claude Opus 5 perform overall in AI benchmarks?

Claude Opus 5 currently ranks #1 out of 215 models on BenchLM's provisional leaderboard with an overall score of 85.88. It is created by Anthropic.

Is Claude Opus 5 good for knowledge and understanding?

Claude Opus 5 ranks #1 out of 53 models in knowledge and understanding benchmarks with an average score of 93.5. It is among the top performers in this category.

Is Claude Opus 5 good for coding and programming?

Claude Opus 5 ranks #9 out of 130 models in coding and programming benchmarks with an average score of 68.8. It is among the top performers in this category.

Is Claude Opus 5 good for mathematics?

Claude Opus 5 has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.

Is Claude Opus 5 good for reasoning and logic?

Claude Opus 5 has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is Claude Opus 5 good for agentic tool use and computer tasks?

Claude Opus 5 ranks #3 out of 129 models in agentic tool use and computer tasks benchmarks with an average score of 69.4. It is among the top performers in this category.

Is Claude Opus 5 good for multimodal and grounded tasks?

Claude Opus 5 ranks #3 out of 32 models in multimodal and grounded tasks benchmarks with an average score of 89.1. It is among the top performers in this category.

Is Claude Opus 5 good for multilingual tasks?

Claude Opus 5 has visible benchmark coverage in multilingual tasks, but BenchLM does not currently assign it a global category rank there.

Does Claude Opus 5 have full benchmark coverage on BenchLM?

Not yet. Claude Opus 5 currently has 79 published benchmark scores out of the 369 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Claude Opus 5?

Claude Opus 5's context window has not been published in a source we can verify yet. The profile leaves this field unavailable instead of borrowing a limit from an earlier model in the family.

Last updated: July 24, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.