Benchmark profile
Data Research and Analysis with Complex Operations (DRACO)
Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.
Data verifiedBenchmark score on DRACO — July 24, 2026
BenchLM mirrors the published score view for DRACO. Claude Opus 5 leads the public snapshot at 88.6%. BenchLM does not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout DRACO
Year
2026
Tasks
Agentic data research and analysis tasks
Format
Normalized rubric score
Difficulty
Professional data analysis
Section 8.10.4 reports the max-effort normalized score using Opus 4.6 as judge and a file-based deliverable protocol. Anthropic warns that the judge change makes the result not directly comparable to the paper headline.
BenchLM freshness & provenance
Version
DRACO 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does DRACO measure?
Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.
Which model scores highest on DRACO?
Claude Opus 5 by Anthropic currently leads with a score of 88.6% on DRACO.
How many models are evaluated on DRACO?
1 AI models have been evaluated on DRACO on BenchLM.