# Data Research and Analysis with Complex Operations (DRACO)

> Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.

Canonical page: https://benchlm.ai/benchmarks/draco

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About DRACO

- Year: 2026
- Tasks: Agentic data research and analysis tasks
- Format: Normalized rubric score
- Difficulty: Professional data analysis
- Paper: [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf)

Section 8.10.4 reports the max-effort normalized score using Opus 4.6 as judge and a file-based deliverable protocol. Anthropic warns that the judge change makes the result not directly comparable to the paper headline.

DRACO is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (4 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 88.6% |
| 2 | [Step 5 Preview](/models/step-5-preview) | StepFun | 83.3% |
| 3 | [Hy4 preview](/models/hy4-preview) | Tencent | 77.2% |
| 4 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 70.4% |

## FAQ

### What does DRACO measure?

Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.

### Which model scores highest on DRACO?

Claude Opus 5 by Anthropic currently leads with a score of 88.6% on DRACO.

### How many models are evaluated on DRACO?

4 AI models have been evaluated on DRACO on BenchLM.

### Does DRACO affect BenchLM's overall score?

Not directly. DRACO is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on DRACO

- [Claude Opus 5 vs Step 5 Preview](/compare/claude-opus-5-vs-step-5-preview)
- [Step 5 Preview vs Hy4 preview](/compare/hy4-preview-vs-step-5-preview)
- [Hy4 preview vs Ling 3.0 Flash](/compare/hy4-preview-vs-ling-3-0-flash)
