# AutoCAD-Bench

> A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.

Canonical page: https://benchlm.ai/benchmarks/autocadbench

- Category: [Agentic](/agentic)
- Last updated: July 2026 public results

## About AutoCAD-Bench

- Year: 2026
- Tasks: 21 2D drawing tasks and 29 3D modeling tasks
- Format: Task completion rate at a 75-point rubric threshold
- Difficulty: Professional CAD computer use
- Paper: [AutoCAD-Bench](https://www.markovstudios.com/research/autocad-bench)

AutoCAD-Bench contains 50 tasks in AutoCAD 2019 on Windows. A rubric combines geometry, dimensions, annotations, and presentation, applies a geometry gate, and counts scores of at least 75 as completed. BenchLM keeps the results display only because model quality, the computer-use harness, and the application environment all affect the result.

AutoCAD-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (7 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 46.0% |
| 2 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 14.0% |
| 3 | [Claude Fable 5](/models/claude-fable) | Anthropic | 10.0% |
| 4 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 8.0% |
| 5 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 0.0% |
| 6 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 0.0% |
| 7 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 0.0% |

## FAQ

### What does AutoCAD-Bench measure?

A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.

### Which model leads the published AutoCAD-Bench snapshot?

GPT-5.6 Sol currently leads the published AutoCAD-Bench snapshot with a score of 46.0%.

### How many models are evaluated on AutoCAD-Bench?

The July 2026 public results contains 7 AI models.

### Does AutoCAD-Bench affect BenchLM's overall score?

Not directly. AutoCAD-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
