Benchmark profile
AutoCAD-Bench
A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.
How BenchLM shows AutoCAD-Bench
BenchLM mirrors Markov Studios' AutoCAD-Bench completion table across 50 tasks: 21 2D drawings and 29 3D models. Agents operate AutoCAD 2019 on Windows through visible mouse and keyboard actions and start each task from a clean drawing.
A task counts as completed when the rubric score reaches 75. The rubric combines geometry, dimensions, annotations, and presentation, with a geometry gate. BenchLM keeps the results display only because they combine model capability with a computer-use harness and application environment.
Completion rate on AutoCAD-Bench — July 2026 public results
BenchLM mirrors the published completion rate view for AutoCAD-Bench. GPT-5.6 Sol leads the public snapshot at 46.0% , followed by GPT-5.6 Terra (14.0%) and Claude Fable 5 (10.0%). BenchLM does not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
openai/gpt-5.6-sol
GPT-5.6 Terra
OpenAI
openai/gpt-5.6-terra
Claude Fable 5
Anthropic
anthropic/claude-fable-5
Completion rate table (7 models)
ScoreThe published AutoCAD-Bench snapshot places GPT-5.6 Sol first at 46.0%. The third row is 36.0 points behind. The broader top-10 range is 46.0 points, so the table still separates the published systems.
7 models have been evaluated on AutoCAD-Bench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AutoCAD-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About AutoCAD-Bench
Year
2026
Tasks
21 2D drawing tasks and 29 3D modeling tasks
Format
Task completion rate at a 75-point rubric threshold
Difficulty
Professional CAD computer use
AutoCAD-Bench contains 50 tasks in AutoCAD 2019 on Windows. A rubric combines geometry, dimensions, annotations, and presentation, applies a geometry gate, and counts scores of at least 75 as completed. BenchLM keeps the results display only because model quality, the computer-use harness, and the application environment all affect the result.
BenchLM freshness & provenance
Version
AutoCAD-Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does AutoCAD-Bench measure?
A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.
Which model leads the published AutoCAD-Bench snapshot?
GPT-5.6 Sol currently leads the published AutoCAD-Bench snapshot with 46.0% completion rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on AutoCAD-Bench?
7 AI models are included in BenchLM's mirrored AutoCAD-Bench snapshot, based on the public leaderboard captured on July 2026 public results.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.