# CADGenBench Generation

> Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.

Canonical page: https://benchlm.ai/benchmarks/cadgenbenchgeneration

- Category: [Coding](/coding)
- Last updated: September 23, 2026 snapshot

## About CADGenBench Generation

- Year: 2026
- Tasks: 49 CAD generation samples
- Format: Validity-gated CAD score
- Difficulty: Professional mechanical CAD
- Paper: [CADGenBench](https://github.com/huggingface/cadgenbench)

The benchmark repository, submission dataset, and hosted leaderboard are public. BenchLM stores the generation score separately from editing and keeps the model-card rows display-only.

CADGenBench Generation is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (24 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Godela](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Godela | 84.5% |
| 2 | [build123d-mcp-v0381-claude-opus-5-xhigh-full-r6](https://github.com/pzfreo/cadgenbench-build123d/tree/0d0a3b6732c11a7b93188ddb5bb904cc9c4ae68f) | Anthropic | 59.6% |
| 3 | [gpt-5.6-sol-xhigh-floor-chain-v0379](https://github.com/pzfreo/cadgenbench-build123d/tree/1c79c494dc6801b9c8bab1e27fbcceba7ee41b02) | OpenAI | 45.1% |
| 4 | [Archie in Forge](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Satvik Adyanthaya | 42.3% |
| 5 | [Grok 4.6 high Mecado Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | xAI | 41.5% |
| 6 | [Grok 4.6 xhigh Mecado Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | xAI | 39.1% |
| 7 | [Claude Fable 5](/models/claude-fable) | Anthropic | 37.3% |
| 8 | [gpt-5.5-build123d-mcp-0.3.59-xhigh-r5](https://github.com/pzfreo/cadgenbench-build123d/tree/0a8a78da20c726181a6210924096a8c208eafd4a) | OpenAI | 37.2% |
| 9 | [GPT-5.6 Sol xhigh Mecado Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | OpenAI | 36.4% |
| 10 | [Claude Opus 5 max Mecado Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Anthropic | 36.1% |
| 11 | [GPT-5.6 Sol HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | OpenAI | 35.5% |
| 12 | [gpt-5.5-build123d-mcp-0.3.56-v2](https://github.com/pzfreo/cadgenbench-build123d/tree/76d6de296a04641d0e7308f08aa9be94a0eca959) | OpenAI | 34.5% |
| 13 | [GPT-5.5 Pro HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | OpenAI | 32.1% |
| 14 | [build123d-mcp + Claude Opus 4.8 (full, 81 fixtures)](https://github.com/pzfreo/cadgenbench-build123d) | Anthropic | 31.5% |
| 15 | [Claude Opus 4.7 HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Anthropic | 29.9% |
| 16 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 29.7% |
| 17 | [Opus 4.8 — Claude Code + text-to-cad skill](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Anthropic | 28.0% |
| 18 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 27.4% |
| 19 | [Claude Opus 4.6 HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Anthropic | 23.6% |
| 20 | [Gemini 3.1 Flash-Lite HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Google | 22.4% |
| 21 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 22.1% |
| 22 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 21.1% |
| 23 | [GLM-4.6V HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Z.AI | 16.4% |
| 24 | [Qwen3-VL 235B-A22B Instruct HF Baseline with Build123d](https://huggingface.co/datasets/HuggingAI4Engineering/cadgenbench-submissions) | Alibaba | 16.3% |

## FAQ

### What does CADGenBench Generation measure?

Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.

### Which model leads the published CADGenBench Generation snapshot?

Godela currently leads the published CADGenBench Generation snapshot with a score of 84.5%.

### How many models are evaluated on CADGenBench Generation?

The September 23, 2026 snapshot contains 24 AI models.

### Does CADGenBench Generation affect BenchLM's overall score?

Not directly. CADGenBench Generation is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
