CADGenBench Generation
We show this table for reference; we do not rank on it.
Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.
Generation score on CADGenBench Generation — September 23, 2026 snapshot
We mirror the published generation score view for CADGenBench Generation. Godela leads the public snapshot at 84.5%, followed by build123d-mcp-v0381-claude-opus-5-xhigh-full-r6 (59.6%) and gpt-5.6-sol-xhigh-floor-chain-v0379 (45.1%). We do not use these results to rank models overall.
Godela
Godela
build123d-mcp-v0381-claude-opus-5-xhigh-full-r6
Anthropic
gpt-5.6-sol-xhigh-floor-chain-v0379
OpenAI
24 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot
Generation score table (24 models)
ScoreHow we show CADGenBench Generation
We mirror 24 completed submissions that the official CADGenBench dataset marks validated. The visible score is the generation-task score; aggregate, editing, validity, version, and data-revision fields remain attached to each row.
CADGenBench Generation is display only. Each submission can include a different agent, CAD library, prompt, and tool configuration, so we do not treat the table as a normalized base-model comparison.
Snapshot
The published CADGenBench Generation snapshot places Godela first at 84.5%. The third row is 39.4 points behind. The broader top-10 range is 48.4 points, so the table still separates the published systems.
24 models have been evaluated on CADGenBench Generation. The benchmark falls in the Coding category. CADGenBench Generation is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About CADGenBench Generation
Year
2026
Tasks
49 CAD generation samples
Format
Validity-gated CAD score
Difficulty
Professional mechanical CAD
The benchmark repository, submission dataset, and hosted leaderboard are public. BenchLM stores the generation score separately from editing and keeps the model-card rows display-only.
Freshness and provenance
Version
CADGenBench Generation 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CADGenBench Generation measure?
Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.
Which model leads the published CADGenBench Generation snapshot?
Godela currently leads the published CADGenBench Generation snapshot with 84.5% generation score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on CADGenBench Generation?
The September 23, 2026 snapshot snapshot contains 24 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.