Benchmark profile
CADGenBench Generation
Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.
How we show CADGenBench Generation
We mirror 19 completed submissions that the official CADGenBench dataset marks validated. The visible score is the generation-task score; aggregate, editing, validity, version, and data-revision fields remain attached to each row.
CADGenBench Generation is display only. Each submission can include a different agent, CAD library, prompt, and tool configuration, so we do not treat the table as a normalized base-model comparison.
Generation score on CADGenBench Generation — August 12, 2026 snapshot
BenchLM mirrors the published generation score view for CADGenBench Generation. build123d-mcp-v0381-claude-opus-5-xhigh-full-r6 leads the public snapshot at 59.6% , followed by gpt-5.6-sol-xhigh-floor-chain-v0379 (45.1%) and Archie in Forge (42.3%). We do not use these results to rank models overall.
build123d-mcp-v0381-claude-opus-5-xhigh-full-r6
Anthropic
pzfreo-build123d-mcp-v0381-claude-opus-5-xhigh-20260807-101147
gpt-5.6-sol-xhigh-floor-chain-v0379
OpenAI
pzfreo-gpt-5-6-sol-xhigh-floor-chain-v0379-20260717-045810
Archie in Forge
Satvik Adyanthaya
satvik-adyanthaya-archie-in-forge-20260711-184330
Generation score table (19 models)
ScoreThe published CADGenBench Generation snapshot places build123d-mcp-v0381-claude-opus-5-xhigh-full-r6 first at 59.6%. The third row is 17.3 points behind. The broader top-10 range is 29.7 points, so the table still separates the published systems.
19 models have been evaluated on CADGenBench Generation. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. CADGenBench Generation is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About CADGenBench Generation
Year
2026
Tasks
49 CAD generation samples
Format
Validity-gated CAD score
Difficulty
Professional mechanical CAD
The benchmark repository, submission dataset, and hosted leaderboard are public. BenchLM stores the generation score separately from editing and keeps the model-card rows display-only.
BenchLM freshness & provenance
Version
CADGenBench Generation 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does CADGenBench Generation measure?
Hugging Face's open benchmark for generating valid and geometrically correct 3D mechanical parts.
Which model leads the published CADGenBench Generation snapshot?
build123d-mcp-v0381-claude-opus-5-xhigh-full-r6 currently leads the published CADGenBench Generation snapshot with 59.6% generation score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on CADGenBench Generation?
19 AI models are included in BenchLM's mirrored CADGenBench Generation snapshot, based on the public leaderboard captured on August 12, 2026 snapshot.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.