Execution-grounded benchmarks for programmatic CAD — evaluating whether models can understand, generate, and edit the parametric programs behind real mechanical parts.
| Artifact | Status |
|---|---|
| BenchCAD (1.0) — 17,900 parts / 106 families / 4 tasks; scored by voxel IoU, no LLM judge | ✅ Released · frozen · cited in frontier-lab system cards |
| benchcad.com · HF dataset · leaderboard | ✅ Live |
| BenchCAD 2.0 — agentic evaluation (tool use, execution feedback, multi-turn refinement) + dataset v2.0: community-grounded parametric designs | 🚧 In preparation — development roadmap · contribute a family |
| BenchCAD-Agent — reference agent & harness for BenchCAD | 📋 Planned |
| BenchCAD-Assembly — assembly-level CAD: multi-part assemblies, mating constraints | 📋 Planned (reserved) |
| BenchCAD-Pro — harder, contamination-resistant tier with private held-out evaluation | 📋 Planned (reserved) |
“BenchCAD” and the BenchCAD-* prefix refer to official artifacts maintained by the BenchCAD authors. If you build on BenchCAD, we'd love that — please:
- name derivative works X-for-BenchCAD (or similar), not BenchCAD-X;
- score with the official scorer and cite the paper (arXiv:2605.10865);
- want your variant to be official? Open an issue — collaborations are welcome and credited (co-authorship for substantial contributions).
Leaderboard numbers are re-graded by us from raw predictions — never self-reported.