发表机构
JoyIndustrial(卓工业)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RealCADBench是面向工业设计意图的参数化CAD建模基准,含12632项多模态任务,评估9个大模型发现无模型在所有指标领先,执行性不足以表征真实CAD建模,模型在多指标上差异显著。
AI 中文摘要
参数化计算机辅助设计(CAD)建模难以用单一指标评估。现有CAD基准通常侧重合成或CAD原生场景、输入模态有限,或仅关注可执行性与交并比(IoU)。我们推出RealCADBench,这是一个针对真实工业设计意图的意图转程序CAD建模基准。它包含来自19个工厂自动化类别的12632项任务,涵盖零件建模与装配建模的文本描述、二维工程图、真实产品图片及渲染图像。我们在1770项任务的评估子集上报告结果:四个输入场景下的1745项零件任务,以及用于所有已报告装配比较的RCB-Assm25(25项装配研究任务)。每种方法生成FreeCAD API Python代码,由共享运行时执行以导出三维模型。我们用可执行性、实体IoU、表面IoU及基于规则的视觉语义一致性Judge对导出模型进行评估。在评估的9个独立前沿大模型中,没有模型在所有四个指标上均领先。在六个前沿规模大模型中,四个零件场景下的可执行性范围为0.565至0.812,实体IoU为0.2841至0.5379,表面IoU为0.112至0.217;场景均衡综合得分最高的模型,与四个组成指标的领先模型并非同一模型。在RCB-Assm25上,带有GPT-5.5的Codex相比独立的GPT-5.5提升了可执行性和两种IoU指标,但使Judge得分降低了6.98个百分点,因此GPT-5.5仍是Judge得分的领先者。我们还观察到反复出现的失败模式,最显著的是缺失精细结构、零件特征丢失及装配放置错误。这些结果表明,仅执行性不足以表征真实CAD建模,前沿模型与智能体在可执行性、IoU及视觉语义一致性上存在显著差异。
英文摘要
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.