AI 中文总结
本文提出CADEngBench双轨基准,从参数化设计、装配推理等维度评估8种多模态代码模型,发现编辑现有CAD更易,复杂编辑与匹配FEA仍难,装配预测存在缺陷,强调CAD评估需关注工程行为而非外观。
AI 中文摘要
CAD模型仅外观正确并不等于工程级,它必须满足设计要求、对参数变化做出可预测响应、支持可控编辑、在指定分析下匹配参考结构响应,并通过有效连接与其他部件相连。本文提出CADEngBench,一个针对上述能力的双轨基准。CADEngBench-P评估300个参数化部件,每个部件用于一个从无到有创建CAD的任务和一个功能编辑任务(共600个任务),评估指标包括边界表示(B-Rep)有效性、工程与可制造性设计(DFM)检查、参数族扰动、功能编辑以及在CalculiX中匹配的线性静力学有限元分析(FEA)。CADEngBench-A评估150个体素对,评估指标包括排名式关节检索、精确面与边定位、关节框架预测以及运动学验证。针对8种多模态、具备代码能力的模型,编辑现有CAD比生成CAD容易得多,而复杂编辑和匹配FEA仍具挑战性;装配预测常能定位相关区域,但无法恢复记录的关节或配合实体。这些结果表明,CAD评估必须测试工程行为而非仅外观。
英文摘要
A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification. Across eight multimodal, code-capable models, editing supplied CAD is substantially easier than generating it, while complex edits and matched FEA remain difficult. Assembly predictions often locate the relevant region but fail to recover the recorded joint or mating entities. These results show that CAD evaluation must test engineering behavior rather than appearance alone.