发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Vision2CAD通过视觉语言模型与确定性CAD内核结合,解决参数化建模中的几何引用与定位难题,显著提升精度并保持参数依赖。
AI 中文摘要
生成参数化CAD模型需要精确的几何形状和稳定的特征依赖关系。现有方法在选取几何参考、解释草图平面局部坐标以及建立对外部投影几何的草图约束方面面临挑战。我们提出了Vision2CAD,一个视觉智能体工具集,它将视觉语言模型(VLM)推理与确定性CAD内核操作相结合。基于ID的接口支持显式几何选择,局部坐标桥将视图坐标转换为草图坐标,投影边定位支持外部草图约束。这些机制在支持的建模操作和约束类型内建立了特征依赖关系。我们还引入了几何显式参考数据集(GERD),该数据集在每一步建模中对齐了命令、几何状态和ID。在GERD-EVL和DeepCAD测试子集上,Vision2CAD将mIoU分别提高了11.1%和5.6%,并将Chamfer距离分别降低了17.3%和41.8%。参数编辑实验和消融研究进一步证明了参数依赖关系的保持。
英文摘要
Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric references, interpreting sketch-plane local coordinates, and establishing sketch constraints to projected external geometry. We present Vision2CAD, a visual agent harness that combines vision-language model (VLM) reasoning with deterministic CAD kernel operations. An ID-based interface supports explicit geometry selection, a local-coordinate bridge converts view coordinates into sketch coordinates, and projected-edge localization supports external sketch constraints. These mechanisms establish feature dependencies within the supported modeling operations and constraint types. We also introduce the Geometry Explicit Reference Dataset (GERD), which aligned commands, geometry states and IDs at every modeling step. On GERD-EVL and a DeepCAD test subset, Vision2CAD improves mIoU by 11.1\% and 5.6\% and reduces Chamfer distance by 17.3\% and 41.8\%, respectively. Parameter-editing experiments and ablation studies further proved the preservation of parametric dependencies.