VGGT-CAD:基于几何约束的参数化CAD三维模型重建
VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding
查看机构详情
- Nanjing University of Science and Technology(南京理工大学)
- KOKONI3D, Moxin (Huzhou) Technology Co., Ltd.(KOKONI3D,魔芯(湖州)科技有限公司)
- Nanjing Institute of Agricultural Mechanization, Ministry of Agriculture and Rural Affairs(农业农村部南京农业机械化研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出VGGT-CAD框架,利用预训练三维几何先验和可变视图聚合模块,从单/多视图观测重建参数化CAD模型,并构建VideoCAD基准验证效果。
中文摘要 AI 辅助
参数化CAD重建需要从视觉观测中恢复精确的几何形状和可编辑的建模操作,在观测有限且模糊的情况下具有挑战性。现有方法主要依赖二维外观线索,缺乏强大的多视图几何先验。在本工作中,我们提出了VGGT-CAD,一个从单视图和多视图观测中进行参数化CAD重建的几何感知框架。我们通过将相机参数编码为条件令牌,并与图像令牌联合建模,将预训练的三维几何先验迁移到CAD重建中。为了处理不同数量的视点,我们引入了一个可变视图的跨视图上下文聚合模块,自适应地融合多视图特征。我们进一步开发了一种无需训练的几何感知视图选择策略,在推理过程中选择互补且可靠的帧。生成的表示通过非自回归解码器解码为CAD命令序列。我们还开发了VideoCAD,一个从现有CAD数据通过多视图重新渲染得到的大规模多视图视频基准。大量实验证明了VGGT-CAD在不同观测配置下进行视觉CAD重建的有效性。
英文摘要
Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.