发表机构
National University of Defense Technology; Hunan University; Shenzhen University(国防科技大学; 湖南大学; 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CGGT,一种曲线接地几何变换器,从稀疏无位姿多视图图像中直接预测相机、深度和曲线掩码,经参数优化重建可编辑3D曲线,并构建Wireframe-100K数据集,在稀疏视图下显著提升精度与效率,且泛化至真实图像。
AI 中文摘要
从二维图像中恢复可编辑的三维参数曲线是计算机图形学中的一个基本挑战,它连接了基于像素的感知和基于向量的CAD建模。现有的基于NeRF和3DGS的方法通常依赖于密集的标定视图、预计算的二维边缘图以及昂贵的逐场景优化,限制了它们对随意拍摄的真实世界输入的适用性。我们提出了CGGT,一种曲线接地几何变换器,它直接从稀疏、无位姿的多视图图像中,在图像空间内接地三维一致的二维曲线实例。CGGT结合了一个用于多视图特征学习的几何感知变换器编码器和一个用于跨视图实例关联的曲线感知掩码注意力解码器。在单次前向传播中,它预测相机参数、密集深度图和实例级二维曲线掩码,然后将这些提升到三维,并通过一个快速的参数优化阶段进行细化,以恢复紧凑、可编辑的三维曲线基元。为了支持结构化曲线学习,我们引入了Wireframe-100K,一个大规模数据集,包含100,000个具有多样拓扑结构的CAD模型、逼真的多视图渲染和准确的参数曲线标注。大量实验表明,我们的框架在重建精度和效率方面都取得了显著提升,特别是在具有挑战性的稀疏视图设置下,以及在将持久的二维结构边缘与由轮廓、纹理和外观变化引起的视图相关图像边缘分离方面。尽管仅在合成数据上训练,CGGT对真实世界图像具有良好的泛化能力,展示了其从无约束视觉输入中进行实用CAD风格线框重建的潜力。
英文摘要
Recovering editable 3D parametric curves from 2D images is a fundamental challenge in computer graphics, bridging pixel-based perception and vector-based CAD modeling. Existing NeRF- and 3DGS-based methods often rely on dense calibrated views, precomputed 2D edge maps, and costly per-scene optimization, limiting their applicability to casually captured real-world inputs. We propose CGGT, a Curve-Grounded Geometry Transformer that directly grounds 3D-consistent 2D curve instances in the image space from sparse, unposed multi-view images. CGGT combines a geometry-aware transformer encoder for multi-view feature learning with a curve-aware masked-attention decoder for cross-view instance association. In a single forward pass, it predicts camera parameters, dense depth maps, and instance-level 2D curve masks, which are then lifted into 3D and refined through a fast parametric optimization stage to recover compact, editable 3D curve primitives. To support structured curve learning, we introduce Wireframe-100K, a large-scale dataset comprising 100,000 CAD models with diverse topologies, realistic multi-view renderings, and accurate parametric curve annotations. Extensive experiments show that our framework achieves substantial improvements in both reconstruction accuracy and efficiency, particularly under challenging sparse-view settings and in separating persistent 3D structural edges from view-dependent image edges caused by silhouettes, textures, and appearance variations. Despite being trained solely on synthetic data, CGGT generalizes well to real-world images, demonstrating its potential for practical CAD-style wireframe reconstruction from unconstrained visual inputs.
CommentsAccepted by SIGGRAPH Asia 2026