AI 中文总结
研究如何利用生成式图像模型进行无训练的原始形状抽象,通过渲染多视图图像、分析语义部分、绘制分割掩码、重新投影和拟合超二次基元的流程,实现与类别无关且方向不变,在相关数据集上Chamfer距离最低。
AI 中文摘要
将三维形状表示为紧凑的几何基元集对机器人技术、模拟和场景理解至关重要。大规模训练的生成式图像模型已成为通用视觉学习者,可直接在图像域中识别和分割对象部分,无需特定任务训练。我们探讨能否直接利用其预训练能力而无需任何训练,并通过无训练方法给出肯定答案。我们的流程包括渲染三维物体的多视图图像,使用视觉语言模型分析其语义部分,促使生成式图像模型绘制颜色编码的部分分割掩码,将其重新投影到几何体上,并通过参数优化为每个部分拟合超二次基元。该方法无学习参数,与类别无关且方向不变。通过真实分割研究表明部分分割是当前精度瓶颈,随着生成模型改进其精度上限会提高。在HumanPrim和Toys4K上,我们的方法平均每个对象使用5 - 9个基元,在所有评估方法中Chamfer距离最低。
英文摘要
Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods depend on their training classes, and optimization-based methods split shapes geometrically rather than into parts. We instead reuse the visual part knowledge of pretrained models without task-specific training or fine-tuning. A vision-language model names parts in multi-view renders, and an unmodified image generator paints color-coded part masks. Reprojection and spatial clustering recover 3D instances, and a classical optimizer fits one tapered and bent superquadric per part. With five to eight primitives per object, the abstractions match the Chamfer distance of the strongest learned baseline on HumanPrim, improve on it by 10% on Toys4K, and have the lowest overlap among compact methods, while chair legs, backrest bars and wheels remain separate primitives. Our accuracy also transfers better than theirs to objects outside the learned methods' ShapeNet training classes. Replacing the generated masks with part labels from the 3D segmentation methods P3-SAM or PartField lowers IoU by 7 to 17 points under the same fitter. Further studies relate the remaining volumetric error to part granularity and to parts that the rendered views observe from one side only.
Comments21 pages, 11 figures, 14 tables