Axolotl3D:用于精确3D形状完成的统一框架
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
浏览论文内容
中文总结 AI 辅助
针对3D生成模型在多场景适用性有限及缺乏统一框架的问题,提出Axolotl3D模型,它基于多种模态条件设定,利用点云和相机参数,通过统一训练策略实现精确3D形状完成,实验验证其在多方面性能出色。
中文摘要 AI 辅助
近期的3D生成模型利用大规模先验和扩散架构从单张图像生成高质量几何形状。但它们假设完全可见性和单视图输入,限制了在多视图、遮挡或编辑场景中的适用性。虽然先前工作分别应对了这些挑战,但缺乏在不同条件信号下可控3D完成的统一框架。我们提出Axolotl3D,一种多模态且感知遮挡的3D生成模型,它联合基于图像、可见性掩码、相机参数和部分点云进行条件设定。点云作为几何锚点促进精确形状完成,相机参数确保在共享3D坐标系中一致的多视图对齐。统一训练策略从大规模3D数据合成不同条件设定,实现强大的跨模态推理。在Toys4K和OmniObject3D上的实验证明了在干净和遮挡设置下的最优性能,以及在真实世界重建和几何一致编辑方面的出色结果。
英文摘要
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。