arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

点扩散曼巴:数据稀缺下单视图三维重建的统一扩散-状态空间建模

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

Wei Zhou, Xinzhe Shi, Xingxing Hao, Xing Hao, Kang Li, Jinye Peng, Ying He

arXiv 2609.25538首次发表:更新:

发表机构

Northwest University; Nanyang Technological University(西北大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对数据稀缺下单视图三维重建的病态问题,提出点扩散曼巴(PDM),融合扩散模型与状态空间模型,通过层次化特征集成和动态加权采样,在ShapeNet和Pix3D上超越现有方法。

AI 中文摘要

尽管单视图三维重建已取得显著进展,但从本质上模糊的二维观测中推断复杂的三维结构仍然是一个病态问题,尤其是在数据稀缺且研究严重不足的情况下。为应对这一挑战,我们提出了点扩散曼巴(PDM),该方法将扩散模型的生成能力与状态空间模型的效率相结合,用于数据稀缺条件下的单视图三维重建。具体而言,PDM采用了一个轻量级重建模块,以有效处理无序点云输入。通过将局部几何聚合模块与Mamba块相结合,我们的方法联合建模了全局几何结构和局部细节。在三维重建中,初始噪声输入中的每个点都需要精确预测,而Mamba模块提取的高层特征仅从稀疏点中捕获抽象语义信息。为弥合这一差距,我们引入了层次化特征集成网络,为每个点融合高层语义和局部几何特征,克服了基于令牌的点云重建的局限性。此外,我们提出了一种动态加权采样策略,通过利用生成先验增强重建质量,自适应地将三维生成与单视图重建统一起来。在ShapeNet和Pix3D基准上的实验结果表明,PDM优于现有最先进方法,为数据稀缺环境下的三维重建提供了有效解决方案。代码可在以下网址获取:此HTTPS URL。

英文摘要

While single-view 3D reconstruction has seen significant progress, extrapolating complex 3D structures from inherently ambiguous 2D observations remains fundamentally ill-posed, particularly in the critically underexplored data-scarce regime. To address this challenge, we propose Point Diffusion Mamba (PDM), a method that integrates the generative power of diffusion models with the efficiency of state-space model for single-view 3D reconstruction under data-scarce conditions. Specifically, PDM employs a lightweight reconstruction module tailored to handle unordered point-cloud inputs effectively. By combining a Local Geometric Aggregation module with Mamba blocks, our approach jointly models global geometric structures and local details. In 3D reconstruction, each point in the initial noisy input requires a precise prediction, yet the high-level features extracted by the Mamba module capture only abstract semantic information from sparse points. To bridge this gap, we introduce the Hierarchical Feature Integration Network, which fuses high-level semantic and local geometric features for each point, overcoming the limitations of token-based point-cloud reconstruction. Furthermore, we propose a Dynamic Weighted Sampling strategy that adaptively unifies 3D generation with single-view reconstruction by leveraging generative priors to enhance reconstruction quality. Experimental results on the ShapeNet and Pix3D benchmarks demonstrate that PDM outperforms state-of-the-art methods, providing an effective solution for 3D reconstruction under data-scarce settings. Code is available at: https://github.com/NWUzhouwei/PDM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑