arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.16211cs.CV

Leveling3D: 通过前馈3D高斯点划法与几何感知生成提升3D重建

Leveling3D: Leveling Up 3D Reconstruction with Feed-Forward 3D Gaussian Splatting and Geometry-Aware Generation

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiming Huang, Baixiang Huang, Beilei Cui, Chi Kit Ng, Long Bai, Hongliang Ren

更新

AI总结:

Leveling3D结合前馈3D重建与几何一致生成,解决传统方法在 extrapolated view 中的缺失区域问题,通过几何感知适配器提升3D重建质量,实现生成与重建的同步优化。

AI中文摘要:

前馈3D重建已革新3D视觉,为下游任务如视图合成提供了强大基础。先前工作通过扩散模型修复损坏的渲染结果,但缺乏几何考虑且无法填补 extrapolated view 中的缺失区域。本文提出Leveling3D,一种整合前馈3D重建与几何一致生成的新流程,通过几何感知适配器将扩散模型的内部知识与前馈模型的几何先验对齐。该适配器可在因3D表示欠约束区域导致的extrapolated novel view的artifact区域进行生成。具体而言,为学习更多样化的分布生成,引入了调色板过滤策略用于训练,以及测试时的掩码细化以防止修复区域的杂乱边界。更重要的是,Leveling3D增强的extrapolated novel views可作为前馈3DGS的输入,提升3D重建性能。我们在公开数据集上实现了SOTA性能,包括视图合成和深度估计等任务。

英文摘要:

Feed-forward 3D reconstruction has revolutionized 3D vision, providing a powerful baseline for downstream tasks such as novel-view synthesis with 3D Gaussian Splatting. Previous works explore fixing the corrupted rendering results with a diffusion model. However, they lack geometric concern and fail at filling the missing area on the extrapolated view. In this work, we introduce Leveling3D, a novel pipeline that integrates feed-forward 3D reconstruction with geometrical-consistent generation to enable holistic simultaneous reconstruction and generation. We propose a geometry-aware leveling adapter, a lightweight technique that aligns internal knowledge in the diffusion model with the geometry prior from the feed-forward model. The leveling adapter enables generation on the artifact area of the extrapolated novel views caused by underconstrained regions of the 3D representation. Specifically, to learn a more diverse distributed generation, we introduce the palette filtering strategy for training, and a test-time masking refinement to prevent messy boundaries along the fixing regions. More importantly, the enhanced extrapolated novel views from Leveling3D could be used as the inputs for feed-forward 3DGS, leveling up the 3D reconstruction. We achieve SOTA performance on public datasets, including tasks such as novel-view synthesis and depth estimation.

补充信息

↑