VS-Splat:用于稀疏视图端到端三维物体重建的体素选择性前馈高斯泼溅
VS-Splat: Voxel-Selective feed-forward Gaussian Splatting for end-to-end 3D object reconstruction from sparse-views
- Sungkyunkwan University (SKKU)(成均馆大学)
- AiM Future
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VS-Splat提出一种端到端前馈高斯泼溅框架,通过可学习体素选择仅在物体区域预测原语,无需三维监督,在稀疏视图重建中优于现有方法。
AI中文摘要:
前馈高斯泼溅模型在从少量二维图像重建三维物体方面已展现出显著效果,即使这些物体是未见过的。由于现有方法通常在整个三维空间中均匀预测高斯原语,大多数原语被放置在非物体区域,这可能阻碍对物体精细细节的表示。本文提出了一种体素选择性高斯泼溅模型(VS-Splat),这是一种新的端到端前馈高斯泼溅框架,仅在被选中的、可能属于物体的体素内预测大量原语,且无需三维结构监督。为实现此目标,我们提出了一种新的可学习体素选择方法,该方法仅利用二维渲染来识别以物体为中心的体素。在三个基准数据集上的稀疏视图渲染实验表明,所提出的VS-Splat优于多种最先进方法。我们进一步证明了其作为现有致密化方法骨干网络的有效性,并展示了一种可选扩展可提高其对不准确相机位姿估计的鲁棒性。
英文摘要:
Feed-forward Gaussian splatting models have demonstrated remarkable effectiveness in reconstructing three-dimensional (3D) objects from a few two-dimensional (2D) images, even if they are unseen. As existing methods typically predict Gaussian primitives uniformly across the 3D space, most primitives are placed in non-object regions. This may hinder the representation of fine object details. This paper proposes a Voxel-Selective Gaussian Splatting model (VS-Splat), a new end-to-endfeed-forward Gaussian splatting framework that predicts many primitives only within selected voxels that are likely to belong to an object, without 3D structural supervision. To achieve this, we propose a new learnable voxel selection approach that identifies object-centric voxels only with 2D rendering supervision. Our sparse-view rendering experiments with three benchmark datasets show that proposed VS-Splat outperforms several state-of-the-art methods. We further demonstrate its effectiveness as a backbone for an existing densification method and show that anoptional extension improves its robustness to inaccurate camera pose estimates.