GIFSplat: 生成先验引导的迭代前馈3D高斯点云从稀疏视角重建
GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views
- La Trobe University(拉特罗布大学)
- Cisco Research(思科研究)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GIFSplat通过生成先验引导的迭代前馈方法,从稀疏视角高效重建3D高斯点云,提升PSNR并保持推理效率。
AI中文摘要:
前馈3D重建在运行时间上相比每场景优化有显著优势,后者在推理时仍然缓慢且在稀疏视角下往往脆弱。然而,现有前馈方法仍有进一步提升的潜力,尤其是在域外数据上,并且在引入生成先验后难以保持第二级推理时间。这些限制源于现有前馈流水线中的一次性预测范式:模型严格受限于容量,缺乏推理时间的细化,并且不适用于持续注入生成先验。我们引入GIFSplat,一种纯前馈迭代细化框架,用于从稀疏未置位视角进行3D高斯点云重建。少量仅前向的残差更新逐步通过渲染证据细化当前3D场景,实现了效率与质量之间的良好平衡。此外,我们从增强的新型渲染中蒸馏出冻结的扩散先验,生成高斯级线索,无需反向传播或持续增加视角集,从而在保持前馈效率的同时实现每场景适应性。在DL3DV、RealEstate10K和DTU上,GIFSplat一致优于最先进的前馈基线,将PSNR提升高达+2.1 dB,并且在不需相机姿态或任何测试时间梯度优化的情况下保持第二级推理时间。
英文摘要:
Feed-forward 3D reconstruction offers substantial runtime advantages over per-scene optimization, which remains slow at inference and often fragile under sparse views. However, existing feed-forward methods still have potential for further performance gains, especially for out-of-domain data, and struggle to retain second-level inference time once a generative prior is introduced. These limitations stem from the one-shot prediction paradigm in existing feed-forward pipeline: models are strictly bounded by capacity, lack inference-time refinement, and are ill-suited for continuously injecting generative priors. We introduce GIFSplat, a purely feed-forward iterative refinement framework for 3D Gaussian Splatting from sparse unposed views. A small number of forward-only residual updates progressively refine current 3D scene using rendering evidence, achieve favorable balance between efficiency and quality. Furthermore, we distill a frozen diffusion prior into Gaussian-level cues from enhanced novel renderings without gradient backpropagation or ever-increasing view-set expansion, thereby enabling per-scene adaptation with generative prior while preserving feed-forward efficiency. Across DL3DV, RealEstate10K, and DTU, GIFSplat consistently outperforms state-of-the-art feed-forward baselines, improving PSNR by up to +2.1 dB, and it maintains second-scale inference time without requiring camera poses or any test-time gradient optimization.