arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.09853cs.CV

DeltaSplat:用于无位姿前馈三维高斯泼溅的迭代高斯细化

DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting

  • Sungkyunkwan University(成均馆大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

Chanung Park, Seunghyeon Song, Joo Chan Lee, Eunbyung Park, Jong Hwan Ko

AI总结:

DeltaSplat通过迭代渲染残差并利用普吕克射线和深度先验细化高斯,以轻量模块纠正无位姿前馈3DGS的相机误差,在DL3DV上提升1.75 dB PSNR并超越真值相机基线。

AI中文摘要:

无位姿前馈三维高斯泼溅(3DGS)在单次网络前向传播中从稀疏、无位姿图像重建场景,免除了相机标定和逐场景优化的需求。然而,相机估计误差会传播到预测的高斯中,并加剧单次预测的几何和光度不准确性。为纠正这些误差,我们引入DeltaSplat,一种用于无位姿前馈3DGS的轻量级高斯细化模块。它在输入上下文视图处迭代渲染当前高斯,并从所得残差中预测每个高斯的更新。然而,仅凭二维残差不足以确定三维校正。因此,DeltaSplat将每次更新基于逐像素普吕克射线和渲染深度作为软几何先验。一个双分支卷积混合器高效编码这些输入,各属性头将融合特征解码为位置、不透明度和颜色更新。该模块仅向主干网络增加约2.2%的参数,并在推理时保持完全前馈。在DL3DV上,DeltaSplat在无位姿设置下达到26.64 dB PSNR,比其最先进的主干网络提升1.75 dB,甚至超过使用真实相机提供的基线;在6-24视图和所有相机模式下均保持一致增益。

英文摘要:

Pose-free feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene from sparse, unposed images in a single network pass, removing the need for camera calibration and per-scene optimization. However, camera estimation errors propagate into the predicted Gaussians and compound the geometric and photometric inaccuracies of single-pass prediction. To correct these errors, we introduce DeltaSplat, a lightweight Gaussian refinement module for pose-free feed-forward 3DGS. It iteratively renders the current Gaussians at the input context views and predicts per-Gaussian updates from the resulting residuals. A 2D residual alone, however, underdetermines the 3D correction. DeltaSplat therefore conditions each update on per-pixel Plücker rays and rendered depth as a soft geometric prior. A dual-branch convolutional mixer efficiently encodes these inputs, and per-attribute heads decode the fused features into position, opacity, and color updates. The module adds only ~2.2% parameters to the backbone and remains fully feed-forward at inference. On DL3DV, DeltaSplat reaches 26.64 dB PSNR in the pose-free setting, improving its state-of-the-art backbone by 1.75 dB and surpassing even baselines supplied with ground-truth cameras; consistent gains hold across 6-24 views and all camera regimes.

补充信息

↑