arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

D^2-4DGS:双深度引导的稀疏相机四维高斯溅射

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

Jijian Zhao

arXiv 2608.01588首次发表:更新:

AI 中文总结

针对稀疏相机动态四维高斯溅射的几何监督不足问题,提出双深度先验引导的D^2-4DGS框架,经9个数据集-视图设置验证,其峰值信噪比平均较最佳对比方法提升1.33dB。

AI 中文摘要

动态四维高斯溅射已成为一种高效的表征方法,用于动态新视图合成,通过显式场景建模和实时渲染实现。然而,现有方法通常需要密集多视图视频以获取足够的几何约束,导致采集成本高昂且限制了稀疏相机的部署。减少输入视图可降低采集成本,但会削弱几何监督,常引发结构缺失和漂浮高斯。深度先验提供几何线索,但单一来源无法同时具备密集覆盖和可靠几何:单目深度提供密集结构但存在尺度歧义与局部偏差,而多视图几何深度提供与重建坐标系一致的不完整锚点。为利用二者互补性,我们提出D^2-4DGS,一种由双源深度先验引导的稀疏相机动态四维高斯溅射框架。我们将单目估计与有效多视图几何深度对齐,并验证其一致性以识别可靠几何锚点,这些经验证的锚点支持感知一致性的剪枝与深度监督;同时,经验证的几何深度与对齐的仅单目估计为欠重建区域的密集化提供候选几何。最终,RGB-D联合优化在稀疏视图监督下提升了外观保真度与几何一致性。在全部9个数据集-视图设置中,D^2-4DGS实现了最高的峰值信噪比(PSNR),在各设置中较最佳对比方法平均提升1.33分贝(dB)。

英文摘要

Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D$^2$-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D$^2$-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑