ProDyGS:来自单个静态单目相机的动态高斯泼溅
ProDyGS: Dynamic Gaussian Splatting from a Single Static Monocular Camera
浏览论文内容
中文总结 AI 辅助
ProDyGS提出一种动态3D高斯泼溅框架,利用单目深度估计生成合成多视角监督,从单个静态相机视频实现高质量新视角合成,在DyNeRF数据集上达到最先进性能。
中文摘要 AI 辅助
我们提出了ProDyGS,一种新颖的动态3D高斯泼溅框架,用于从单个静态相机拍摄的视频中进行高质量的新视角合成。现有方法依赖多视角设置或显著的相机运动来提供几何约束,而我们的方法则解决了完全缺乏多视角监督这一具有挑战性的场景。我们通过深度引导的代理图像合成来生成合成的多视角监督,从而克服了这一限制。具体来说,我们使用基础单目深度网络估计时间上一致的深度图,然后构建3D高斯表示,从任意视角生成代理图像。一个形变网络通过利用这种增强监督来扭曲规范高斯,学习时间动态。在DyNeRF数据集上的实验表明,我们的方法在仅需单目深度估计作为外部监督的情况下达到了最先进的性能,优于依赖场景流等更强先验的方法。
英文摘要
We present ProDyGS, a novel dynamic 3D Gaussian Splatting framework for high-quality novel view synthesis from videos captured by a single static camera. While existing methods rely on multi-view setups or significant camera motion for geometric constraints, our approach addresses the challenging scenario where multi-view supervision is completely absent. We overcome this limitation by generating synthetic multi-view supervision through depth-guided proxy image synthesis. Specifically, we estimate temporally consistent depth maps using foundational monocular depth networks, then construct 3D Gaussian representations that generate proxy images from arbitrary viewpoints. A deformation network learns temporal dynamics by warping canonical Gaussians using this augmented supervision. Experiments on the DyNeRF dataset demonstrate that our method achieves state-of-the-art performance while requiring only monocular depth estimation as external supervision, outperforming approaches that rely on stronger priors such as scene flow.