arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

4DGS-Fixer:基于视频扩散先验的迭代细化生成式稀疏视角4D高斯泼溅

4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

Haitao Huang, Shenghao Zhao, Boyuan Tian, Shin-Fang Chng, Songlin Yang, Sheila Lim, Huangying Zhan, Yi Xu, Anyi Rao, Frank Guan

arXiv 2609.21176首次发表:更新:

发表机构

Goertek Alpha Labs; Singapore Institute of Technology; The Hong Kong University of Science and Technology(歌尔阿尔法实验室; 新加坡理工大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稀疏视角动态场景合成中几何初始化差和不适定问题,提出基于视频扩散先验的迭代细化框架,通过深度融合初始化与伪监督细化,实现近2 dB PSNR提升。

AI 中文摘要

本文解决了从稀疏视角视频进行动态场景合成的挑战。现有方法采用几何先验、自适应优化或密度控制策略来改进稀疏观测下的4D高斯建模。然而,它们无法从根本上解决由观测不足和场景信息缺失引起的不适定问题。此外,稀疏视角4D高斯泼溅(4DGS)常常遭受较差的几何初始化:仅凭少量输入视角,COLMAP通常重建出稀疏且不完整的点云,留下大片场景区域缺乏足够的高斯支撑,使得这些区域难以通过后续优化恢复。为解决这些局限,我们提出了一种基于视频扩散模型的新型迭代细化框架,以提高动态4D场景的完整性和一致性。具体而言,我们首先估计多视角深度图并将其融合为密集点云,为动态4DGS表示提供更完整的几何初始化。然后,我们采用预训练的视频恢复模型来细化沿新颖相机轨迹在不同时间步渲染的序列。恢复后的序列作为伪监督,用于正则化并迭代细化4DGS表示。在广泛使用的基准数据集上的实验表明,我们的方法大幅优于现有基线,相比之前最佳方法实现了近2 dB的PSNR提升。

英文摘要

This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.

CommentsAccepted to SIGGRAPH Asia TC

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑