arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21849cs.CV

GaussVid:结合3D感知视频扩散先验的稀疏视图高斯溅射

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

Xinhui Liu, Can Wang, Wei Jiang, Wei Wang, Dong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出结合3D感知视频扩散先验的GaussVid框架,构建3DGS视频数据集并引入相机条件几何先验,提升稀疏视图3DGS重建的像素、结构保真度与多视图一致性。

中文摘要 AI 辅助

3D高斯溅射(3DGS)在新视角合成领域已取得显著成功,但稀疏视图下的重建结果常存在明显伪影。尽管近期视频扩散模型为3DGS重建提供了强大的时空先验,直接对其进行重建微调效果不佳,因为这些模型缺乏对底层多相机几何结构的感知,会导致多视图不一致问题。本研究提出一种新型3D感知视频重建框架,旨在提升稀疏3DGS重建的质量。具体而言,我们构建了大规模3DGS视频数据集以支持针对性微调。为弥合2D视频生成与3D多视图约束之间的差距,我们引入了相机条件几何先验。通过将首帧和末帧作为边界锚点并编码对应的相机关系,我们明确地将空间结构注入视频生成流程。这种边界锚定、相机感知的先验引导网络进行基于几何的重建,确保跨视角保持一致性。大量实验表明,在基于视频先验的重建方法中,我们的方法在像素级和结构级保真度(PSNR/SSIM)上达到最优,且提升了多视图一致性,同时在感知质量(LPIPS)上保持竞争力。

英文摘要

3D Gaussian Splatting (3DGS) has achieved remarkable success in novel view synthesis; however, reconstructions under sparse views often exhibit noticeable artifacts. While recent video diffusion models provide strong spatio-temporal priors for 3DGS restoration, directly fine-tuning them for restoration is suboptimal, as they lack awareness of the underlying multi-camera geometry, resulting in multi-view inconsistencies. In this work, we propose a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction. Specifically, we construct a large-scale 3DGS video dataset to enable specialized fine-tuning. To bridge the gap between 2D video generation and 3D multi-view constraints, we introduce a camera-conditioned geometric prior. By using the first and last frames as boundary anchors and encoding the corresponding camera relationships, we explicitly inject spatial structure into the video generation pipeline. This boundary-anchored, camera-aware prior guides the network toward geometrically grounded restoration that remains coherent across viewpoints. Extensive experiments show that, among video-prior restoration methods, our approach attains the best pixel- and structure-level fidelity (PSNR/SSIM) and improves multi-view consistency, while remaining competitive in perceptual quality (LPIPS).

发表机构

  • The University of Hong Kong(香港大学)
  • School of Computing and Data Science, The University of Hong Kong(香港大学计算与数据科学学院)
  • Futurewei Technologies Inc(华为主流技术公司)

机构由 AI 辅助整理,请以论文原文为准。

↑