发表机构
Tsinghua University; Beijing MEET YUAN Co., Ltd(清华大学; 北京美图元科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对六相机稀疏视角动态重建,提出结合区域自适应空间先验、运动一致时序先验与扩散模型生成辅助的4D高斯溅射框架,显著提升重建质量并获赛道第一。
AI 中文摘要
我们提出了一个用于SIGGRAPH Asia 2026体积视频挑战赛稀疏视角赛道的4D高斯溅射框架,该赛道要求仅从六个宽基线相机进行动态场景重建。为了在如此稀疏的视角下实现鲁棒的动态重建,我们的框架整合了三个组成部分。(1)区域自适应空间先验:我们使用前景掩码来引导高斯初始化,并通过掩码投票分别控制动态前景和静态背景的致密化。背景几何通过单目深度进行正则化,并对齐到度量尺度。(2)运动一致的时序先验:我们通过帧插值在中间时刻提供监督,并用估计的光流约束投影高斯运动。(3)生成辅助:我们在最宽的角间隙中放置虚拟相机,并使用基于扩散的模型(以相机姿态为条件)恢复其渲染图像。恢复的图像作为伪监督迭代地纳入训练。在验证集上,我们的框架将全帧PSNR从基线的25.60 dB提升到29.75 dB。在官方测试基准上,它实现了30.04 dB的全帧PSNR和27.88 dB的前景PSNR,在稀疏视角赛道中总体排名第一。
英文摘要
We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.
Comments4 pages, 5 figures, Accepted to SIGGRAPH Asia 2026 Workshops (SA Workshops '26)