UniFusion:通过统一时空深度对齐的稀疏视角4D重建
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
浏览论文内容
中文总结 AI 辅助
针对稀疏视角视频4D重建中单目深度不一致的问题,提出统一时空深度对齐框架,利用时空神经场联合解决跨视角和时间不一致,无需分割掩码,提升动态高斯溅射重建的精度与一致性。
中文摘要 AI 辅助
在本文中,我们解决了从稀疏视角视频进行4D重建这一具有挑战性的问题。这种设置通常依赖单目深度估计为重建模型提供先验。一个关键挑战源于有限的跨视角重叠和时间变化,使得单目深度预测在视角和时间上不一致。现有方法在分离的阶段中对齐空间和时间维度,需要前景分割掩码,同时未能利用时间线索进行跨视角对齐。与这些方法相反,我们提出了一个统一的时空深度对齐框架,联合解决跨视角和跨时间的不一致性,而无需区分前景/背景。我们的方法将跨视角和时间的深度图表示为时空神经场集合。这种表示不仅实现了快速收敛,而且隐式地捕获了深度图之间的时空相关性,无需依赖外部分割/跟踪模型。我们还提出了多视角深度顺序损失,同时利用经典的尺度和平移不变损失来进一步提高最终深度质量。对齐后的深度用于初始化并监督高斯溅射模型进行4D重建。在Ego-Exo4D和EgoHuman上的实验表明,我们改进的深度对齐显著有益于基于动态高斯溅射的重建方法,在新型时间/视角合成和几何精度/一致性方面表现优异。
英文摘要
In this paper, we address the challenging problem of 4D reconstruction from sparse-view videos. This setup usually relies on monocular depth estimation to provide priors for the reconstruction model. A key challenge arises from limited cross-view overlap and temporal variation, making monocular depth predictions inconsistent across views and time. Existing methods align spatial and temporal dimensions in separate stages, requiring foreground segmentation masks while failing to leverage temporal cues for cross-view alignment. Contrary to these methods, we propose a unified spatial-temporal depth alignment framework that jointly resolves cross-view and cross-time inconsistencies without distinguishing foreground/background. Our method represents depth maps across views and time as a set of spatio-temporal neural fields. This representation not only yields fast convergence, but also captures spatio-temporal correlation among depth maps implicitly, without dependence on external segmentation/tracking models. We also propose a multi-view depth-order loss while leveraging the classic scale-and-shift-invariant loss to further improve the final depth quality. The aligned depths initialize and supervise Gaussian splatting models for 4D reconstruction. Experiments on Ego-Exo4D and EgoHuman demonstrate that our improved depth alignment substantially benefits dynamic Gaussian-splatting-based reconstruction methods for novel-time/view synthesis and geometry accuracy/consistency.
发表机构
- State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能全国重点实验室,北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。