arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PASTEL:用于单目4D场景重建的全景对齐

PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction

Yuankun Yang, Yi Wei, Bo Bai, Wenyang Zhou, Li Zhang

arXiv 2609.06099首次发表:更新:

发表机构

School of Data Science, Fudan University; Central Media Technology Institute, Huawei(复旦大学数据科学学院; 华为中央媒体技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PASTEL通过全景对齐将单目视频的4D重建扩展至不可见区域,利用2D方向轨迹规划优化探索,显著提升重建性能并超越现有方法。

AI 中文摘要

从随意拍摄的单目视频中重建4D场景对于虚拟现实(VR)和具身人工智能等应用至关重要。近年来,4D重建和新视角合成方面的进展极大地推动了这一能力的发展。然而,现有的重建方法通常无法恢复超出可见相机范围之外的区域。因此,我们引入了一种新范式,通过将来自单目输入的可见区域重建与超出可观察相机边界的不可见区域生成相结合,实现4D场景合成。我们提出了全景对齐策略性利用生成先验(PASTEL)。具体而言,PASTEL提出了全景场景对齐,这是一种新颖的表示方法,将棘手的3D“不可见区域”探索重新表述为可处理的2D方向轨迹规划。这是通过将视点规划从6自由度搜索减少到具有明确可见性边界的2D方向搜索来实现的。通过在该全景空间中操作,我们的方法策略性地识别出能够最大化超出可观察边界探索同时最小化视点偏差的相机轨迹。实验结果表明,PASTEL不仅能推断出超出输入单目视频可观察边界的合理场景内容,还能显著提升单目4D重建性能。在DyCheck IPhone数据集上,PASTEL在全图像PSNR上比之前最先进的方法高出0.9dB。

英文摘要

Reconstructing 4D scenes from casually captured monocular video is vital for applications in virtual reality (VR) and embodied AI. Recent advances in 4D reconstruction and novel view synthesis have substantially propelled this capability. However, existing reconstruction methods generally cannot recover regions beyond visible camera limits. Consequently, we introduce a new paradigm that achieves 4D scene synthesis by combining visible-region reconstruction from monocular input with invisible-region generation beyond observable camera boundaries. We present Panoramic Alignment for Strategic Exploitation of Generative Priors (PASTEL). Specifically, PASTEL proposes panoramic scene alignment, a novel representation that reformulates the intractable 3D "invisible region" exploration into a tractable 2D directional trajectory planning. This is achieved by reducing the viewpoint planning from 6-DoF search to a 2D directional search with explicit visibility boundaries. By operating within this panoramic space, our method strategically identifies camera trajectories that maximize exploration beyond observable boundaries while minimizing viewpoint deviation. Experimental results show that PASTEL can not only extrapolate plausible scene content beyond the observable boundaries of input monocular videos, but also substantially boost monocular 4D reconstruction performance. PASTEL outperforms the previous state-of-the-art method by 0.9dB in full-image PSNR on the DyCheck IPhone dataset.

CommentsAccepted to ECCV 2026. 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑