arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03931cs.CVcs.LG

用于从多视图图像生成场景的稀疏自回归建模

Sparse auto-regressive modeling for scene generation from multi-view images

  • NAVER LABS Europe(NAVER欧洲实验室)
  • NAVER LABS(NAVER实验室)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel, Wonjune Cho, Bardienus Pieter Duisterhof, Vincent Leroy, Jerome Revaud

AI总结:

针对从多视图图像生成完整3D场景的挑战,提出无需真实3D数据监督的SPAR3S模型,通过掩码自回归Transformer实现高效空间一致的场景补全,在合成室内场景及RealEstate10k数据集上表现优于现有方法。

AI中文摘要:

从稀疏、无约束的视图生成完整的3D场景是3D视觉领域的基础挑战,该任务需要在观测内容之外进行推理,同时保持计算上的可处理性。现有的前馈重建方法固有地局限于输入图像中可见的内容,而3D生成建模则因密集体积表示的高计算成本以及大规模3D监督数据的稀缺性而受到阻碍。我们引入SPAR3S,这是一种用于条件场景补全的稀疏体素对齐3D潜在生成模型,无需真实3D数据进行监督。我们的核心见解是将3D场景生成建模为结构化、紧凑的体素对齐3D潜在空间,其中仅表示已占据的体素。我们通过可微分3D高斯溅射技术,利用光度监督直接从多视图图像学习该稀疏潜在空间。给定从稀疏输入视图编码的部分观测体素集,场景补全可简化为预测缺失的潜在标记及其在体素网格中的空间支撑。为此,我们训练了一个掩码自回归Transformer,该模型联合建模体素占据和潜在标记值,实现未观测区域的高效且空间一致的生成。我们在合成室内场景上验证了该方法的有效性,取得了比现有工作更高的新视图质量。我们进一步在RealEstate10k数据集上验证了其泛化能力,凸显了其对真实世界数据的适用性。

英文摘要:

Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.

补充信息

↑