AI 中文总结
本文提出一种基于单张图像生成完整3D场景网格的方法,通过自适应分块、显式2D-3D对应及合成户外数据,重新设计物体中心生成器,在室内外场景中均优于现有基线。
AI 中文摘要
单图像场景生成旨在从单张图像生成完整的3D场景网格,包括相机未观察到的表面。虽然预训练的3D物体生成器编码了强大的形状先验,但它们主要针对固定规范体积中的孤立物体设计,且由于户外场景的多样化3D数据相当有限,它们大多聚焦于室内场景。在这项工作中,我们提出了一种方法,重新设计这种以物体为中心的生成器(例如Trellis 2),使其在保留先验的同时适用于室内和室外场景。我们通过以下方式实现:(a) 将场景划分为自适应的块,这些块的大小相对于相机的距离进行缩放;附近的块尺寸较小以保留更精细的细节,而远处的结构(例如建筑物)则由大块覆盖;(b) 通过提升图像特征并让模型感知自由空间、观察表面和未观察区域,使生成器捕获显式的2D-3D对应关系;(c) 合成约4000个户外场景以拓宽训练数据,因为现有场景数据集大多为室内。在Tanks and Temples、ScanNet++和野外图像上的实验表明,我们的方法在几何精度和感知质量方面,在室内和室外场景中均优于所有基线。
英文摘要
Single-image scene generation aims to produce a complete 3D scene mesh from a single image, including surfaces the camera did not observe. While pretrained 3D object generators encode a strong shape prior, they are mainly designed for isolated objects in a fixed canonical volume and focus mostly on indoor scenes, since diverse 3D data for outdoor scenes are quite limited. In this work, we present a method that redesigns such an object-centric generator, e.g., Trellis 2, to work on both indoor and outdoor scenes while retaining its prior. We accomplish this by (a) partitioning the scene into adaptive chunks that scale relative to the distance to the camera; nearby chunks have a smaller size to keep the finer detail, while distant structures, e.g., buildings, are covered by large chunks; (b) making the generator capture explicit 2D-3D correspondence by lifting image features and making the model aware of the free space, observed surface, and unobserved region; (c) synthesizing around 4,000 outdoor scenes to broaden the training data, as existing scene datasets are largely indoor. Experiments on Tanks and Temples, ScanNet++, and in-the-wild images show that our method outperforms all baselines in geometric accuracy and perceptual quality across both indoor and outdoor scenes.
CommentsProject page: https://build-rome.github.io/