AI 中文总结
本文提出Scenix框架,通过可执行场景程序实现稀疏视图3D室内场景重建,构建了含约11万场景的数据集,引入观测一致性监督,在多类案例上完成相关评估。
AI 中文摘要
从少量未校准的RGB视图合成结构化且可编辑的3D室内场景,不仅需要生成高质量的单个资产,还需推断房间结构、关联不完整观测中的物体,并恢复全局一致的空间配置。此前方法主要聚焦于文本输入的3D场景生成,或需要带额外先验的连续视觉输入,例如人工标注的掩码或精确的3D布局,这些方法人力成本高,难以应用于一般场景。本文提出Scenix,一个通过可执行场景程序实现的稀疏视图3D场景重建框架,可执行场景程序是一种结构化表示,能直接实例化为可编辑的3D场景。给定稀疏视图,Scenix通过感知驱动的资产实例化和闭环空间细化来预测可执行场景程序。为支持该任务,本文构建了包含约11万个合成及真实室内场景的数据集,该数据集具备多视图图像、房间结构、以物体为中心的描述以及度量空间标注。本文还引入了观测一致性监督,使每个目标场景与其输入视图中的视觉证据对齐。在保留的XScene场景、真实室内图像以及分布外的SpatialGen案例上进行的实验,对结构化场景预测、物体 grounding(关联)和空间细化进行了评估。
英文摘要
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.
Comments4 figures 5 table 9 pages