arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Scenix:基于可执行场景程序的稀疏视图3D场景重建

Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

Kai Li, Lutao Jiang, Zhenyang Li, Jiayu Dong, Jierui Zhang, Yingda Yin, Runze Zhang, Kai Yan, Xiaoyang Huang, Keyang Luo, Xin Wang, Xiangyu Zhao, Weikai Chen

arXiv 2608.07012首次发表:更新:

AI 中文总结

本文提出Scenix框架,通过可执行场景程序实现稀疏视图3D室内场景重建,构建了含约11万场景的数据集,引入观测一致性监督,在多类案例上完成相关评估。

AI 中文摘要

从少量未校准的RGB视图合成结构化且可编辑的3D室内场景,不仅需要生成高质量的单个资产,还需推断房间结构、关联不完整观测中的物体,并恢复全局一致的空间配置。此前方法主要聚焦于文本输入的3D场景生成,或需要带额外先验的连续视觉输入,例如人工标注的掩码或精确的3D布局,这些方法人力成本高,难以应用于一般场景。本文提出Scenix,一个通过可执行场景程序实现的稀疏视图3D场景重建框架,可执行场景程序是一种结构化表示,能直接实例化为可编辑的3D场景。给定稀疏视图,Scenix通过感知驱动的资产实例化和闭环空间细化来预测可执行场景程序。为支持该任务,本文构建了包含约11万个合成及真实室内场景的数据集,该数据集具备多视图图像、房间结构、以物体为中心的描述以及度量空间标注。本文还引入了观测一致性监督,使每个目标场景与其输入视图中的视觉证据对齐。在保留的XScene场景、真实室内图像以及分布外的SpatialGen案例上进行的实验,对结构化场景预测、物体 grounding(关联)和空间细化进行了评估。

英文摘要

Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.

Comments4 figures 5 table 9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑