arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SceneReGen:基于单图像的3D场景生成式重建

SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image

Zefan Tian, Yuteng Ye, Yiheng Zhang, Yuhang Yang, Xueqiang Lv, Shizhou Zhang, Le Liu, Di Xu

arXiv 2608.23930首次发表:更新:

发表机构

Huawei; Northwestern Polytechnical University(华为; 西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SceneReGen是基于DiT的生成式重建框架,通过选择性姿态分解解决物体与场景的坐标系鸿沟,在3D-FUTURE数据集上的场景级指标表现最优,可应用于自动驾驶等场景。

AI 中文摘要

单图像3D场景重建需要补全部分观测到的物体,并将它们连贯地放置在与观测对齐的共享场景坐标系中。物体级生成先验具备强大的补全能力,但其中心对齐、尺度归一化的输出通常以物体自身坐标系表示,这在物体生成与场景重建之间形成了根本性的表示鸿沟。我们提出SceneReGen,这是一个生成式重建框架,它将场景重建重新定义为在与观测对齐的共享场景坐标系中生成并组装完整物体资产的过程。SceneReGen通过选择性姿态分解解决生成-重建鸿沟:每个物体的观测方向直接编码在生成的网格中,而平移和尺度则从实例级和全局场景证据中估计。给定场景图像和实例掩码,几何编码器提取密集线索;可学习形状查询对预训练的基于DiT的3D生成器进行条件约束,以生成具有观测方向的完整网格,而位置查询则融合物体和场景特征,将它们组装到共享坐标系中。在3D-FUTURE评估子集上,SceneReGen在评估方法中取得了最佳的场景级CD、场景级F-Score和3D边界框IoU,与最佳物体级CD持平,在物体级F-Score中排名第二。在自动驾驶和具身AI场景中的定性输出进一步表明,以资产为中心的重建在室内家具之外的场景中也具有潜力。

英文摘要

Single-image 3D scene reconstruction must complete partially observed objects and place them coherently in a shared observation-aligned scene frame. Object-level generative priors offer strong completion ability, but their centered, scale-normalized outputs are typically expressed in an object frame, creating a fundamental representation gap between object generation and scene reconstruction. We introduce SceneReGen, a generative reconstruction framework that reinterprets scene reconstruction as the generation and assembly of complete object assets in a shared observation-aligned scene frame. SceneReGen addresses the generation-reconstruction gap through selective pose factorization: each object's observed orientation is encoded directly in the generated mesh, while translation and scale are estimated from instance-level and global scene evidence. Given a scene image and instance masks, a geometry encoder extracts dense cues; learnable shape queries condition a pretrained DiT-based 3D generator to produce complete meshes in their observed orientations, while position queries fuse object and scene features to assemble them in the shared frame. On the 3D-FUTURE evaluation subset, SceneReGen achieves the best scene-level CD, scene-level F-Score, and 3D bounding-box IoU among the evaluated methods, ties the best object-level CD, and ranks second in object-level F-Score. Qualitative outputs in autonomous-driving and embodied-AI scenes further illustrate the potential of asset-centric reconstruction beyond indoor furniture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑