发表机构
KAIST AI(韩国科学技术院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tetris3D通过显式条件化周围物体几何与物理关系,实现单图像三维场景中物体形状和姿态的连贯生成,并引入120万场景的ComOb数据集,在生成质量和物理稳定性上达到最优。
AI 中文摘要
我们提出了Tetris3D,一种用于单图像三维场景重建的生成框架,能够恢复在物理和几何上作为一个场景连贯的物体。现有方法通常独立生成物体或隐式地耦合它们,为确保相互作用的相邻物体之间细粒度的空间兼容性提供的指导有限。为了解决这一问题,我们明确地将每个物体的生成条件设定为周围物体的几何形状及其物理关系,引导其形状和姿态在场景中保持几何和物理上的合理性。此外,我们引入了ComOb,一个基于物理模拟的数据集,包含120万场景,涵盖不同物体类别间的物理交互,并带有每个物体的网格和成对物理关系标注。在合成场景和真实世界场景上的全面实验表明,即使交互区域被遮挡,Tetris3D也能恢复连贯的物体形状和姿态,并在生成质量和物理稳定性方面达到了最先进的性能。
英文摘要
We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.
CommentsProject page: https://cvlab-kaist.github.io/Tetris3D/