GOPI:面向生成的单视图RGB-D室内场景家具插入的3D位姿推理
GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
- School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
- Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
- X-Era AI Lab(X-Era AI实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对单视图RGB-D室内场景家具插入的不确定性问题,提出两阶段框架GOPI,通过数据驱动迭代推理实现合理3D位姿估计,结合几何引导生成策略,在3D放置和图像合成任务上均优于基线方法。
AI中文摘要:
我们研究在室内场景图像中插入新家具的问题。然而,在被遮挡的单视图2D图像平面条件下,插入家具相对于场景的物理尺度无法被唯一确定,仅从图像证据出发,基于物理的家具放置任务是不确定的。因此,我们将该任务重新表述为3D位姿推理与几何引导图像生成的结合,其中估计几何上合理的3D放置对于可靠合成至关重要。为此,我们提出了一个两阶段框架。对于3D放置,我们引入了GOPI,这是一个面向生成的3D位姿推理框架,通过数据驱动的迭代推理解决单视图家具插入的不确定性,生成几何上合理的物体放置。对于图像生成,我们开发了一种几何引导的条件策略,将推断出的3D位姿投影到图像平面作为像素级对齐的约束,确保合成图像与底层3D几何的一致性。实验结果从3D位姿估计和图像合成两个角度验证了所提框架。对于3D放置,GOPI生成的位姿比直接回归和普通基线具有更强的几何可行性和与参考布局的更好一致性;对于图像合成,我们的方法在不同家具尺度下保持了与投影3D几何的对齐,在测试的家具尺度上显示出稳定的投影-生成对齐。
英文摘要:
We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furniture relative to the scene cannot be uniquely determined, making physically grounded furniture placement underdetermined from image evidence alone. We therefore reformulate the task as a combination of 3D pose inference and geometry-guided image generation, where estimating a geometrically plausible 3D placement is essential for reliable synthesis. To this end, we propose a two-stage framework. For 3D placement, we introduce GOPI, a generation-oriented 3D pose inference framework that addresses the underdetermined nature of single-view furniture insertion through data-driven iterative inference, producing geometrically plausible object placements. For image generation, we develop a geometry-guided conditioning strategy that projects the inferred 3D pose into the image plane as a pixel-aligned constraint, enforcing consistency between the synthesized image and the underlying 3D geometry. Experimental results validate the proposed framework from both 3D pose estimation and image synthesis perspectives. For 3D placement, GOPI produces poses with stronger geometric feasibility and better consistency with reference layouts than direct regression and vanilla baselines. For image synthesis, our method preserves alignment with the projected 3D geometry across different furniture scales, showing stable projection-generation alignment across the tested furniture scales.