发表机构
University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出HARMONY分层智能体推理框架,结合VLM语义推理与几何基础模型,实现从单目图像到高保真3D场景的合成,优于现有重建基线。
AI 中文摘要
组合式3D场景重建近期从两个方向进行了探索:智能体推理提供了对空间关系的语义理解,但缺乏与输入图像的精确对齐;视觉几何基础模型从输入图像预测密集点图,但重建质量有限。因此,从单目图像恢复具有准确物体间关系和高保真重建质量的完整3D场景仍然具有挑战性。本文提出HARMONY,一个利用智能体推理和视觉几何基础的分层思维链框架。给定室内场景图像,从空3D平面图开始,HARMONY首先针对参考图像校准相机以建立语义基础的空间框架,然后使用智能体VLM推理恢复3D房间布局和初始放置顺序。随后,它按分层顺序放置物体,从墙面元素、独立家具到家具上的依赖装饰。我们还对家具使用深度优先遍历,使每次放置都基于先前解析的结构,并通过反思性反馈循环避免错误累积。在VLM每次放置物体后,我们使用点云估计进行基于几何的细化,使渲染图像与输入更好地对齐。HARMONY能生成与参考图像语义一致且感知对齐的3D场景,将单图像组合重建扩展到复杂室内场景图像。在合成和真实世界图像上的实验表明,HARMONY优于评估的重建基线,而与GPT-6 Astra的定性比较显示更忠实的物体排列和更好的场景细节保留。
英文摘要
Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality remains challenging. In this paper, we present HARMONY, a hierarchical chain-of-thought framework that leverages both agentic reasoning and visual geometry foundation. Given an image of an indoor scene, starting from an empty 3D floorplan, HARMONY first calibrates the camera against the reference image to establish a semantically-grounded spatial frame, then uses agentic VLM reasoning to recover the 3D room layout and an initial placement order. It then places the objects in a hierarchical order, from wall-mounted elements, free-standing furniture, to dependent decorations on top of furniture. We also use depth-first traversal for furniture so each placement conditions on previously resolved structure and a reflective feedback loop to avoid error accumulation. After each object placement by VLM, we use the point cloud estimations to perform geometry-based refinement so that the rendered image aligns better with the input. HARMONY can produce 3D scenes that are semantically consistent and perceptually aligned with the reference image, extending single-image compositional reconstruction to complex indoor scene images. Experiments on synthetic and real-world images demonstrate that HARMONY outperforms the evaluated reconstruction baselines, while qualitative comparisons with GPT-6 Astra suggest more faithful object arrangements and better preservation of scene details.
CommentsProject Page: http://cwchenwang.github.io/harmony