arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39590cs.CV

SPOON:面向未标定多视角图像的一致组合式三维场景生成

SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images

Guibiao Liao, Mochu Xiang, Heng Li, Ken Deng, Zijie Wang, Guanbin Li, Ping Tan, Shenghua Gao, Yizhou Yu

首次发表
浏览论文内容

中文总结 AI 辅助

SPOON通过引导-路由-协调范式,利用重建几何协调多视角姿态,实现一致的三维场景生成,在ARSG-110K上显著降低倒角距离。

中文摘要 AI 辅助

组合式三维场景生成旨在从视觉观测中恢复完整的三维物体形状及其空间排列。最近的图像条件三维生成器为生成高质量物体几何提供了强大的先验,使得复杂场景的生成日益可行。因此,一个核心挑战是如何将这些生成的资产在空间上组织成全局一致的场景,同时与多视角观测保持一致。现有方法要么将场景布局与物体生成纠缠在一起,要么从特定视角的观测中分别估计空间位置,其中姿态假设可能在不同视角间保持模糊且不一致,常常导致物体-相机混杂的混乱结果。我们提出了SPOON,一个将多视角组合式三维生成重新定义为场景级、基于几何的姿态推理的框架。SPOON不独立处理特定视角的物体姿态假设,而是通过引导-路由-协调范式,利用重建衍生的多视角几何来协调它们。这逐步将物体姿态和相机配置组织成一致的场景级空间排列。在ARSG-110K和MIDI-3D-Front上的大量实验表明,在不同数量的输入视角下,物体放置和场景组成均取得了一致的改进。在ARSG-110K上,与强基线相比,SPOON将场景级和物体级的倒角距离分别降低了12.7%和17.7%。

英文摘要

Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.

发表机构

  • The University of Hong Kong(香港大学)
  • Shenzhen Loop Area Institute(深圳河套学院)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Sun Yat-sen University(中山大学)
  • TranscEngram

机构由 AI 辅助整理,请以论文原文为准。

↑