Mira-Scene:用于生成式3D场景的像素对齐布局
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene Reconstruction
浏览论文内容
中文总结 AI 辅助
Mira-Scene提出用密集有界的规范坐标图(CCM)替代稀疏姿态回归,结合多模态扩散Transformer,实现高精度的组合式3D场景重建,显著提升布局准确性。
中文摘要 AI 辅助
单图像3D物体生成现在可以产生高保真资产,但将它们准确放置到连贯的场景布局中仍然是一个开放的挑战。一个核心困难在于物体布局的表示方式。整体方法将放置吸收到场景级生成过程中,牺牲了物体级细节。组合方法通过将几何与布局解耦来保持物体保真度,但通常将布局参数化为稀疏、无界的姿态变量,这些变量难以学习,并且在稀缺的场景级数据下泛化能力差。我们提出了Mira-Scene,一个组合式3D场景重建框架,用密集、有界的对应关系恢复取代稀疏姿态回归。其核心是规范坐标图(CCM),一个像素对齐的场,将每个可见物体像素映射到物体有界规范空间中的表面坐标。当与来自单目几何估计的场景空间点云图(PCM)配对时,CCM产生密集的规范到场景对应关系,通过鲁棒的几何对齐从中恢复物体变换。由于CCM在有界规范空间中操作,它提供了一个稳定的预测目标,可以从可扩展的物体级3D数据中训练,而无需场景级布局标注。Mira-Scene进一步引入了一个多模态扩散Transformer,联合生成物体几何和CCM,使用具有共享注意力和位置编码的模态特定专家流来促进几何-布局一致性。在室内、室外、合成和野外场景上的实验表明,Mira-Scene在布局准确性上大幅优于强基线,在3D-IoU上相对提升39.8%,在2D-IoU上相对提升16.5%,相对于SAM3D,仅使用有限的开源训练数据。
英文摘要
Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challenge. A central difficulty lies in how object layout is represented. Holistic methods absorb placement into a scene-level generation process, sacrificing object-level detail. Compositional methods preserve object fidelity by decoupling geometry from layout, but typically parameterize layout as sparse, unbounded pose variables that are difficult to learn and generalize poorly under scarce scene-level supervision. We present Mira-Scene, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery. At its core is the Canonical Coordinate Map (CCM), a pixel-aligned field that maps each visible object pixel to a surface coordinate in the object's bounded canonical space. When paired with a scene-space Point Cloud Map (PCM) from monocular geometry estimation, CCM induces dense canonical-to-scene correspondences from which object transformations are recovered through robust geometric alignment. Because CCM operates in bounded canonical space, it provides a stable prediction target that can be trained from scalable object-level 3D data without requiring scene-level layout annotations. Mira-Scene further introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency. Experiments on indoor, outdoor, synthetic, and in-the-wild scenes show that Mira-Scene substantially outperforms strong baselines in layout accuracy, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.
发表机构
- The University of Hong Kong(香港大学)
- VAST
机构由 AI 辅助整理,请以论文原文为准。