发表机构
TU Wien; Fondazione Bruno Kessler; Imperial College London(维也纳工业大学; 布鲁诺·凯斯勒基金会; 伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何弥合几何基础模型与3D场景生成模型的范式差异,通过在VGGT潜在空间进行黎曼流匹配,利用其3D先验,不依赖明确下游表示,在多个数据集上取得良好性能,确立该方法为3D生成的可行范式。
AI 中文摘要
几何基础模型,如视觉几何基础变压器(VGGT),可从未姿态图像中提供强大的3D先验。然而,此类模型仅以前馈、确定性方式运行,无法生成超出输入视图直接支持的合理几何。另一方面,3D场景生成模型必须依赖强大的几何先验,才能从稀疏输入中产生连贯输出。我们通过在VGGT的潜在空间中直接进行流匹配来弥合这两种范式,利用其学习到的3D先验,而不依赖任何明确的下游表示。这需要尊重潜在几何,因为VGGT令牌占据高维超球体的乘积,标准欧几里得流匹配在此失效。我们在与VGGT多尺度编码器对齐的四个超球体的乘积流形上定义了一个黎曼流匹配框架,使生成的令牌保持在冻结解码头所需的有效数据流形上。在RealEstate10K、ScanNet++和ETH3D上,我们的方法在逐视图外观和聚合3D几何方面均取得了优于近期场景生成基线的强大性能,确立了几何基础模型上的潜在空间流匹配作为3D生成的可行范式。
英文摘要
Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support. Generative models for 3D scenes, on the other hand, must rely on strong geometric priors to produce coherent outputs from sparse inputs. We bridge these two paradigms by performing flow matching directly in VGGT's latent space, leveraging its learned 3D priors without committing to any explicit downstream representation such as Gaussians, meshes, or video-VAE latents. This requires respecting the latent geometry: VGGT tokens occupy a product of high-dimensional hyperspheres on which standard Euclidean flow matching fails. We address this with a Riemannian Flow Matching framework defined on a product manifold of four hyperspheres, aligned with VGGT's multi-scale encoder, which keeps generated tokens on the valid data manifold required by the frozen decoding heads. On RealEstate10K, ScanNet++ and ETH3D, our method achieves strong performance against recent scene generation baselines in both per-view appearance and aggregated 3D geometry, establishing latent-space flow matching on geometric foundation models as a viable paradigm for 3D generation. The project page can be found $\href{https://lisaweijler.github.io/geometry-grounded-rfm/}{\text{here}}$.