AI 中文总结
VGGT-Diff提出一种几何路由的多视角扩散模型,通过将VGGT的视觉几何潜变量注入视频扩散模型,结合置信度感知路由器与点轨迹残差一致性,在稀疏视角新视角合成中实现插值和外推的先进性能。
AI 中文摘要
我们提出了VGGT-Diff,一种几何路由的多视角扩散模型,用于稀疏视角的新视角合成。现有的新视角合成(NVS)方法面临一个根本性的权衡:基于重建的方法保留观察到的几何结构,但难以合成未见区域,而基于扩散的方法提供强大的生成先验,却依赖于隐式的源到查询对应关系。VGGT-Diff通过将来自VGGT-{\Omega}的视觉几何潜变量路由到预训练的视频扩散模型中,弥合了这两种机制。每个视觉标记与一个3D点和置信度相关联,然后通过一个置信度感知的视觉几何路由器(VGR)转换为查询对齐的潜变量条件,该路由器保留前表面和后表面的证据。这些条件引导联合目标视角去噪,而点轨迹残差一致性(PTRC)沿着可靠的3D轨迹正则化预测的干净残差,提高多视角稳定性。我们进一步引入了鲁棒的几何条件化,将训练时的正则化与推理时的引导相结合,以提高鲁棒性。实验表明,在不同视角难度下的插值和外推任务中,我们的方法达到了具有竞争力或最先进的性能。我们的代码可在https://this https URL获取。
英文摘要
We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.
CommentsProject page: https://chenkangjie1123.github.io/VGGT-Diff, Code at: https://github.com/chenkangjie1123/VGGT-Diff