发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; The Chinese University of Hong Kong; The University of Hong Kong; Fudan University; Zhejiang University(上海交通大学; 上海人工智能实验室; 香港中文大学; 香港大学; 复旦大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoVerse在几何潜在空间中结合视频生成先验与全局空间记忆,实现稀疏图像下的世界一致新视角合成,显著提升视觉质量与几何一致性。
AI 中文摘要
从稀疏图像进行新视角合成必须在忠实重建已观测区域与合理补全未见内容之间取得平衡,同时保持跨视角的世界一致性。现有的基于几何的方法能保留已观测场景结构,但往往难以补全未见区域,而视频生成模型虽提供丰富的表观先验,但在顺序视角生成过程中会累积不一致性。我们提出GeoVerse,一个通过在预训练3D基础模型的几何潜在空间内进行生成,并注入来自视频生成模型的表观先验,来合成世界一致新视角的框架。具体而言,GeoVerse从Wan2.2 VACE中提取多级特征,并通过ControlNet风格的适配器将其注入几何潜在扩散模型,从而融入视频学习的表观先验以增强结构补全。为强制跨视角一致性,一个全局空间记忆持续聚合已观测和已合成的内容,通过重投影目标对齐的引导,将后续预测锚定到共享场景表示上。跨多个数据集的广泛实验表明,视觉质量和几何一致性均有提升,在DL3DV上PSNR提高2.23 dB,在Mip-NeRF360上ATE降低32.4%,优于GLD。
英文摘要
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
CommentsProject Page: https://geoverse-nvs.github.io/