发表机构
Peking University; Rabbitpre AI(北京大学; Rabbitpre AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UniWorld-View 是结合显式 3D 指导与视频扩散建模的统一框架,可从单目输入实现可控大基线新视图合成,在基准测试中展现出优异的可控性、几何一致性与视觉保真度。
AI 中文摘要
社交媒体上大量随手拍摄的单目视频和图像为沉浸式内容创作提供了宝贵资源,从这类稀疏观测中生成新视图可极大提升用户体验。然而,当输入覆盖范围极有限时,生成具有精确相机控制的 photorealistic( photorealistic 保留)和几何一致的视图仍具挑战性。基于重建的方法如 NeRF 和 3D Gaussian Splatting(3DGS)在稀疏输入下性能严重下降,且无法显式处理遮挡。生成式方法降低了数据需求,但由于几何指导不准确或隐式,在大基线视图合成上仍存在困难。为克服这些限制,我们提出 UniWorld-View,一个从单目输入进行可控大基线新视图合成的统一框架。UniWorld-View 将显式 3D 指导与生成式扩散建模相结合,以实现精确的相机控制和几何一致的视图生成。几何指导通过感知遮挡的点云渲染策略获得,该策略解决了可见性歧义并为基于扩散的合成提供准确先验。通过将此渲染策略与强大的视频扩散骨干网络耦合,UniWorld-View 即使在极端相机运动和宽基线变化下也能实现高保真新视图生成,还可为下游动态 3DGS 重建提供多视图视频。在 WorldScore 基准和零样本 NVS 基准上的实验证明了 UniWorld-View 在可控性、几何一致性和视觉保真度方面的有效性。
英文摘要
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.
CommentsProject Homepage: https://zhouhyocean.github.io/uniworld-view/ Code: https://github.com/PKU-YuanGroup/UniWorld-View