Video4Spatial: 通过基于场景的视频生成实现视觉空间智能
Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- Netflix
- Nanyang Technological University(南洋理工大学)
- University of Oxford(牛津大学)
- Eyeline Studios
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Video4Spatial通过仅使用视频数据训练的生成模型,实现了复杂空间任务的视觉空间智能,展示了视频生成模型在空间推理中的潜力。
AI中文摘要:
我们研究视频生成模型是否能仅通过视觉数据表现出视觉空间智能,这种能力是人类认知的核心。为此,我们提出了Video4Spatial框架,证明仅基于视频场景上下文的视频扩散模型可以完成复杂的空间任务。我们验证了两个任务:场景导航——在遵循摄像机姿态指令的同时保持场景3D几何的一致性,以及物体定位——需要语义定位、指令遵循和规划。这两个任务均使用仅视频输入,不借助深度或姿态等辅助模态。通过简单的但有效的框架设计和数据整理,Video4Spatial展示了从视频上下文中强大的空间理解能力:它能够端到端地规划导航和定位目标物体,遵循摄像机姿态指令的同时保持空间一致性,并泛化到长上下文和非领域环境。总体而言,这些结果推动视频生成模型向通用视觉空间推理迈进。
英文摘要:
We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.