发表机构
Tsinghua University; NVIDIA; Stanford University(清华大学; 英伟达; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对空间超感知中多模态模型基准测试局限于合成视频和家庭场景的问题,引入VSI-Super-Wild基准,受人类认知启发探究世界状态三元组,通过大量真实视频问答对测试发现模型不足及失败模式,为空间超感知发展指明方向。
AI 中文摘要
人类能够有效地解析从数小时到数年的连续感官流,构建一个基于空间推理和预测的内部世界模型。为模仿这种能力,空间超感知挑战多模态模型超越语言理解,实现真正的世界建模。然而,其基准测试依赖合成长视频,多限于家庭场景,对现实世界的连续性和多样性探索不足。为此,我们引入VSI-Super-Wild,一个用于评估野外不同场景中长时间空间超感知的大规模基准。受人类构建经验的认知研究启发,我们系统探究世界状态的三元组:智能体、物体和环境。VSI-Super-Wild包含6980个人工验证的问答对,源自442个跨越8个场景类别的真实世界视频。结果显示,尽管静态图像理解有进展,但模型在需要连贯跟踪世界状态随时间变化的任务上持续失败。我们刻画了性能如何随世界状态复杂性和时间跨度下降,并诊断出四种失败模式。这种分类揭示模型缺乏将物体、智能体和环境绑定成统一空间世界模型的机制,这一根本差距为空间超感知指明了前进方向。
英文摘要
Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and prediction. To mimic this capacity, spatial supersensing challenges multimodal models to move beyond linguistic understanding toward true world modeling. However, their benchmark relies on synthetic long videos, formed by concatenating random short clips, and is mostly limited to household scenes, leaving real-world continuity and diversity underexplored. To address the gap, we introduce $\textbf{VSI-Super-Wild}$, a large-scale benchmark for evaluating spatial supersensing over long temporal horizons in diverse in-the-wild scenes. Notably, inspired by cognitive studies on how humans structure experience, we systematically probe the full triad of world state: the agent (observer), objects (scene items), and the environment (places and global layout). In total, VSI-Super-Wild contains $\textbf{6,980}$ human-verified question-answer pairs derived from $\textbf{442}$ real-world videos spanning 8 scene categories, including long-form recordings exceeding 4 hours. Results on VSI-Super-Wild expose a fundamental disconnect: despite advances in static image understanding, models consistently fail at tasks that require coherent world-state tracking over time. We characterize how performance degrades with world-state complexity and temporal horizon, and diagnose four failure modes: spatial collapse, semantic shortcuts, insufficient update, and instance confusion. This taxonomy reveals that models lack mechanisms to bind objects, agents, and environments into a unified spatial world model, a fundamental gap that defines the path forward for spatial supersensing.
CommentsAccepted to ECCV 2026. Project page: https://vsi-super-wild.github.io/