arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34630cs.CV

Long Time No See: 基准测试视频大语言模型在自我中心视频中的视野外时空推理能力

Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos

Fangzhou Ma, Ivo Alexander Ban, Eren Homburg, Gabriele Goletto, Rémi Pautrat, Mahdi Rad, Chiara Plizzari, Marc Pollefeys

首次发表
浏览论文内容

中文总结 AI 辅助

针对自我中心视频中物体移出视野后的时空推理,提出Beyond3D基准,含9000个问题,评测九个VLM,最佳准确率42.2%,揭示当前模型在此任务上的显著不足。

中文摘要 AI 辅助

现实世界的人工智能系统必须对不再可见的物体进行推理:一个增强现实助手引导用户回到之前使用过的物体,一个家用机器人取回某人放下的物品。这不仅需要回忆物体最后出现的位置,还需要在物体被移动时更新其状态,并在物体离开视野后保留该更新。我们将其称为视野外时空推理。我们引入了Beyond3D,这是第一个在动态自我中心视频中隔离这种能力的VQA基准:每个查询都针对一个已被重新放置且此后离开视野的物体。我们根据HD-EPIC注释创建问题,为每个动态物体构建一个可见性轨迹,利用其3D位置、相机姿态和场景几何来理解每一时刻该物体是可见、被遮挡还是处于视野外。Beyond3D包含来自9名参与者的135个视频中的8种类型的9,000个问题,组织成一个推理链:视觉基础(目标当前是否可观察)、时间基础(它最后可见和最后放置的时间)、场景定位(哪个固定装置锚定该位置)以及3D空间感知(它相对于当前视点或场景中另一物体的位置)。我们基准测试了九个通用和空间专门的VLM。最佳模型达到42.2%,而随机猜测为29.7%,仅文本基线达到31.9%,最大的失败在于恢复物体最后可见的时间,这表明对于当前的VLM来说,跟踪物体在视野外的移动仍然远未解决。

英文摘要

Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved and retaining that update once it leaves view. We refer to this as out-of-sight spatiotemporal reasoning. We introduce Beyond3D, the first VQA benchmark to isolate this ability in dynamic egocentric video: every query targets an object that has been relocated and has since left the field of view. We create our questions from HD-EPIC annotations, building a visibility track for each dynamic object from its 3D position, the camera pose, and the scene geometry to understand at each moment whether it is visible, occluded, or out of view. Beyond3D comprises 9,000 questions in eight types over 135 videos from nine participants, organized as one reasoning chain: visual grounding (is the target observable now), temporal grounding (when it was last visible and last placed), scene localization (which fixture anchors that location), and 3D spatial perception (where it lies relative to the current viewpoint or another object in the scene). We benchmark nine general-purpose and spatially specialized VLMs. The best model reaches 42.2% against 29.7% chance and text-only baselines reaching 31.9%, with the largest failures in recovering when an object was last visible, showing that tracking object movement out of sight remains far from solved for current VLMs.

发表机构

  • Microsoft Spatial AI Lab(微软空间智能实验室)
  • Bocconi University(博科尼大学)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑