arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MotionBlind:探究视频大语言模型中的运动理解错觉

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas

arXiv 2609.09528首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MotionBlind基准证明视频大语言模型无法真正理解运动,即使能识别物体也无法区分速度等物理属性,仅Gemini3.1 Pro部分通过,揭示其作为世界模型感知前端的不足。

AI 中文摘要

视频大语言模型(Video-LLMs)越来越多地被用作世界模型的感知前端,这一角色假设它们能够读取运动信息:物体移动的速度、方向以及被推动的力度。我们证明它们无法做到这一点。一个视频大语言模型可以观看同一人在同一房间的两段视频片段,并能说出两段视频中的每个物体,但仍然无法判断哪段视频中的运动更快。我们引入了MotionBlind,一个基于自录视频的对比基准,用于测试物理上可感知的运动(速度、幅度和方向),这些是世界模型必须预测的变量。每个实例是一对仅在运动上不同的近乎相同的视频片段。每个片段附带两个互补的是/否问题,每个实例共四个问题,模型只有在四个问题全部回答正确时才能得分。我们报告实例准确率(IAcc),其随机猜测基线为6.25%。单帧、外观和仅语言捷径都会使性能降至该基线。MotionBlind补充了最近的TimeBlind基准。我们对六个开源和两个前沿视频大语言模型进行了受控研究,变化条件包括:视频是否存在、帧是否按正确时间顺序显示,以及帧的采样方式(1到24帧,四种选择策略)。开源模型接近6.25%的随机基线,且模型规模增大并无帮助。移除视频会使所有模型的IAcc降至零,打乱帧顺序则使IAcc降至随机水平,因此该任务确实需要按顺序的视频输入。增加帧数或更智能的帧选择都无法缩小差距,因为这些方法改变的是看到的帧,而非是否读取运动。只有Gemini3.1 Pro在整体上通过了该基准,但即使是它也未能通过速度测试。一个无法区分同一动作两种速度的前端,尚不能作为世界模型可靠的监督、奖励或评估来源。

英文摘要

Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show they cannot. A Video-LLM can watch two clips of the same person in the same room, name every object in both, and still fail to say which clip moves faster. We introduce MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion(speed, magnitude, and direction), the variables a world model must predict. Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark. We run a controlled study of six open and two frontier Video-LLMs, varying whether the video is present, whether frames are shown in the correct temporal order, and how frames are sampled (1 to 24 frames, four selection strategies). Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed. A frontend that cannot tell two speeds of the same action apart is not yet a trustworthy source of supervision, reward, or evaluation for a world model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑