运动即提示:通过运动引导的跨帧视觉提示增强多模态大语言模型的运动推理能力
Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting
浏览论文内容
中文总结 AI 辅助
本文针对多模态大语言模型处理视频时丢失帧间关键运动信息的问题,提出Motion-as-Prompt框架,通过标记轨迹增强视觉提示,在CLEVRER等数据集上提升GPT-5.5的运动推理准确率且不影响非运动理解
中文摘要 AI 辅助
以运动为核心的视频推理是机器人操纵、自主导航等交互应用的基础。然而,多模态大语言模型(MLLMs)通常通过稀疏均匀采样处理视频,以控制视觉 token 和注意力成本,该策略可能会丢弃采样帧之间的关键过渡,限制对物体运动、碰撞及因果交互的推理。为缓解该问题,本文提出Motion-as-Prompt(MaP,运动即提示),一种轨迹引导的跨帧视觉提示框架。MaP 恢复密集点轨迹,选择具有运动信息的帧,并将连续采样帧之间积累的轨迹直接标记在视觉输入上,使原本隐藏的位移、方向变化和交互对冻结的 MLLMs 可见。在 CLEVRER 和 Something-Something-v2 上的实验表明,MaP 可一致提升平均运动推理准确率,使 GPT-5.5 的准确率分别提升 4.2% 和 8.9%。值得注意的是,这些改进在不降低非运动理解能力的情况下获得,凸显了 MaP 的鲁棒性。上述结果表明,MaP 为增强以运动为核心的视频推理提供了一种简单有效的解决方案,无需模型训练或架构修改。项目页面:this https URL
英文摘要
Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs) typically process videos through sparse uniform sampling to control visual-token and attention costs. This strategy may discard critical transitions between sampled frames, limiting reasoning about object movement, collisions, and causal interactions. To mitigate this issue, we propose Motion-as-Prompt (MaP), a track-guided cross-frame visual prompting framework. MaP recovers dense point trajectories, selects motion-informative frames, and marks the trajectories accumulated between consecutive sampled frames directly onto the visual inputs, making otherwise hidden displacement, direction changes, and interactions observable to frozen MLLMs. Experiments on CLEVRER and Something-Something-v2 show that MaP consistently improves average motion-reasoning accuracy, yielding gains of 4.2% and 8.9% for GPT-5.5, respectively. Notably, these improvements are obtained without degrading non-motion understanding, highlighting the robustness of MaP. These results demonstrate that MaP provides a simple and effective solution for enhancing motion-centric video reasoning without model training or architectural modification. Project page:https://github.com/SunVictor23/MaP.