COMET:面向视频多模态大语言模型的对比运动增强时序推理
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
COMET是一种视频多模态大语言模型时序增强框架,通过显式时序表示等方法,在Qwen3-VL-8B等模型的动作、时序推理任务上取得性能提升,且可跨模型家族通用。
中文摘要 AI 辅助
视频多模态大语言模型虽已取得显著进展,但细粒度的运动-时序理解仍较薄弱。核心瓶颈不仅在于帧采样稀疏,还缺乏完整的时序建模流水线,无法显式表示帧间变化、实现外观-运动交互,也难以优化时序方向敏感性。我们提出COMET,一种基于时序的框架,通过显式时序表示、外观-运动融合和方向感知优化,系统性增强视频多模态大语言模型。架构上,COMET引入基于泰勒帧差的时序运动分支,通过时序注意力偏置增强的交叉注意力将其运动证据注入外观流;优化阶段,COMET结合时序先验蒸馏与前向-反向TC-GRPO,将时序顺序转化为直接学习信号,强化模型对时序运动分支编码的方向运动模式的利用。该方法实现了一致的整体改进,且具有显著的运动-时序偏向:在Qwen3-VL-8B上,动作中心任务(STAR、SSv2)平均提升4.9%,时序推理任务(NExT-QA、CLEVRER、LLaVA-178K)较BL-GRPO提升2.1%,静态感知任务(PerceptionTest)保持相当性能;相同增益模式可迁移至InternVL2.5-8B,表明COMET可跨模型家族通用。
英文摘要
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COMET introduces a temporal motion branch built on Taylor frame differences and injects its motion evidence into the appearance stream via temporal attention bias-enhanced cross-attention. For optimization, COMET combines temporal prior distillation with a forward-reverse TC-GRPO stage that turns temporal order into a direct learning signal and strengthens the model's use of directional motion patterns encoded by the temporal motion branch. The method achieves consistent overall improvements with a pronounced motion-temporal bias: on Qwen3-VL-8B, action-centric tasks (STAR, SSv2) improve by 4.9% on average, temporal reasoning tasks (NExT-QA, CLEVRER, LLaVA-178K) by 2.1% over BL-GRPO, while static perception tasks (PerceptionTest) remain on par. The same gain pattern also transfers to InternVL2.5-8B, indicating that COMET generalizes across model families.
发表机构
- Peking University(北京大学)
- Nanyang Technological University(南洋理工大学)
- National University of Singapore(新加坡国立大学)
- Sun Yat-Sen University(中山大学)
- Pingan Technology(平安科技)
机构由 AI 辅助整理,请以论文原文为准。