arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21030cs.CVcs.CLcs.LG

COMET:面向视频多模态大语言模型的对比运动增强时序推理

COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

Chenghua Zhu, Zhaolu Kang, Qifan Shi, Siyan Wu, Kehan Jiang, Lei Wei, Lianyu Hu, Guangyuan Dong, Mingbo Yang, Rui Lu, Guibo Luo

首次发表
浏览论文内容

中文总结 AI 辅助

COMET是一种视频多模态大语言模型时序增强框架,通过显式时序表示等方法,在Qwen3-VL-8B等模型的动作、时序推理任务上取得性能提升,且可跨模型家族通用。

中文摘要 AI 辅助

视频多模态大语言模型虽已取得显著进展,但细粒度的运动-时序理解仍较薄弱。核心瓶颈不仅在于帧采样稀疏,还缺乏完整的时序建模流水线,无法显式表示帧间变化、实现外观-运动交互,也难以优化时序方向敏感性。我们提出COMET,一种基于时序的框架,通过显式时序表示、外观-运动融合和方向感知优化,系统性增强视频多模态大语言模型。架构上,COMET引入基于泰勒帧差的时序运动分支,通过时序注意力偏置增强的交叉注意力将其运动证据注入外观流;优化阶段,COMET结合时序先验蒸馏与前向-反向TC-GRPO,将时序顺序转化为直接学习信号,强化模型对时序运动分支编码的方向运动模式的利用。该方法实现了一致的整体改进,且具有显著的运动-时序偏向:在Qwen3-VL-8B上,动作中心任务(STAR、SSv2)平均提升4.9%,时序推理任务(NExT-QA、CLEVRER、LLaVA-178K)较BL-GRPO提升2.1%,静态感知任务(PerceptionTest)保持相当性能;相同增益模式可迁移至InternVL2.5-8B,表明COMET可跨模型家族通用。

英文摘要

Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COMET introduces a temporal motion branch built on Taylor frame differences and injects its motion evidence into the appearance stream via temporal attention bias-enhanced cross-attention. For optimization, COMET combines temporal prior distillation with a forward-reverse TC-GRPO stage that turns temporal order into a direct learning signal and strengthens the model's use of directional motion patterns encoded by the temporal motion branch. The method achieves consistent overall improvements with a pronounced motion-temporal bias: on Qwen3-VL-8B, action-centric tasks (STAR, SSv2) improve by 4.9% on average, temporal reasoning tasks (NExT-QA, CLEVRER, LLaVA-178K) by 2.1% over BL-GRPO, while static perception tasks (PerceptionTest) remain on par. The same gain pattern also transfers to InternVL2.5-8B, indicating that COMET generalizes across model families.

发表机构

  • Peking University(北京大学)
  • Nanyang Technological University(南洋理工大学)
  • National University of Singapore(新加坡国立大学)
  • Sun Yat-Sen University(中山大学)
  • Pingan Technology(平安科技)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑