arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.14935cs.CV

Video-CoE:通过事件链强化视频事件预测

Video-CoE: Reinforcing Video Event Prediction via Chain of Events

  • AMAP, Alibaba Group(阿里妈妈,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Qile Su, Jing Tang, Rui Chen, Lei Sun, Xiangxiang Chu

更新

中文总结 AI 辅助

本文提出Video-CoE方法,通过构建时间事件链提升视频事件预测的准确性,实验表明其优于现有主流大语言模型。

中文摘要 AI 辅助

尽管在各种视频任务中应用大语言模型(MLLMs)取得了进展,但视频事件预测(VEP)仍然相对研究较少。VEP需要模型对视频进行细粒度的时间建模,并建立视频与未来事件之间的逻辑关系,而当前的MLLMs仍然难以应对。在本文中,我们首先对当前领先的MLLMs在VEP任务上的全面评估,揭示了其预测不准确的原因,包括缺乏对未来事件预测的逻辑推理能力以及对视觉信息利用不足。为了解决这些挑战,我们提出了\textbf{C}hain\textbf{o}f\textbf{E}vents(\textbf{CoE})范式,通过构建时间事件链来隐式地促使MLLM关注视觉内容及视频与未来事件之间的逻辑联系,并通过多种训练协议激励模型的推理能力。在公共基准上的实验结果表明,我们的方法优于领先的开源和商业MLLMs,在VEP任务上建立了新的状态-of-the-art。代码和模型将很快发布。

英文摘要

Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal modeling of videos and establish logical relationships between videos and future events, which current MLLMs still struggle with. In this work, we first present a comprehensive evaluation of current leading MLLMs on the VEP task, revealing the reasons behind their inaccurate predictions, including lack of logical reasoning ability for future events prediction and insufficient utilization of visual information. To address these challenges, we propose \textbf{C}hain \textbf{o}f \textbf{E}vents (\textbf{CoE}) paradigm, which constructs temporal event chains to implicitly enforce MLLM focusing on the visual content and the logical connections between videos and future events, incentivizing model's reasoning capability with multiple training protocols. Experimental results on public benchmarks demonstrate that our method outperforms both leading open-source and commercial MLLMs, establishing a new state-of-the-art on the VEP task. Codes and models will be released soon.

补充信息

↑