arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.21776cs.CV

Video-R1:增强多模态大语言模型中的视频推理能力

Video-R1: Reinforcing Video Reasoning in MLLMs

  • CUHK MMLab(香港中文大学多模态实验室)
  • CUHK (SZ)(香港中文大学(深圳))
  • Tsinghua University(清华大学)
  • UCAS(中国科学院大学)
  • CUHK HCCL(香港中文大学高性能计算实验室)

机构由 AI 辅助整理,请以论文原文为准。

Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue

更新

AI总结:

针对视频推理缺乏时序建模、高质量数据稀缺的问题,研究提出T-GRPO算法并结合图像推理数据构建数据集,推出Video-R1模型,在多个视频基准上性能显著提升,7B版本甚至超越GPT-4o。

AI中文摘要:

受DeepSeek-R1通过基于规则的强化学习(RL)激发推理能力的成功启发,我们提出Video-R1,这是首个系统性探索R1范式以激励多模态大语言模型(MLLMs)中视频推理的尝试。然而,将采用GRPO算法的RL训练直接应用于视频推理存在两大主要挑战:(1)缺乏针对视频推理的时序建模;(2)高质量视频推理数据稀缺。为解决这些问题,我们首先提出T-GRPO算法,该算法鼓励模型利用视频中的时序信息进行推理。此外,我们并未仅依赖视频数据,而是将高质量的图像推理数据纳入训练过程。我们构建了两个数据集:用于有监督微调(SFT)冷启动的Video-R1-CoT-165k,以及用于RL训练的Video-R1-260k,两者均包含图像和视频数据。实验结果表明,Video-R1在VideoMMMU、VSI-Bench等视频推理基准,以及MVBench、TempCompass等通用视频基准上均取得显著提升。值得注意的是,Video-R1-7B在视频空间推理基准VSI-bench上达到37.1%的准确率,超过了商业专有模型GPT-4o。所有代码、模型和数据均发布于:https://github.com/tulerfeng/Video-R1。

英文摘要:

Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly applying RL training with the GRPO algorithm to video reasoning presents two primary challenges: (i) a lack of temporal modeling for video reasoning, and (ii) the scarcity of high-quality video-reasoning data. To address these issues, we first propose the T-GRPO algorithm, which encourages models to utilize temporal information in videos for reasoning. Additionally, instead of relying solely on video data, we incorporate high-quality image-reasoning data into the training process. We have constructed two datasets: Video-R1-CoT-165k for SFT cold start and Video-R1-260k for RL training, both comprising image and video data. Experimental results demonstrate that Video-R1 achieves significant improvements on video reasoning benchmarks such as VideoMMMU and VSI-Bench, as well as on general video benchmarks including MVBench and TempCompass, etc. Notably, Video-R1-7B attains a 37.1% accuracy on video spatial reasoning benchmark VSI-bench, surpassing the commercial proprietary model GPT-4o. All code, models, and data are released in: https://github.com/tulerfeng/Video-R1.

补充信息

↑