arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.19524cs.CVcs.MA

VideoChat-M1: 通过多智能体强化学习实现视频理解的协作策略规划

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

Boyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan, Chenyun Yu, Yi Huang, Zijun Liu, Yafei Wen, Xiaoxin Chen, Yang Liu, Peng Li, Yali Wang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

VideoChat-M1通过多智能体强化学习实现视频理解的协作策略规划,显著提升复杂视频任务的性能。

中文摘要 AI 辅助

通过利用工具增强的多模态大语言模型(MLLMs),多智能体框架正在推动视频理解的进步。然而,大多数方法采用静态且不可学习的工具调用机制,这限制了发现对于时间或空间复杂视频中必要的多样化线索。为了解决这一挑战,我们提出了一种新的视频理解多智能体系统,即VideoChat-M1。与使用单一或固定策略不同,VideoChat-M1采用了一种独特的协作策略规划(CPP)范式,包含三个关键过程。(1)策略生成:每个智能体生成其独特的工具调用策略,以适应用户的查询;(2)策略执行:每个智能体依次调用相关工具以执行其策略并探索视频内容;(3)策略通信:在策略执行的中间阶段,智能体相互交互以更新各自的策略。通过这种协作框架,所有智能体协同工作,根据同伴提供的上下文洞察动态改进各自的首选策略,以有效响应用户的查询。此外,我们为CPP范式配备了简洁的多智能体强化学习(MARL)方法。因此,策略智能体团队可以联合优化以提升VideoChat-M1的性能,由最终答案奖励和中间协作过程反馈引导。大量实验表明,VideoChat-M1在八个基准测试中四个任务上均取得了SOTA性能。值得注意的是,在LongVideoBench上,我们的方法比SOTA模型Gemini 2.5 Pro高出3.6%,比GPT-4o高出15.6%。

英文摘要

By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. To address this challenge, we propose a novel Multi-agent system for video understanding, namely VideoChat-M1. Instead of using a single or fixed policy, VideoChat-M1 adopts a distinct Collaborative Policy Planning (CPP) paradigm with multiple policy agents, which comprises three key processes. (1) Policy Generation: Each agent generates its unique tool invocation policy tailored to the user's query; (2) Policy Execution: Each agent sequentially invokes relevant tools to execute its policy and explore the video content; (3) Policy Communication: During the intermediate stages of policy execution, agents interact with one another to update their respective policies. Through this collaborative framework, all agents work in tandem, dynamically refining their preferred policies based on contextual insights from peers to effectively respond to the user's query. Moreover, we equip our CPP paradigm with a concise Multi-Agent Reinforcement Learning (MARL) method. Consequently, the team of policy agents can be jointly optimized to enhance VideoChat-M1's performance, guided by both the final answer reward and intermediate collaborative process feedback. Extensive experiments demonstrate that VideoChat-M1 achieves SOTA performance across eight benchmarks spanning four tasks. Notably, on LongVideoBench, our method outperforms the SOTA model Gemini 2.5 pro by 3.6% and GPT-4o by 15.6%.

发表机构

  • Shenzhen Key Lab of Computer Vision and Pattern Recognition(深圳计算机视觉与模式识别重点实验室)
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • VIVO AI Lab(VIVO人工智能实验室)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Shenzhen Campus of Sun Yat-sen University(孙逸仙大学深圳校区)
  • Shanghai Jiao Tong University(上海交通大学)
  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
  • Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(清华大学计算机科学与技术系,人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑