arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.23478cs.CV

Video-R2: 通过强化一致性和基于现实的推理提升多模态语言模型

Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models

Muhammad Maaz, Hanoona Rasheed, Fahad Shahbaz Khan, Salman Khan

首次发表 更新
浏览论文内容

中文总结 AI 辅助

Video-R2通过强化学习提升多模态模型在视频推理中的时间对齐和逻辑一致性,提高准确性和可信度。

中文摘要 AI 辅助

动态视觉内容的推理仍然是多模态大语言模型的核心挑战。最近的思考模型为可解释性生成明确的推理轨迹;然而,其推理往往看似合理,但逻辑上不一致或在视觉证据上缺乏依据。我们通过两个诊断指标识别并正式化这些问题:Think Answer Consistency (TAC),衡量推理与答案的一致性,以及Video Attention Score (VAS),捕捉推理依赖视觉还是文本线索的程度。在11个视频推理基准上的分析显示,当前模型严重依赖语言先验而非视觉内容。为解决这一问题,我们提出了一种强化学习方法,以增强时间精度和推理一致性。我们的方法结合了时间戳感知的监督微调与由新型Temporal Alignment Reward (TAR)引导的Group Relative Policy Optimization (GRPO)。这一双阶段训练后阶段鼓励时间对齐和因果连贯的视频推理。所得到的模型Video R2在多个基准上实现了更高的TAC、VAS和准确性,证明了时间对齐和推理一致性改进导致更准确和可信的视频理解。代码:https://github.com/mbzuai-oryx/Video-R2

英文摘要

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretability; however, their reasoning often appears convincing while being logically inconsistent or weakly grounded in visual evidence. We identify and formalize these issues through two diagnostic metrics: Think Answer Consistency (TAC), which measures the alignment between reasoning and answers, and Video Attention Score (VAS), which captures the extent to which reasoning depends on visual versus textual cues. Analysis across 11 video reasoning benchmarks shows that current models rely heavily on linguistic priors rather than visual content. To address this, we propose a reinforcement learning approach that enhances both temporal precision and reasoning consistency. Our approach combines timestamp aware supervised fine tuning with Group Relative Policy Optimization (GRPO) guided by a novel Temporal Alignment Reward (TAR). This dual step post training stage encourages temporally aligned and causally coherent video reasoning. The resulting model, Video R2, achieves consistently higher TAC, VAS, and accuracy across multiple benchmarks, demonstrating that improvements in temporal alignment and reasoning coherence lead to more accurate and trustworthy video understanding. Code: https://github.com/mbzuai-oryx/Video-R2

发表机构

  • Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
  • Linköping University(林肯皮大学)
  • Australian National University(澳大利亚国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑