arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.01460cs.CV

通过结构化奖励强化视频MLLMs的一致性

Reinforcing Consistency in Video MLLMs with Structured Rewards

  • Rutgers University(罗格斯大学)
  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

Yihao Quan, Zeru Shi, Jinman Zhao, Ruixiang Tang

中文总结 AI 辅助

研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在视频理解上取得显著进展,但输出常存在视觉和时间 grounding 不足的问题。本文通过组合一致性审计,分解caption为事实和时间主张,发现正确高层预测缺乏可靠底层证据。传统句子级监督不足,强化学习中句子级奖励过于粗略。本文提出结构化奖励,包含场景图奖励、时间奖励和视频基础VQA奖励,提升视频理解效果。

英文摘要

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, seemingly plausible outputs often suffer from poor visual and temporal grounding: a model may fabricate object existence, assign incorrect attributes, or collapse repeated events while still producing a globally reasonable caption or answer. We study this failure mode through a compositional consistency audit that decomposes a caption into supporting factual and temporal claims, investigating whether a correct high-level prediction is actually backed by valid lower-level evidence. Our top-down audit reveals that even correct root relational claims often lack reliable attribute and existence support. This indicates that standard sentence-level supervision is a weak proxy for faithful video understanding. Furthermore, when turning to reinforcement learning (RL) for better alignment, standard sentence-level rewards often prove too coarse to accurately localize specific grounding failures. To address this, we replace generic sentence-level rewards with a structured reward built from factual and temporal units. Our training objective integrates three complementary components: (1) an instance-aware scene-graph reward for factual objects, attributes, and relations; (2) a temporal reward for event ordering and repetition; and (3) a video-grounded VQA reward for hierarchical self-verification. Across temporal, general video understanding, and hallucination-oriented benchmarks, this objective yields consistent gains on open-source backbones. These results suggest that structured reward shaping is a practical route to more faithful video understanding.

补充信息

↑