arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.17555cs.CV

GraphThinker: 通过事件图思维强化时间感知的视频推理

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Zixu Cheng, Da Li, Jian Hu, Yuhang Zang, Ziquan Liu, Shaogang Gong, Wei Li

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出GraphThinker,通过构建结构化事件表示并强化视觉 grounding,减少视频推理中的时间幻觉。在RexTime和VidHalluc数据集上取得显著提升。

中文摘要 AI 辅助

视频推理需要对视频中对象和事件的时序依赖及事件级关系进行细粒度理解。当前多模态大语言模型(MLLMs)在视频推理中易产生严重的时序幻觉。其根本原因在于视觉-时序 grounding 较弱以及缺乏显式结构来建模事件关系。模型常依赖辅助文本,如密集描述,而非明确将推理锚定在实际视觉证据上。然而,这些文本表示本质上是无结构的,无法提供引导模型推理所需的显式因果约束。在本文中,我们提出GraphThinker,一种强化微调方法,通过构建视频的结构化事件表示并强制视觉 grounding,共同减少推理幻觉。具体而言,我们利用MLLM构建事件基视频场景图(EVSG),捕捉内事件和外事件关系,引导结构化视频推理过程。此外,我们通过引入新颖的视觉注意力奖励在强化微调中解决弱 grounding 问题,鼓励模型主动关注可靠的视觉线索。在RexTime数据集上,GraphThinker在IoU=0.3时的时刻局部化任务上实现超过4%的提升。在VidHalluc数据集上,GraphThinker在减少时间序列幻觉方面比最先进方法提升9.8%,在减少动作幻觉的二元问答任务中提升7.6%。

英文摘要

Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language Models (MLLMs) are prone to severe temporal hallucinations in video reasoning. An underlying cause of these hallucinations is weak visual-temporal grounding and the lack of explicit structure for modelling event relations. Models often rely on auxiliary text, such as dense captions, rather than explicitly anchoring their reasoning in actual visual evidence. However, these textual representations are inherently unstructured and fail to provide explicit causal constraints needed to guide the model's reasoning. In this work, we propose GraphThinker, a reinforcement finetuning method that constructs a structured event representation of a video and enforces visual grounding to jointly reduce reasoning hallucinations. Specifically, we employ an MLLM to construct an Event-based Video Scene Graph (EVSG) that captures both intra- and inter-event relations, guiding a structured video reasoning process. Moreover, we address the weak grounding issue by introducing a novel visual attention reward during reinforcement finetuning that encourages the model to actively attend to reliable visual cues. On the RexTime dataset, GraphThinker achieves an over 4% improvement in IoU=0.3 for moment localisation. On the VidHalluc dataset, GraphThinker achieves a 9.8% improvement in reducing temporal sequence hallucination and a 7.6% gain in Binary QA in reducing action hallucination, compared to the state-of-the-art methods.

发表机构

  • Queen Mary University of London(伦敦玛丽女王大学)
  • Samsung AI Centre Cambridge(剑桥三星人工智能中心)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑