arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VepAgent:通过工具增强强化学习桥接因果转换以进行视频事件预测

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang, Fangfang Lin, Xinyi Long, Yin Wang, Zijian Song, Rui Lan, Shi Qiu, Boyuan Pan, Yang Luo, Yuyin Zhou

arXiv 2610.06293首次发表:更新:

发表机构

Nankai University; Carnegie Mellon University; Xiaohongshu Inc.; Columbia University; Santa Clara University; New York University; University of California, Santa Cruz(南开大学; 卡内基梅隆大学; 小红书公司; 哥伦比亚大学; 圣克拉拉大学; 纽约大学; 加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VepAgent智能体框架,结合因果转换推理与工具增强强化学习,构建思维链数据集和诊断工具库,在FutureBench和NEPBench上实现视频事件预测的最先进性能。

AI 中文摘要

多模态大语言模型(MLLMs)在视频理解方面已展现出显著潜力,然而在应用于视频事件预测(VEP)时,它们对回顾性总结和以文本为中心的先验的依赖往往限制了其桥接未观察到的因果转换的能力。为解决这一问题,我们提出了VepAgent,一个将因果转换推理与工具增强强化学习(RL)相结合的智能体框架,以实现稳健的VEP。与先前从历史依赖中被动投影未来轨迹的方法不同,我们的方法显式建模从终端观察状态到未来事件的逻辑进展。具体而言,我们首先构建了futurebench-4K,一个用于监督微调(SFT)的高质量思维链数据集,通过结构化未观察中间状态的推导,有效弥合了因果逻辑差距。随后,我们开发了一个诊断工具库,集成了状态跟踪、帧检索和区域放大功能,使智能体能够在推理过程中通过外部工具动态增强推理,以恢复缺失的时空证据并解决视觉歧义。此外,我们提出了一种复合奖励机制,联合优化预测准确性、因果连贯性和可靠先验,促使智能体依赖真实的视觉基础而非表面的文本相似性。在FutureBench和NEPBench数据集上的广泛评估表明,我们的方法达到了最先进的性能,显著优于更大的MLLMs,验证了我们智能体化、面向未来的推理范式的实证有效性。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑