arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09985cs.CLcs.CV

VLX-VR:一种智能体感知的视频推理模型

VLX-VR: An Agentic-Aware Video Reasoning Model

Sheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

VLX-VR通过“思考-记忆-观测”循环和强化学习实现智能体感知的视频推理,在MINERVA上取得78.79%的准确率,并展现出跨时长的稳定性。

中文摘要 AI 辅助

真实世界中的视频理解需要整合分布在视频中的视觉、音频、文本和时间证据。然而,许多流程使用固定的视频上下文和单次推理,这限制了在观测不完整、模糊或冲突时自适应的证据获取能力。我们提出了VLX-VR,一种智能体感知的视频推理模型,它在一个由“思考-记忆-观测”循环定义视频推理框架内进行训练。在每一步中,VLX-VR确定所需证据,调用read_memory或write_memory,整合返回的观测结果,并决定是继续还是生成任务输出。我们使用包括视频和智能体轨迹在内的多模态数据训练VLX-VR,通过强化学习来学习证据获取、记忆使用和终止。在MINERVA上,VLX-VR在我们比较的模型中达到了最先进的性能,准确率为78.79%。在原始的三个时长组中,其准确率分别为76.70%、78.73%和80.92%,跨时长准确率方差为2.97平方百分点。在正确回答的样本中,VLX-VR的96.20%推理轨迹与MINERVA参考推理轨迹及其描述的证据一致,而大约75.80%的所有评估样本同时满足答案正确性和这种基于证据的轨迹标准。这些结果表明,VLX-VR在不同时长下表现出强大的性能和广泛稳定的行为,而计数、状态变化、因果推理和空间感知仍然具有挑战性。

英文摘要

Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete, ambiguous, or conflicting. We present VLX-VR, an agentic-aware video reasoning model trained within a video reasoning framework defined by a Think--Memory--Observation loop. At each step, VLX-VR determines the needed evidence, invokes read_memory or write_memory, incorporates the returned Observation, and decides whether to continue or produce the task output. We train VLX-VR with multimodal data, including videos and agent trajectories, using reinforcement learning to learn evidence acquisition, memory use, and termination. On MINERVA, VLX-VR achieves state-of-the-art performance among the models included in our comparison, with 78.79% accuracy. Under the original three duration groups, its accuracies are 76.70%, 78.73%, and 80.92%, with a cross-duration accuracy variance of 2.97~$\mathrm{pp}^2$. On correctly answered samples, 96.20% of VLX-VR's reasoning traces are consistent with the MINERVA reference reasoning traces and the evidence described by them, while approximately 75.80% of all evaluated samples satisfy both answer correctness and this evidence-grounded trace criterion. These results show strong performance and broadly stable behavior across durations, while counting, state changes, causal reasoning, and spatial perception remain challenging.

发表机构

  • Om AI Research(Om AI 研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑