超越视频:统一视频推理与深度研究的开放世界视频智能体
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
浏览论文内容
中文总结 AI 辅助
本研究提出统一视频推理与深度研究的VideoRover框架,构建相关数据集与基准,其8B-RL模型在无工具直接回答场景性能接近专有模型,优于同工具套件的更大开源模型。
中文摘要 AI 辅助
开放世界视频理解通常要求模型定位稀疏的视觉证据,并获取视频及其参数化记忆中不存在的外部知识。虽然Thinking-with-Videos支持主动时间感知,Deep Research支持多步骤信息搜索,但这两种能力通常是独立开发的。我们提出VideoRover,一个统一的视频深度研究框架,该框架迭代协调视频裁剪、多模态搜索和网页浏览。给定视频-问题对,VideoRover利用每个工具的结果选择下一个动作,因此定位的视频片段指导外部检索,而检索到的证据会触发进一步的视频检查和验证。为开发该能力,我们构建了自动化数据整理流水线,生成了26000条经过验证的SFT轨迹和3000个具有挑战性的RL实例。我们还推出VideoRover-Bench,一个按视频时长和研究难度分层的基准。在VideoDR和VideoRover-Bench上的实验表明,我们的VideoRover-8B-RL在不使用工具的直接回答设置中达到了与专有模型相当的性能,同时在配备相同工具套件的情况下优于更大的开源模型。消融研究和训练动态进一步验证了主动视频定位、外部检索和长视野强化学习的互补作用。
英文摘要
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are typically developed in isolation. We introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover uses each tool result to select the next action, so localized video clips guide external retrieval and retrieved evidence triggers further video inspection and verification. To develop this capability, we construct an automated data curation pipeline, producing 26K verified SFT trajectories and 3K challenging RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show that our VideoRover-8B-RL achieves performance comparable to proprietary models in the direct-answer setting without tool use while outperforming larger open-source models equipped with the same tool suite. Ablation studies and training dynamics further validate the complementary roles of active video grounding, external retrieval, and long-horizon reinforcement learning.
发表机构
- Shandong University(山东大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Beihang University(北京航空航天大学)
- City University of Hong Kong(香港城市大学)
- Nanyang Technological University(南洋理工大学)
- Kuaishou Technology(快手科技)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。