VideoScout:面向长视频理解的自适应推理节奏智能体主动探索学习
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
浏览论文内容
中文总结 AI 辅助
针对长视频理解中稀疏证据定位难的问题,提出VideoScout多轮推理智能体,通过自适应观看节奏实现序列证据获取,结合两阶段训练与DAPO强化学习,7B模型在基准上表现强劲。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在短视频理解方面取得了显著进展,但由于视觉上下文窗口有限,其在长视频上的表现仍然受限。现有方法依赖于均匀帧采样或近期提出的从粗到细的智能体缩放,这两种方法都难以在足够长的视频中定位稀疏且决定性的证据。我们将长视频理解形式化为一个序列证据获取(SEA)问题,其中智能体沿时间轴逐轮读取视频,在每一轮决定观看速度、保留哪些证据、何时重新访问不确定片段,以及何时停止并作答。受此观点启发,我们提出了VideoScout,一个通过自适应推理节奏实例化SEA范式的多轮推理智能体。具体而言,通过动态控制观看节奏,VideoScout能够在有界视觉上下文窗口内高效遍历长视频,使智能体能够访问更多视频内容,同时平衡内容分析深度与阅读效率。为训练VideoScout,我们构建了VideoScout-66K,一个包含从10K个答案验证轨迹中提取的超过66K个高质量探索轮次的数据集,并采用两阶段流程:冷启动监督微调教会智能体每轮的输出格式,而解耦片段与动态采样策略优化(DAPO)算法则执行轨迹级强化学习,其复合奖励同时考虑答案准确性、输出格式符合性以及智能体观看进度与教师答案时机之间的时间对齐(通过交并比(IoU)衡量)。在长视频理解和推理基准上的大量实验表明,我们7B模型与现有经过训练的7B智能体模型相比取得了强劲的性能。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or recent coarse-to-fine agentic zooming, both of which struggle to localize sparse, decisive evidence in sufficiently long videos. We formulate long video understanding as a \textbf{Sequential Evidence Acquisition (SEA)} problem, in which an agent reads the video turn by turn along the temporal axis, deciding at each turn how fast to watch, what evidence to retain, when to revisit uncertain segments, and when to stop and answer. Inspired by this view, we propose \textbf{VideoScout}, a multi-turn reasoning agent that instantiates the SEA paradigm through adaptive reasoning pacing. Specifically, by dynamically controlling the viewing pace, VideoScout enables efficient traversal of long videos within a bounded visual context window, allowing the agent to access more video content while balancing content analysis depth with reading efficiency. To train VideoScout, we construct VideoScout-66K, a set of over 66K high-quality exploration turns from 10K answer-verified trajectories, and adopt a two-stage pipeline: cold-start supervised fine-tuning teaches the agent per-turn output format, while the Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) algorithm performs trajectory-level reinforcement learning with a composite reward that jointly considers answer accuracy, output format compliance, and the temporal alignment between the agent's viewing progress and the teacher's answer timing measured by intersection-over-union (IoU). Extensive experiments on long video understanding and reasoning benchmarks demonstrate that our 7B model achieves strong performance compared with existing trained 7B agentic models.