arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VideoResearchAgent:面向开放网页视频研究的接地任务合成与仿真到现实强化学习

VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research

Yuhang Zhou, Fei Li, Yuxi Wu, Bin Zhu, Jingjing Chen

arXiv 2610.04911首次发表:更新:

发表机构

Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI; Singapore Management University(复旦大学可信具身智能研究院; 上海多模态具身智能重点实验室; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VideoResearchAgent框架,通过可控任务合成、本地视频模拟器加速及检索域随机化强化学习,实现开放网页视频研究,在Video-BrowseComp上以40.48%准确率媲美Gemini-3-Flash-Preview,并降低74.9%令牌消耗。

AI 中文摘要

现有的深度研究智能体主要针对基于文本和图像的网页来源设计,而视频推理系统通常假设相关视频已提前提供。我们研究开放网页视频研究,即智能体必须自主发现相关视频、浏览其时间内容,并将答案锚定在视觉证据上。大规模训练此类智能体具有挑战性,因为实时视频交互缓慢且不可靠,而固定的本地仿真可能引入检索特定的捷径,这些捷径无法迁移到开放网页。我们引入VideoResearchAgent,一个可扩展的训练框架以解决这些挑战。首先,我们提出可控任务合成流程,从带时间戳的视觉证据中合成多跳研究任务,同时过滤仅文本的捷径。其次,我们构建一个字段对齐的本地视频模拟器,保留面向部署的搜索和观看交互,同时将视频搜索加速34.5-64.6倍。第三,我们引入检索域随机化GRPO(RDR-GRPO),在训练期间多样化候选排名、干扰项、元数据和结果结构,以减少对模拟检索的过拟合。在Video-BrowseComp上,使用Qwen3.5-4B训练的VideoResearchAgent达到40.48%的准确率,与Gemini-3-Flash-Preview相当,同时相对于未训练模型将累计API令牌消耗减少74.9%。这些结果共同为开放网页视频研究建立了一个准确且高效的训练配方。

英文摘要

Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑