arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25559cs.CVcs.AI

AdaVDR:面向视频深度研究的自适应工具使用与反思

AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

  • Alibaba Group(阿里巴巴集团)
  • Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xintong Zhang, Xiaomeng Fan, Shilin Yan, Ekko He, Zicheng Liu, Zijian Zou, Guannan Zhang, Yuwei Wu, Zhi Gao, Hongwei Xue

AI总结:

提出AdaVDR自适应视频深度研究智能体,通过自适应工具调用与反思解决视频深度研究的不当工具调用等问题,构建相关基准并在VDR-EE、VideoDR上取得优于基础模型的效果。

AI中文摘要:

视频深度研究通过联合理解视频内容并从开放网络中检索外部知识来回答复杂问题。然而,多样的问题和视频需要不同的工具使用策略,不当的工具调用会产生错误结果。不确定的定位与检索也会使不必要的交互成本高昂且易出错,增加延迟与推理错误。为应对这些挑战,我们提出AdaVDR,这是一种具备自适应工具调用与反思能力的自适应视频深度研究智能体。AdaVDR根据任务及其能力选择工具,仅在不可靠的中间结果需要修正时回溯。为实现这些能力,我们开发了一种视频深度研究数据构建流程:首先在多样视频中发现与检索相关的事件和实体,通过定位与外部检索获取详细信息以构建高质量问答对;针对每个问答,特定任务的提示将信息获取过程组织为工具使用轨迹,使不同问题与视频类型可遵循不同的定位与检索策略。我们进一步引入模型条件工具必要性过滤,该方法会针对目标模型的视频理解与内部知识评估工具调用,移除模型可绕过的工具或工具链,从而生成适配目标模型视频理解能力与知识的轨迹。利用该流程,我们构建了训练数据及基准VDR-EE,其涵盖以实体为中心和以事件为中心的问题。我们执行监督微调,随后采用具备冗余感知奖励的强化学习,以强化自适应工具调用与反思能力。实验表明,我们的方法在评估的开源模型中于VDR-EE上表现最佳,且在VideoDR上较其基础模型有显著提升。

英文摘要:

Video deep research answers complex questions by jointly understanding video content and retrieving external knowledge from the open Web. However, diverse questions and videos require different tool-use strategies, and inappropriate tool calls can produce incorrect results. Uncertain grounding and retrieval also make unnecessary interactions costly and error-prone, increasing latency and reasoning errors. To address these challenges, we propose AdaVDR, an adaptive video deep research agent with adaptive tool invocation and reflection. AdaVDR selects tools according to the task and its capabilities, and backtracks only when unreliable intermediate results require correction. To enable these capabilities, we develop a video deep research data construction pipeline. We first discover retrieval-relevant events and entities in diverse videos and acquire detailed information through grounding and external retrieval to construct high-quality QA pairs. For each QA, task-specific prompts organize the information acquisition process into a tool-use trajectory, allowing different question and video types to follow different grounding and retrieval strategies. We further introduce model-conditioned tool necessity filtering, which evaluates tool calls against the target model's video understanding and internal knowledge, removing tools or tool chains the model can bypass. This yields trajectories tailored to the target model's video understanding capability and knowledge. Using this pipeline, we construct training data and VDR-EE, a benchmark covering entity-centric and event-centric questions. We perform supervised fine-tuning followed by reinforcement learning with a redundancy-aware reward to strengthen adaptive tool invocation and reflection. Experiments show that our method performs best among the evaluated open-source models on VDR-EE and substantially improves over its base models on VideoDR.

↑