arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VideoResearcher:面向长视频理解的自我改进工具设计

VideoResearcher: Self-Improving Tool Design for Long-Video Understanding

Dingqiang Ye, Dongdi Zhao, Kaishen Wang, Qingqiao Hu, Jingchen Sun, Yijun Liang, Yuqi Jia, Yiqiao Huang, Yunjie Tian, Jiaxing Zhang, Chuanyang Jin, Ke Zhang, Vishal M. Patel, Di Fu

arXiv 2609.19664首次发表:更新:

发表机构

Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VideoResearcher提出无需训练的多智能体框架,通过双重循环自主设计、测试和优化视频理解工具,在自我改进智能体中达到最先进性能,接近人工设计上限。

AI 中文摘要

视频智能体在长视频理解方面已取得实质性进展。然而,有效的视频智能体系统需要昂贵且耗时的人工设计和反复试验。当前的自我改进方法要么优化影响较小的提示词,要么重新组合预定义的微工具,要么在框架优化中难以收敛。为弥补这一差距,我们针对高影响力的视频工具提出VideoResearcher,这是一个无需训练的多智能体框架,能够像人类研究员一样自主设计、测试和优化用于视频理解的工具。VideoResearcher通过双重求解与进化循环运作:它分析工具使用轨迹以识别能力差距,协调专门智能体开发和验证可执行工具,并重用进化后的工具以加强后续视频推理中的证据获取。通过迭代式工具优化和验证,它无需更新模型参数即可逐步加强证据获取。VideoResearcher在自我改进智能体中达到最先进性能,并接近人工设计的上限,展示了一种无需训练的长视频理解范式,通过自主工具开发扩展智能体能力,同时减少昂贵的人工工程。

英文摘要

Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current self-improvement methods either refine low-impact prompts, recombine predefined micro-tools, or struggle with convergence in harness optimization. To bridge this gap, we target high-impact video-tool with VideoResearcher, a training-free multi-agent framework that autonomously designs, tests, and refines tools for video understanding, like a human researcher. VideoResearcher operates through dual Solving and Evolving loops: it analyzes tool-use trajectories to identify capability gaps, coordinates specialized agents to develop and validate executable tools, and reuses evolved tools to strengthen evidence acquisition in subsequent video reasoning. Through iterative tool refinement and validation, it progressively strengthens evidence acquisition without updating model parameters. VideoResearcher achieves state-of-the-art performance among self-improving agents and approaches the human-designed upper bound, demonstrating a training-free paradigm for long-video understanding that expands agent capabilities through autonomous tool development while reducing costly manual engineering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑