arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VidHarness:面向成本高效长视频理解的智能体框架演化

VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding

Susan Liang, Jianmin Wu, Daxiang Dong

arXiv 2609.38413首次发表:更新:

发表机构

Baidu, Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VidHarness通过蒙特卡洛树搜索自动演化视频理解智能体框架,结合多保真度验证和混合框架路由,在多个基准上以更低成本超越手工智能体。

AI 中文摘要

视觉语言模型(VLM)能够回答关于长达一小时视频的问题,但处理每一帧的成本高得令人望而却步,尽管问题的证据通常仅跨越几秒钟。视频智能体,即围绕冻结的VLM构建的框架程序,通过选择性观察视频来解决这一问题,然而现有的框架是由专家通过缓慢的构建-测试循环手工制作的。我们提出VidHarness,一个自动化框架设计的框架,用于成本高效的长视频理解,其中框架提议者基于演化环境的执行反馈迭代地演化框架。为了摆脱贪婪细化的局部最优,我们将演化组织为蒙特卡洛树搜索(MCTS),并为了降低评估成本,我们整合了不确定性感知的多保真度验证,该验证在少量问题上筛选新框架,仅提升有前景的框架。由于最佳框架随帧预算而变化,我们进一步引入了框架混合体,将每个问题路由到针对其预算专门化的框架。VidHarness在LongVideoBench、Video-MME和Video-Holmes上取得了新的最先进结果,比最强的手工视频智能体高出最多11.2个百分点,并且在知识密集型基准Video-MMMU和MMVU上以少于均匀采样一半的帧实现了泛化。

英文摘要

Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the video selectively, yet existing harnesses are hand-crafted by experts through slow build-and-test cycles. We propose VidHarness, a framework that automates harness design for cost-efficient long video understanding, in which a harness proposer iteratively evolves harnesses based on execution feedback from an evolution environment. To escape the local optima of greedy refinement, we organize the evolution as Monte Carlo tree search (MCTS), and to reduce the evaluation cost, we integrate uncertainty-aware multi-fidelity validation, which screens new harnesses on a few questions and promotes only the promising ones. Since the best harness varies with the frame budget, we further introduce a mixture-of-harness that routes each question to a harness specialized for its budget. VidHarness sets new state-of-the-art results on LongVideoBench, Video-MME, and Video-Holmes, outperforms the strongest hand-crafted video agent by up to $11.2$ points, and generalizes to the knowledge-intensive benchmarks Video-MMMU and MMVU with fewer than half of the frames of uniform sampling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑