arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12818cs.CVcs.AI

在线视频智能体框架用于长视频理解

Online Video Agent Harness for Long Video Understanding

Sen Yang, Boqiang Duan, Jing Yang, Weihao Bo, Jie Liu, Boyuan Tong, Ze Feng, Wenkang Zhang, Jingdong Wang, Hua Wu

首次发表
浏览论文内容

中文总结 AI 辅助

VideoXAgent提出纯在线视频智能体框架,按需调用专家工具聚合证据,以约50k token上下文在长视频基准上媲美前沿模型,仅用15%上下文达MINERVA水平。

中文摘要 AI 辅助

长视频理解通常表现为一种视觉上的“大海捞针”问题:与查询相关的证据稀疏地分布在长时间跨度中,而将密集帧打包进单个VLM上下文会引发“上下文腐烂”和高昂成本。现有的视频智能体往往依赖与查询无关的离线预处理或临时工具集,这可能会遗漏查询特定的细节并浪费计算资源。在本工作中,我们提出了VideoXAgent,一个纯粹在线的视频智能体框架,用于长视频理解,它从给定的视频文件和用户查询出发,规划并分解任务,按需调用专门的专家工具,并聚合多模态证据以产生最终答案,同时解决观察结果之间的冲突。为了支持这种按需调用,我们设计了一套异构专家工具,由数据驱动的原子能力分类法指导,涵盖脚本、VLM和领域模型(例如检测、OCR、ASR、人脸识别)。该框架还强制执行客观证据提示和预算感知控制,以抑制幻觉和非终止。在Video-MME-Long、LongVideoBench-Long、LVBench和MINERVA上,VideoXAgent在更小的上下文占用下与前沿LMM和视频智能体竞争——即使在一小时长的视频上,每个样本的智能体上下文也约为50k个token。特别是在MINERVA等复杂视频推理基准上,它仅使用1,024帧密集打包基线上下文的约15%就达到了同等水平。值得注意的是,该框架即使使用视觉较弱甚至仅文本的编排器也保持有效,这表明强大的长视频理解能力可以从渐进式的智能体证据寻求中涌现,而不是将整个视频打包进单个上下文。项目页面:此https URL

英文摘要

Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline preprocessing or ad hoc tool sets, which can miss query-specific details and waste computation. In this work, we present VideoXAgent, a purely online video-agent harness for long video understanding that starts from the given video file and user query, plans and decomposes the task, invokes specialized expert tools on demand, and aggregates multimodal evidence to produce a final answer while resolving conflicts among observations. To support this on-demand invocation, we design a suite of heterogeneous expert tools guided by a data-driven taxonomy of atomic capabilities, spanning scripts, VLMs, and domain models (e.g., detection, OCR, ASR, face recognition). The harness further enforces objective evidence prompting and budget-aware control to curb hallucination and non-termination. Across Video-MME-Long, LongVideoBench-Long, LVBench, and MINERVA, VideoXAgent is competitive with frontier LMMs and video agents under a smaller context footprint---about 50k tokens of agent context per sample, even on hour-long videos. In particular, on complex video-reasoning benchmarks such as MINERVA, it matches this level while using only about 15\% of the context of a 1,024-frame dense-packing baseline. Notably, the harness remains effective with a visually weak or even text-only orchestrator, suggesting that strong long-video understanding can emerge from progressive agentic evidence seeking rather than from packing the full video into a single context. Project page: https://go-agent-x.github.io/video_agent_harness/

发表机构

  • Baidu Inc(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑