arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40048cs.CV

CoEvoWhen:超长视频时间定位的策略-工具协同进化

CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对超长视频时间定位中现有智能体方法依赖预定义策略和工具的问题,提出策略-工具协同进化框架,从VLM推理轨迹联合进化策略与工具,在不更新参数下提升定位精度并降低推理成本,实验验证其有效性与泛化性。

中文摘要 AI 辅助

超长视频时间定位要求在有限的视觉预算下平衡长距离证据搜索与细粒度事件理解,然而现有的智能体方法仍主要依赖预定义策略和工具能力。受此启发,我们提出了一种新颖的策略-工具协同进化框架,该框架从视觉语言模型(VLM)的智能体推理轨迹中联合进化高层策略和可执行的媒体工具,在不更新模型参数的情况下形成可复用的技能。在进化过程中,外部技能更新器提炼长视频时间定位中可迁移的任务经验,相应地优化长距离基于图像的观测与细粒度基于视频的观测的编排。伴随这些策略更新,更新器利用其编码能力升级现有工具或创建新工具,使工具适应长视频证据获取。配备进化后的技能,VLM在进化策略的指导下自主编排工具,协调图像和视频观测进行智能体推理,而无需依赖单独的更强规划模型。涵盖五个基准和三个VLM的大量实验表明,策略-工具协同进化持续提升超长视频中的时间定位精度,同时降低推理时的视觉标记成本,并且进化后的技能在无需额外任务特定进化的情况下,在通用长视频问答中带来显著的性能提升,证明了我们框架在长视频理解中的有效性和泛化能力。

英文摘要

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.

发表机构

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑