MetaVideoAgent:面向长视频理解的自动视频智能体进化框架
MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding
浏览论文内容
中文总结 AI 辅助
MetaVideoAgent是自动进化视频智能体的框架,在VA-EvoBench上经4次迭代后将宏平均准确率从38.44%提至51.47%,性能优于现有最优固定设计智能体且成本更低。
中文摘要 AI 辅助
长视频理解需要在长时长多模态视频中定位与问题相关的稀疏证据。真实世界的视频分布在模态特定信息密度、内容结构和证据模式上存在差异,导致固定设计的视频智能体在不匹配时会产生冗余处理或失效。将自动智能体进化从文本扩展到视频具有挑战性,因为完整执行长视频会使候选验证成本高昂,失败会在耦合的证据处理阶段间传播,且复杂的预处理、感知工具和定位策略使得可靠实现代码级更新十分困难。我们提出MetaVideoAgent,一个能自动为目标分布进化视频智能体的框架。它通过稀疏采样的帧及关联查询分析信息密度和证据需求以指导初始设计,再将局部失败压缩为可独立执行的最小验证任务。它构建基于证据的黄金路径(Gold Paths),审计学生智能体轨迹,聚合样本间的重复失败并将其归因于责任模块。模块化智能体表示将每次更新限制在主要责任模块及其必要依赖项中。我们还推出VA-EvoBench,包含8种视频分布,每种分布都有独立的进化集和保留集。每种分布经过4次进化迭代后,MetaVideoAgent提升了所有初始智能体,使宏平均准确率从38.44%升至51.47%,每种分布的平均进化成本为354万token。进化后的智能体比所有对比视频智能体中最强的固定设计视频智能体性能高6.39个百分点,且在每个问题使用的token和视频帧数量最少。我们将发布所有代码和数据以支持可复现研究。
英文摘要
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched. Extending automated agent evolution from text to video is challenging because full long-video execution makes candidate validation expensive, failures propagate across coupled evidence-processing stages, and complex preprocessing, perception tools, and localization strategies make code-level updates difficult to implement reliably. We introduce MetaVideoAgent, a framework that automatically evolves a video agent for a target distribution. It profiles information density and evidence requirements from sparsely sampled frames and associated queries to guide initial design, then compresses localized failures into independently executable minimal validation tasks. It constructs evidence-grounded Gold Paths, audits Student trajectories, aggregates recurring failures across samples, and attributes them to responsible modules. A modular agent representation constrains each update to the primary responsible module and its necessary dependencies. We further introduce VA-EvoBench, covering eight video distributions with separate evolution and held-out splits. With four evolution iterations per distribution, MetaVideoAgent improves every initial agent and raises macro-average accuracy from 38.44% to 51.47%, at an average evolution cost of 3.54M tokens per distribution. The evolved agents outperform the strongest prior fixed-design video agent by 6.39 percentage points while using the fewest tokens and video frames per question among the compared video agents. We will release all code and data to support reproducible research.
发表机构
- Alibaba Group(阿里巴巴集团)
- Zhejiang University(浙江大学)
- Beihang University(北京航空航天大学)
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。