发表机构
Alibaba Group; University of Science and Technology of China; Shanghai Jiao Tong University; University of Chinese Academy of Sciences; Beihang University(阿里巴巴集团; 中国科学技术大学; 上海交通大学; 中国科学院大学; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VideoEvolve提出协同演化记忆与检索的自演化框架,通过交替式强化学习及瓶颈与能力感知反馈,将下游经验转化为可迁移能力,提升长视频理解。
AI 中文摘要
长视频理解日益依赖外部记忆将海量视觉流组织成紧凑的表示。然而,大多数基于记忆的方法动态调整针对不同问题的信息检索方式,却在很大程度上固定了记忆内容。这种不匹配使得缺失细节难以恢复,而存储的信息仅在其能被可靠检索时才具有价值。为解决此问题,我们提出VideoEvolve,一种新颖的自演化框架,协同演化记忆与检索以用于长视频理解。具体而言,从粗粒度的低帧率概览开始,VideoEvolve将用于选择性记忆增强的Memory Evolver与用于在演化记忆上进行自适应检索的Retrieval Evolver耦合。随后,我们通过交替式智能体强化学习(Agentic RL)协同演化这两个Evolver,在更新其中一个的同时冻结另一个。为引导这种交替演化,瓶颈感知演化反馈(BEF)识别当前瓶颈在于记忆还是检索,并将优化导向更受限的一侧。此外,VideoEvolve引入能力感知演化反馈(CEF),以缓解下游反馈将记忆过度特化于固定训练问题集的问题,将训练转向尚未充分发展但可学习的长视频能力。通过将Agentic RL与BEF和CEF相结合,VideoEvolve将下游推理经验转化为可迁移的能力更新,为从静态长视频系统走向经验驱动、自我改进的多模态智能提供了具体路径。在多个长视频理解基准上的大量实验证明了VideoEvolve的有效性。
英文摘要
Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.