arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38119cs.CV

VideoLoop:循环工作记忆对抗长视频智能体中的语义抖动

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

发表机构罗切斯特大学 · 微软
查看机构详情
  • University of Rochester(罗切斯特大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Jianming Xu, Jinfa Huang, Jingyang Lin, Zhengyuan Yang, Jiebo Luo

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视频智能体仅追加记忆导致的语义抖动问题,提出双循环架构VideoLoop,通过检索与重写有界工作记忆,在多个基准上平均提升4.2个百分点。

中文摘要 AI 辅助

长视频理解要求多模态智能体在许多推理步骤中迭代地收集证据。然而,大多数现有的智能体方法遭受语义抖动:随着仅追加的工作记忆增长,对关键证据的注意力崩溃,智能体失去对其已发现内容的访问能力。首先,我们提供了一个结构性论证,表明仅追加的记忆可以纳入新观察到的目标证据,但在没有重写操作符的情况下,无法移除累积的噪声或防止有序的上下文增长。其次,受此分析启发,我们提出了VideoLoop,一种具有两个耦合循环的多模态智能体。外循环对视频进行推理,而内循环在每一步之后,从过去观察和中间分析的无界文件系统中检索工件,并重写一个有界的工作记忆。大量实验证明了VideoLoop的有效性,它以即插即用的方式改进了四个流行的LVLM骨干,在VideoMME(长)上比基线平均提高了4.2个百分点。对工作记忆的进一步分析表明,VideoLoop缓解了语义抖动:在VideoMME(长)最难的四分之一问题上,一个仅读取智能体上下文的盲审者正确回答了81.1%,而仅追加的智能体为60.9%。使用Gemini 3.1 Pro,VideoLoop在VideoMME(长)上达到88.3%,在VideoMMMU上达到88.8%,在LongVideoBench(长)上达到80.9%。

英文摘要

Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).

补充信息

↑