arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35497cs.CV

Sprout:在推理中构建动态记忆用于智能体视频理解

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Wei Chen, Xuanyu Zheng, Yancheng Long, Haoyang Xu, Kaiyu Jiang, Bin Wen, Tingting Gao, Han Li, Long Chen

首次发表
浏览论文内容

中文总结 AI 辅助

Sprout提出在推理中动态构建视频记忆的智能体框架,通过时间树按需记录片段并更新记忆,降低上下文成本并保持或提升长视频问答准确性。

中文摘要 AI 辅助

长视频理解依赖于视频记忆来克服多模态大语言模型的上下文限制。现有方法遵循“先构建后推理”的流程:先为整个视频离线构建记忆,然后将其作为静态来源进行推理。在实践中,一个长视频会被多个问题共享,这一流程在两端都代价高昂:当问题较少时,为整个视频构建记忆的成本远高于回答这些问题;当问题较多时,记忆从不更新,因此回答问题时所学到的内容会丢失,无法用于下一个问题。为缓解这些问题,我们引入了Sprout,一个在推理中构建记忆的智能体框架:一棵时间树,随着问题的回答而萌发出详细的节点。智能体以低帧率逐段观看视频,在当前问题可被回答时停止,将每个片段记为树的一个粗粒度节点,并以更高帧率重新访问关键区间,用恢复的细节来细化这棵树。一旦一个片段被记录为文本,其视频输入便从上下文历史中移除,而原始视频仍可通过视频工具访问。记忆树和先前的问答记录在问题之间持续存在,因此记忆是在线的、动态的:从第一个问题开始构建,并由之后的每个问题更新。我们发现,用文本记忆替换累积的视频输入能显著减少上下文使用量,同时保持准确性,在某些设置下还有轻微提升。在三个模型的基准测试中,Sprout相对于代表性的离线记忆方法取得了相当或更好的准确性,且无需前置构建阶段,每个问题的上下文成本也更低。

英文摘要

Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.

发表机构

  • HKUST(香港科技大学)
  • Kling AI(可灵AI)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑