FlashBack:在流式视觉语言模型中知晓何时记忆
FlashBack: Knowing When to Remember in Streaming Vision-Language Models
浏览论文内容
中文总结 AI 辅助
FlashBack提出无需训练的流式视觉语言模型记忆框架,通过语义判断查询是否需要历史信息,并采用隔离的回忆轨迹结合Side-KV路径,在保持实时感知的同时提升长时程记忆任务性能。
中文摘要 AI 辅助
流式视觉语言模型必须在有限的计算预算下处理持续增长的视频流,这造成了实时感知与长期记忆之间的持续张力。检索历史信息提供了一种自然的补救措施,然而历史回忆并非总是有益的:不必要的历史信息可能会将无关背景引入当前推理,并干扰原生的实时感知。因此,有效的流式记忆不仅应解决记住什么的问题,还应解决何时以及如何访问记忆的问题。为此,我们引入了FlashBack,一种无需训练的框架,用于流式视觉语言模型中的选择性、多层级记忆。在检索历史之前,FlashBack利用冻结的流式视觉语言模型的语义理解来推断查询是否需要历史证据。这一评估决定了推理是保持在原生轨迹上,还是调用一个独立的回忆轨迹。回忆轨迹通过查询局部的Side-KV路径将近期上下文与检索到的长期记忆相结合,在保持局部时间连续性的同时不修改持久的原生状态。我们在StreamingVLM和Mage-VL-4B上实例化了FlashBack,并在OVO-Bench和StreamingBench上进行了评估。结果显示,在多个长时程和依赖记忆的任务上取得了改进,同时很大程度上保持了实时感知,性能与强大的基于训练的流式方法相当,且无需额外训练。我们的代码将在稍后公布。
英文摘要
Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.
发表机构
- Rightly Robotics
机构由 AI 辅助整理,请以论文原文为准。