发表机构
University of Virginia; Google(弗吉尼亚大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
KeyRec提出免训练的有界视觉记忆框架,通过视觉缓存与事件库区分近期与历史信息,以10%令牌预算在流式和长视频基准中取得最优压缩性能。
AI 中文摘要
视觉语言模型越来越多地被用于理解长视频和连续流。然而,密集的视觉令牌会随着视频时长而累积,使得长上下文推理变得极其昂贵。现有的免训练视觉令牌选择方法通过保留信息性令牌来降低这一成本,但可能会丢失连贯的事件证据,并且难以区分详细的近期观察与长期历史。我们提出了KeyRec,一个用于构建有界视觉记忆的免训练框架。在查询无关的写入过程中,KeyRec将细粒度的近期观察保存在视觉缓存中,并将历史证据组织成结构化的事件库。候选事件根据其相对于先前存储事件的新颖性被提出,并通过在线添加-合并-驱逐更新进行维护。当问题到来时,一个纯文本路由器在近期记忆和事件记忆之间自适应地分配固定的读取预算,而无需重新处理历史帧。KeyRec作用于面向模型的视觉嵌入,并支持模块化编码器-投影仪视觉语言模型以及无编码器和投影仪的NEO-ov架构。在四个流式和长视频基准测试以及三个视觉语言模型主干上,KeyRec仅使用密集解码器面向视觉令牌预算的10%,就在15个设置中的13个中取得了最佳的压缩性能。它在实时问题上比最强的压缩基线高出2.21至18.37个百分点,在六个长视频设置中的五个中取得了最佳压缩结果,并在每个NEO-ov 2B设置中表现最佳。
英文摘要
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.