ReMem:具有长上下文保留能力的流式视频理解
ReMem: Streaming Video Understanding With Long Context Retention
浏览论文内容
中文总结 AI 辅助
针对流式视频理解中长上下文丢失问题,提出免训练的ReMem方法,利用流式上下文记忆和检索视觉记忆增强VLM,在多个基准上达到SOTA。
中文摘要 AI 辅助
尽管当前的视觉语言模型(VLMs)在广泛的视频理解任务上表现出色,但它们主要针对离线场景设计,难以处理需要低延迟响应的在线流式视频。已有若干研究探索了记忆和令牌压缩策略,试图将离线VLM适配到流式视频理解任务中。然而,通过我们的探测实验,我们发现大多数现有工作在输入流长度增加时,会逐渐丢失长上下文信息。为解决这一问题,我们提出ReMem,一种新颖的免训练适配技术,使VLM能够处理任意长度的流式视频,同时提升其长上下文信息保留能力。ReMem从两个视角利用记忆,并实现为两个核心组件。流式上下文记忆(SCM)通过查询无关的注意力机制持续压缩历史上下文。检索视觉记忆(RVM)随后从记忆中检索最显著、与查询相关的上下文,以增强VLM的输入。综合实验表明,所提出的ReMem在多种广泛使用的基准测试上取得了最先进(SOTA)的性能,涵盖流式视频和通用长视频理解任务。
英文摘要
Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.
发表机构
- Nanyang Technological University(南洋理工大学)
- Shanghai Jiaotong University(上海交通大学)
- University of the Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。