arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27881cs.CV

StreamEMS:面向视觉语言模型的自演进记忆方案的流视频理解

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

发表机构中国科学院深圳先进技术研究院 · 上海人工智能实验室
查看机构详情
  • Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • Shanghai AI Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yuxin Liu, Peiqin Zhuang, Yali Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出StreamEMS机制,通过语义与先验感知两个演进模块优化记忆表征,在OVO-Bench等数据集上优于现有方法,高token下降率下仍具优势。

中文摘要 AI 辅助

近年来,许多流视频理解方法被提出,这些方法通过构建外部记忆来存储历史数据以减少计算量。大多数方法聚焦于优化当前数据(写入)注入过程以及从记忆中检索有用的历史数据(读取),却忽视了进一步提升记忆本身表征能力的机会。本研究中,我们提出了StreamEMS,这是一种用于改进流视频理解的通用机制,它通过自演进记忆方案重构存储在记忆中的历史数据,以实现更具信息性和鲁棒性的记忆表征。具体而言,我们首先引入语义演进模块,该模块通过利用从粗到细逐步缩小语义尺度所发现的有价值记忆实体,将记忆演进为信息密度更高的表征。此外,我们还引入了先验感知演进模块,该模块通过利用先验记忆分布来优化当前记忆状态,从而将记忆演进为更鲁棒的表征。我们在广泛使用的流视频理解数据集OVO-Bench和StreamingBench上验证了所提设计的有效性,结果表明我们的方法优于其他方法。此外,即使在高token使用率下降率的设置下,我们方法的优势也始终明显,这表明我们的方法在释放记忆自身潜力方面具有有效性和鲁棒性。

英文摘要

Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooking the opportunity to further enhancing the representational capability of memory itself. In this work, we present StreamEMS, a general mechanism for improving streaming video understanding by re-structuring the historical data stored in memory through self-evolving memory scheme, enabling more informative and robust memory representations. Specifically, we first introduce a Semantic Evolution Module to evolve the memory into more information-dense representations by exploiting informative memory entities discovered via progressively shrinking semantic scales from coarse to fine. In addition, we further introduce a Prior-informed Evolution Module to evolve memory into more robust representations by leveraging prior memory distributions to refine the current memory state. We validate the effectiveness of our proposed designs on widely-used streaming video understanding datasets, i.e., OVO-Bench and StreamingBench, and the results showcase that our method performs better than other methods. Moreover, the advantage of our method becomes consistently evident even under high token usage drop rate settings, indicating the effectiveness and robustness of our method in unleashing the potential of the memory itself.

↑