arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27470cs.CVcs.CLcs.LG

DeltaS:读取门控线性注意力状态以在流式视频中进行KV缓存驱逐

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

  • Maum AI Inc.(Maum AI公司)
  • Seoul National University(首尔国立大学)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim

中文总结 AI 辅助

针对流式视频中KV缓存驱逐需在问题前决策的挑战,提出基于门控线性注意力状态漂移的查询无关、免训练方法DeltaS,以低开销显著提升长视频基准性能。

中文摘要 AI 辅助

近期视频语言模型日益采用混合架构,交错使用线性注意力和全注意力层,以实现高效的长上下文处理。虽然线性注意力的循环状态大小固定,但全注意力的KV缓存随视频流持续增长,因此在有限内存预算下必须进行驱逐。流式处理中的关键挑战在于驱逐必须在问题到达之前发生,因此决定保留什么内容时无法借助问题信息。现有的驱逐方法从KV缓存本身推导令牌分数,利用位置、注意力或键值表示,而基于注意力的分数进一步需要代理查询或额外计算。混合骨干网络提供了另一种信号来源。在门控增量线性注意力中,循环状态通过每个输入与可从状态中已检索内容之间的残差进行更新,因此其在帧块上的变化反映了该块带来的新信息量。我们提出DeltaS,一种与查询无关、无需训练的方法,保留引起较大归一化状态变化(即状态漂移)的视频块。在预算和保留策略固定的受控比较中,状态漂移优于基于位置、注意力和键值的信号。该信号仅占前向传播的1.9%,DeltaS在六个长视频基准上平均超过最强的查询无关有界内存基线2.1个百分点,在最长的基准上超过5.6个百分点。这些结果表明混合架构的两种记忆可以协同工作。代码可在https://github.com/MaumAI-Company/DeltaS获取。

英文摘要

Recent video-language models increasingly adopt hybrid architectures that interleave linear and full attention layers for efficient long-context processing. While the recurrent state of linear attention remains fixed in size, the KV cache of full attention continues to grow with the video stream, making eviction necessary under a bounded memory budget. The key challenge in streaming is that eviction must occur before the question arrives, so what to retain has to be decided without the question. Existing eviction methods derive token scores from the KV cache itself, using position, attention, or key-value representations, and attention-based scores further require proxy queries or extra computation. Hybrid backbones offer another source of signal. In gated-delta linear attention, the recurrent state is updated by the residual between each input and what can already be retrieved from the state, so its change over a chunk of frames reflects how much new information the chunk brings. We propose DeltaS, a query-agnostic, training-free method that retains video chunks inducing larger normalized state change, or state drift. In a controlled comparison with the budget and retention policy held fixed, state drift outperforms position-, attention-, and key-value-based signals. With a signal costing only 1.9% of the forward pass, DeltaS surpasses the strongest query-agnostic bounded-memory baseline by 2.1 points on average across six long-video benchmarks and by 5.6 points on the longest benchmark. These results suggest that the two memories of hybrid architectures can work cooperatively. Code is available at https://github.com/MaumAI-Company/DeltaS.

补充信息

↑