跳出框框思考:滑动窗口KV推理中信息的保留与传递
Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
浏览论文内容
中文总结 AI 辅助
本研究通过五个开源模型实验,探讨滑动窗口KV推理中窗口外信息的保留与传递,发现保留旧状态可提升检索,且Muse Glimmer和Mistral 7B因滑动窗口注意力展现出最强的潜在信息中继能力。
中文摘要 AI 辅助
滑动窗口KV推理是指在增量处理序列时,仅保留最近键和值状态的固定大小缓存。该方法可在推理时应用于预训练的因果Transformer,无需额外训练,且随着处理更多token,其KV缓存内存保持固定。由于缓存状态是在较早token的上下文中计算的,它们可能携带超出当前窗口的信息,并将其传递给后续状态。本研究使用五个开源权重模型(涵盖Qwen、Llama、Mistral和Muse Glimmer)进行了一系列实验。我们调查了源自即时上下文窗口之外的信息是否能在滚动KV缓存中持续存在,并保持对检索的有用性。初步结果显示,与从原始token重新计算最终固定窗口相比,保留先前计算的状态可提高所有测试模型的检索性能。我们随后测量了这种效应的延伸范围,发现Muse Glimmer和Mistral 7B表现出最强的“潜在信息中继”:即使相关源token已离开缓存,它们也能恢复信息。这两个模型在其公开架构中都融入了滑动窗口注意力,这一关联促使我们测试使用滑动窗口进行训练是否能促进更可靠的信息保留。
英文摘要
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens are processed. Because cached states are computed in the context of earlier tokens, they may carry information from beyond the current window and transmit it to later states. This study presents a series of experiments using five open-weight models spanning Qwen, Llama, Mistral, and Muse Glimmer. We investigate whether information originating outside the immediate context window can persist through a rolling KV cache and remain useful for retrieval. Initial results show that retaining previously computed states improves retrieval across the models tested compared with recomputing the final fixed window from raw tokens. We then measure how far this effect extends and find that Muse Glimmer and Mistral 7B show the strongest \emph{latent information relay}: they can recover information even after the relevant source tokens have left the cache. Both models incorporate sliding-window attention in their published architectures, an association that motivates testing whether training with sliding windows promotes more reliable information retention.