arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sigmoid注意力作为学习型KV缓存驱逐的更优基础

Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction

Isaac, Li

arXiv 2608.23296首次发表:更新:

发表机构

University of Pittsburgh(匹兹堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文探究注意力基础对学习型KV缓存驱逐软-硬转换的影响,通过2×2×2对比实验发现,Sigmoid注意力经学习型硬驱逐后,其门控模型删除KV条目时PPL变化可忽略,且比H₂O、KeyDiff等方法表现更优。

AI 中文摘要

学习型KV缓存驱逐常面临软-硬不匹配问题:训练时,可微分门控通常会衰减token贡献,而推理时仅当KV条目被物理移除时才能节省内存。本文探究注意力基础是否会影响该软-硬转换,使用在OpenWebText上训练的GPT-2规模Transformer,对注意力类型、学习型门控和位置编码进行受控的2×2×2对比。尽管Sigmoid注意力作为稠密语言模型表现更差,但学习型硬驱逐改变了有用的操作点:Sigmoid门控模型删除KV条目时,相对于自身无驱逐参考的困惑度(PPL)变化可忽略。在相同稠密主干的匹配存活缓存协议下,学习型Sigmoid门控比H₂O和KeyDiff实现获得更低的PPL,而Softmax门控无法一致优于这些事后方法。结果表明,注意力归一化会显著影响训练时的软门控能否顺利迁移到硬KV删除。

英文摘要

Learned KV-cache eviction often faces a soft-to-hard mismatch: during training, differentiable gates typically attenuate token contributions, whereas inference saves memory only when KV entries are physically removed. We ask whether the attention substrate affects this soft-to-hard transition. Using GPT-2-scale Transformers trained on OpenWebText, we run a controlled $2\times2\times2$ comparison over attention type, learned gating, and positional encoding. Although sigmoid attention is worse as a dense language model, learned hard eviction changes the useful operating points: sigmoid-gated models delete KV entries with negligible PPL change relative to their own no-eviction references. Under a matched live-cache protocol on the same dense backbones, learned sigmoid gates obtain lower PPL than our H$_2$O and KeyDiff implementations, whereas softmax gates do not uniformly beat these post-hoc methods. The results suggest that attention normalization can substantially affect whether a training-time soft gate transfers cleanly to hard KV deletion.

CommentsAccepted at the ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑