arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TwinKV:一种通过成对键冗余实现KV缓存驱逐的可组合修复通路

Small Frequency Corrections Can Change What Survives KV Cache Compression

Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng, Junyan Zhang, Yubo Gao, Xuming Hu

arXiv 2608.27128首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Bosum Institute of Management Science(香港科技大学(广州); 博深管理科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TwinKV是一种无训练、无注意力的KV缓存可组合修复通路,通过识别上下文键的近重复项交换冗余供体与孤儿令牌,在多个长上下文基准上提升了多数驱逐策略的性能,但对少样本分类任务无帮助。

AI 中文摘要

长上下文推理受限于键值(KV)缓存的内存占用,尤其对于资源预算紧张的小型模型。现有KV缓存驱逐方法通过模型的注意力分布对令牌打分,或在无注意力变体中通过每个键与全局参考点的距离打分。采用受控留一法探测,我们发现注意力大小与令牌对答案的因果贡献无关(斯皮尔曼相关系数ρ=-0.004),这对主流驱逐方法的前提提出了挑战。我们提出TwinKV,一种无需训练、无注意力的冗余信号,用于检测某令牌的键在上下文其他位置是否存在近重复项。TwinKV不替代现有策略,而是作为可组合修复通路:给定策略固定的保留集合,它识别无存活重复项的被驱逐令牌(“孤儿”)和信息在其他位置重复的保留令牌(“冗余供体”),并在保持原始预算和打分规则的前提下交换二者。我们将TwinKV与四种近期驱逐策略组合,在LongBench、LooGLE、RULER数据集及短上下文MMLU-Pro无危害控制任务上,以0.3、0.5、0.7的压缩比进行实验。在Qwen3-4B模型上,TwinKV提升了两种策略的多数配置,对第三种策略提升接近均匀,仅对第四种已接近性能上限的自适应基线有少数配置帮助;三种非上限策略的增益在最宽松压缩比下最小。但在搭配Llama-3.2-1B模型的RULER任务中,第四种策略在所有评估单元均有提升,因其单独评分仍有较大改进空间。更广泛而言,Llama-3.2-1B在LongBench上的平均增益较小,但在LongBench和LooGLE上的改进单元比例高于Qwen3-4B,且在RULER任务上取得明确胜利。我们还发现,在少样本分类示例这类任务结构中,TwinKV对两种模型均无帮助。

英文摘要

Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value energy measures entry strength but does not distinguish isolated keys from those with many similar neighbors. We introduce TwinKV, a training-free method that discounts value energy by nonlocal post-RoPE key frequency. Prefix attention allocates head capacities, while retained entries preserve their original keys and values under an exact storage budget. Across four language models, TwinKV exceeds five evaluated compressed baselines in mean score on LongBench, LooGLE, and RULER at 50\% KV removal. Component controls isolate the frequency contribution. On Llama-3.2-1B RULER at 75\% removal, normalized frequency weights average 0.95, yet change 7\% of nonprotected retained positions and improve value-only retention by about 5.5 points under both uniform and adapted capacities. Permuting the weights within heads weakens this gain. These results show that modest frequency corrections can change retention and answering outcomes, with effects that depend on the model and task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑