周期性弱点:分块KV缓存压缩中的相位敏感性
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
- Princeton University(普林斯顿大学)
- Stanford University(斯坦福大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究揭示分块KV缓存压缩引入的相位敏感性导致检索性能周期性波动,通过预训练系列模型和因果干预分析其机制,强调评估需跨相位测量以避免平均分数掩盖位置失败。
中文摘要 AI 辅助
分块KV缓存压缩通过以固定步幅将连续令牌窗口压缩为更少的缓存条目,降低了长上下文推理的内存和注意力成本。这种压缩还引入了一个新的位置坐标:令牌的相位,即其相对于压缩窗口边界的位置。我们发现了使用此类压缩的模型中存在系统性不对称:相同的信息在一个相位容易检索,在另一个相位却难以检索。我们将这种检索性能的周期性变化称为相位敏感性。在采用此类压缩的大型开放权重模型中,长上下文检索准确率在不同相位间可相差高达40个百分点,揭示了平均基准分数可能掩盖的周期性弱点。为了研究这一行为,我们从零开始预训练了一系列跨越多种KV压缩设计的Transformer模型,并在各变体中重现了相位敏感性。利用这些模型中的因果干预进行的机制分析揭示了相位特化:不同的注意力组件在从不同源相位检索信息时贡献不对称。我们进一步分析了理想化检索模型,展示了梯度流动力学可能有利于尖锐的相位特化。因此,评估采用分块KV缓存压缩的模型需要跨压缩相位进行测量:高平均准确率可能与系统性的位置失败共存。
英文摘要
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.