发表机构
Salesforce AI Research; University of Illinois Urbana-Champaign(Salesforce AI研究院; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Random Attention方法不计算KV缓存token的得分,均匀随机驱逐token,在四个模型六个推理任务上性能与现有最强方法相当,且vLLM部署吞吐量提升32%-43%,揭示了推理轨迹的冗余抵御驱逐的机制。
AI 中文摘要
大型语言模型在需要扩展推理的任务上表现出卓越性能,但长推理链会使KV缓存成为严重的内存瓶颈。现有的KV缓存压缩方法都遵循同一范式:通过某种对后续重要性的估计为缓存token打分,保留得分最高的token。我们的研究表明,该选择信号几乎没有作用。Random Attention(随机注意力)保留提示(prompt),并在每个注意力头内均匀随机驱逐token,完全不计算得分;在四个模型和六个推理任务上,它的性能与现有最强驱逐方法相当,且在vLLM部署中的吞吐量比后者高32%-43%。控制实验解释了原因:1)提示是缓存中脆弱的部分,选择器之间的差距主要在于其选择信号是否恰好保留了提示;2)推理轨迹通过两个层面的冗余来抵御驱逐:文本层面(模型在工作时会重申仍需处理的内容)和注意力头层面(每个头都保留自身的轨迹副本),因此一旦提示安全,随机抽取就能保留模型仍需的足够副本,无需得分来挑选。我们的代码可在该URL公开获取。
英文摘要
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at https://github.com/SalesforceAIResearch/Random-Attention.