从未来回溯:基于反事实惊奇的键值缓存管理
Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
浏览论文内容
中文总结 AI 辅助
该研究针对LLM的KV缓存内存占用问题,提出基于反事实惊奇的KV驱逐方案及快速单层近似方法,在基准测试中性能优于或可与现有最优方法竞争。
中文摘要 AI 辅助
近年来,通过压缩和驱逐策略进行的键值(KV)缓存管理已成为重要研究方向。大型语言模型(LLM)及其多模态变体在输出生成过程中的计算需求,可通过缓存后续缩放点积注意力操作所需的先前键值计算结果得到部分缓解,但这也带来了新问题:生成的KV缓存大小随上下文长度线性增长,当提示词或生成输出较长时会快速耗尽所有可用GPU内存。KV缓存管理会定期从缓存中修剪条目,以减少内存占用,同时保留足够信息保证生成准确性,附带优势是推理速度更快。我们提出一种简单却有效的KV驱逐方案,其核心见解是:可通过较新token准确预测的过去token是冗余的,可将其关联的键值从缓存中移除。为对驱逐条目评分,我们按原始顺序对token运行模型,复用KV缓存中已存储的键值表示,并应用反事实注意力掩码,使每个位置仅关注其未来上下文。该方案属于分布内方法,与实际缓存内容直接关联,无需额外训练。为进一步降低成本,我们还提出一种快速单层近似方法,将反事实传递限制在最后一个Transformer层,在精度损失极小的情况下,大幅提升了每次刷新周期的速度。我们在多种开源LLM和基准数据集上对该策略进行评估,结果显示其性能优于或可与其他最先进方法竞争,参考代码可在指定URL获取。
英文摘要
Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years. Computational demands of large language models (LLMs) and their multi-modal variants during output generation can be partially alleviated by caching previous key and value calculations needed by subsequent scaled dot-product attention operations. However, this leads to another problem: the size of the resulting KV cache grows linearly with context length and quickly consumes all available GPU memory when either the prompt or the generated output are long. KV cache management periodically prunes entries from the cache thereby reducing its memory footprint while attempting to retain sufficient information for accurate generation. A by-product is faster inference speed. We propose a simple yet effective KV eviction scheme motivated by the insight that past tokens which can be well-predicted from more recent tokens are redundant and their associated keys and values can be removed from the cache. To score entries for eviction we run the model on the tokens in their original order, reusing the key and value representations already stored in the KV cache, and applying a counter-causal attention mask so that each position attends only to its future context. This is in-distribution, tied directly to the actual cache contents, and requires no additional training. To further reduce cost, we additionally propose a fast single-layer approximation that restricts the counter-causal pass to the last transformer layer, achieving a significant speedup per refresh cycle at marginal accuracy cost. We evaluate our strategy on various open-source LLMs and benchmark datasets showing competitive or improved performance over other state-of-the-art methods. Reference code is available at https://github.com/metacognitionai/counter_causal.
发表机构
- Metacognition AI(元认知人工智能)
- Australian National University(澳大利亚国立大学)
- Adelaide University(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。