发表机构
Microsoft Applied Sciences Group (ASG)(微软应用科学组(ASG))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AttSVD通过基于提示注意力几何的在线截断SVD压缩KV缓存,保留所有令牌并自适应调整秩,在多种模型和基准上以最多50%内存保持与密集缓存相当的性能。
AI 中文摘要
自回归Transformer的键值(KV)缓存随上下文长度线性增长,并在长上下文场景中主导内存占用。大多数无需训练的内存优化方法会驱逐低重要性令牌,这是沿序列轴的一种不可逆选择。我们则保留每个令牌,并沿“特征”轴更廉价地存储它们。为此,我们提出AttSVD,一种新的“可解释”低秩压缩方法,其基源自每个提示自身的注意力几何结构:一种在线、逐提示的截断SVD,仅保留注意力实际读取的方向,从而按保留秩的比例削减每个头的持久KV内存。我们提出两种解码时缓存策略——累积式和流式,分别适用于短生成和长生成场景。此外,我们提出两项改进使压缩具有自适应性。一个逐矩阵能量规则独立地调整logit空间和注意力质量的大小。一个注意力感知的基截断仅在注意力实际读取的空间中进行,同时保留注意力logits和注意力输出。这些因素还提供了免费的、逐头的可解释性洞察,揭示有效秩和注意力消耗的几何结构。在多种模型上,在智能体基准和完整LongBench套件上,AttSVD与密集缓存保持同等性能,同时仅使用最多50%的KV缓存内存。
英文摘要
The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "interpretable" low-rank compression whose basis is derived from each prompt's own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes. Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory.