发表机构
Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对循环语言模型KV缓存膨胀问题,提出混合潜在注意力(HLA),通过滑动窗口加潜在表示压缩缓存,提升吞吐量并保持高准确率。
AI 中文摘要
循环语言模型对每个词元重复应用相同的层堆栈T次,这在不增加参数的情况下加深了模型,但使其键值(KV)缓存扩大了T倍。更大的缓存限制了GPU一次能解码的序列数量,并减慢了每一步解码速度,因为每一步都需要读取整个缓存。我们提出了混合潜在注意力(HLA),它在最近W个词元的滑动窗口内保留精确的键和值,并将每个更早的词元存储为紧凑的潜在表示,每个循环的查询直接读取该潜在表示,而无需重建键和值。我们在Ouro循环模型(T=4)上对HLA进行上训练,参数量分别为1.4B和2.6B,保持预训练权重冻结,仅训练新增参数以复现原始注意力。缓存每个词元缩小10.7倍,使每个GPU的并发序列数增加4.0-8.8倍,解码吞吐量在1K词元上下文下提升2.5倍,在16K上下文下提升高达7.4倍。HLA在数学、知识和推理基准上保留了原始准确率的97%以上,在长达16K词元的长上下文检索中保留了96-100%。经过监督微调后,它在竞赛级数学上的表现与微调后的原始模型相当。
英文摘要
Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.