arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分组值注意力:通过按需键重建实现高效KV缓存

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

Vishesh Tripathi, Abhay Kumar, Ramsha Khan

arXiv 2609.13285首次发表:更新:

发表机构

FrontiersMind(FrontiersMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出分组值注意力(GVA),通过存储分组值并利用线性映射按需重建键,减少KV缓存约45-47%的标量,在350M参数规模下达到接近GQA的基准准确率,并开发自定义解码内核以加速推理。

AI 中文摘要

KV缓存是Transformer解码的主要瓶颈:其内存占用和缓存读取流量随序列长度增长而增加。分组查询注意力(GQA)通过共享键值头来降低这一成本,但仍需在每一步存储键和值。我们提出了分组值注意力(GVA),它存储分组值,并通过学习到的线性映射重建内容键。在推理时,该映射可被吸收进查询中,从而在预期的解码路径中无需实例化内容键。一个小的共享解耦RoPE通道通过单独缓存的 positional key 保留位置信息。在所研究的配置中,与匹配的GQA相比,该表示将持久缓存标量减少了约45-47%。在350M参数规模、使用30B FineWeb-Edu token的情况下,16维位置变体在五个任务上的平均准确率达到44.18,而GQA为44.36,MLA为43.88。这些结果表明,在更紧凑的缓存表示下,基准准确率接近GQA。为了将这种紧凑表示转化为更快的自回归推理,我们开发了自定义解码内核,目前正在评估其端到端推理性能,并计划很快开源发布。

英文摘要

The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with sequence length. Grouped-query attention (GQA) reduces this cost by sharing key-value heads, but still stores both a key and a value at every step. We introduce Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map. At inference, the map can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path. A small shared decoupled RoPE channel retains positional information through a separately cached positional key. For the configurations studied, this representation reduces persistent cache scalars by approximately 45-47% relative to matched GQA. At the 350M-parameter scale with 30B FineWeb-Edu tokens, the 16-dimensional positional variant reaches 44.18 average accuracy across five tasks, compared with 44.36 for GQA and 43.88 for MLA. These results demonstrate near-GQA benchmark accuracy with a more compact cache representation. To translate this compact representation into faster autoregressive inference, we have developed custom decoding kernels and are currently evaluating their end-to-end inference performance with an open-source release planned soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑