发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出RAM-Net作为连接KV缓存量化与线性注意力的桥梁,实现Transformer到RAM-Net的权重迁移,在9个预训练Transformer模型上以5亿token预算恢复了教师模型87.1%的平均准确率提升。
AI 中文摘要
KV缓存量化与线性注意力是解决Transformer存储及计算成本问题的两种代表性方法。KV缓存量化将单个KV条目压缩为离散编码,但保留所有条目;而线性注意力会反复聚合多个历史KV贡献为固定大小的连续状态,但可能引入干扰。这种对比提出了一个问题:能否在单一机制中连接逐KV压缩与多KV聚合,以实现高效注意力?我们通过离散地址空间上的软分配,确定RAM-Net就是这样的桥梁。这些分配决定了与每个地址关联的连续槽状态的循环更新。在受限的RAM-Net构造下,我们证明软地址分配将硬量化匹配扩展为可分离的读写重叠,该重叠可局部近似全注意力相似度并支持循环聚合。这些连接进一步通过基于软量化中间构造的新路径,实现Transformer到RAM-Net的权重迁移。在9个参数规模从0.3B到7B的预训练Transformer模型上,RAM-Net在6个常识与知识任务中,仅使用每个模型5亿token的预算,平均恢复了教师模型相对于随机猜测的87.1%准确率提升。
英文摘要
KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.