发表机构
Amazon Web Services; University of Amsterdam; Amazon; Technical University Berlin(亚马逊网络服务; 阿姆斯特丹大学; 亚马逊; 柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种适配任意KV缓存策略的稀疏注意力模型微调方法,仅需中等硬件预算,性能优于精确注意力训练模型,还提供H2O稀疏注意力高效实现及开源库KeysAndValues支持。
AI 中文摘要
已有大量研究通过稀疏注意力处理键值(KV)缓存的选择与压缩,以在不占用过多硬件资源的情况下实现Transformer语言模型的长上下文推理。本文提出一种适用于稀疏注意力模型的微调新方法,该方法可适配任意KV缓存策略,仅需中等硬件预算(如1块40GB显存的Nvidia A100 GPU),支持模型与策略协同适配,性能常优于采用精确注意力(序列并行)训练的模型。本文还提供H2O稀疏注意力(实验中表现最优的策略)的高效实现,其带有专用缩放点积注意力内核支持;新开源长上下文推理与微调库KeysAndValues(对应链接)为本文所有方法提供易用且高性能的代码支持。
英文摘要
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.
Comments42 pages, 1 figure