arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习如何遗忘:面向长上下文稀疏注意力的微调

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

Matthias Seeger, Zeyu Zhang, Vihang Patil, Konstantinos Benidis, Sebastian Schelter

arXiv 2608.19920首次发表:更新:

发表机构

Amazon Web Services; University of Amsterdam; Amazon; Technical University Berlin(亚马逊网络服务; 阿姆斯特丹大学; 亚马逊; 柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种适配任意KV缓存策略的稀疏注意力模型微调方法,仅需中等硬件预算,性能优于精确注意力训练模型,还提供H2O稀疏注意力高效实现及开源库KeysAndValues支持。

AI 中文摘要

已有大量研究通过稀疏注意力处理键值(KV)缓存的选择与压缩,以在不占用过多硬件资源的情况下实现Transformer语言模型的长上下文推理。本文提出一种适用于稀疏注意力模型的微调新方法,该方法可适配任意KV缓存策略,仅需中等硬件预算(如1块40GB显存的Nvidia A100 GPU),支持模型与策略协同适配,性能常优于采用精确注意力(序列并行)训练的模型。本文还提供H2O稀疏注意力(实验中表现最优的策略)的高效实现,其带有专用缩放点积注意力内核支持;新开源长上下文推理与微调库KeysAndValues(对应链接)为本文所有方法提供易用且高性能的代码支持。

英文摘要

A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.

Comments42 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑