arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

滑动窗口注意力优于线性注意力

Sliding-window beats linear attention

Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais

arXiv 2608.28444首次发表:更新:

发表机构

Microsoft; Applied Sciences Group (ASG)(微软公司; 应用科学集团(ASG))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比发现,带汇点的滑动窗口注意力(SWA)性能优于后训练的线性注意力模型,在长上下文推理任务中性能高2-10倍,且无需后训练、成本低,推荐用SWA替代线性注意力模型。

AI 中文摘要

由于二次注意力的特性,大语言模型(LLMs)会消耗大量内存和能量,每生成一个新token的成本都高于前一个,每新增一个token,键(keys)和值(values)都必须无限期存储在内存中,这是不可持续的。为解决二次缩放问题,研究人员提出了多种替代方案,其中之一是改造LLMs以使用线性注意力,该方案因有望以低成本实现最先进性能、解决二次缩放问题而受到大量关注。然而,这一研究方向尚未与更简单的基线进行恰当比较。本研究表明,带“汇点(sinks)”的滑动窗口注意力(Sliding Window Attention, SWA)的性能与后训练的线性注意力模型相当或更优,在多种下游任务的多个LLMs上均观察到这一结果。在长上下文推理任务(针在干草堆Needle-in-a-Haystack任务、BABILong任务)中,SWA的性能大幅更高,比线性注意力高出2至10倍。SWA无需后训练,速度极快且内存需求低,因此是一种极其廉价且可靠的解决方案。为降低推理内存成本,强烈建议切换至SWA而非后训练的线性模型。线性注意力模型可能已展现出一些潜力,但它们可能需要从头训练或大量后训练才能达到SWA的性能。

英文摘要

Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑