切换线性注意力
Switching Linear Attention
查看机构详情
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出切换线性注意力(SwiLA),通过在线期望最大化在混合线性回归中实现动态组件选择,在保持线性注意力固定内存的同时增强表达能力,并在多个基准上接近或超越 softmax 注意力。
中文摘要 AI 辅助
设计具有高效推理能力的表达性序列层是现代机器学习中的一个核心挑战。标准 softmax 注意力通过丰富的非线性 token 交互实现了优异的序列建模性能,但它需要一个随序列长度线性增长的键值缓存,这限制了其可扩展性。线性注意力支持具有恒定内存占用的高效循环计算,但其降低的表达能力往往导致较差的建模性能。我们引入了切换线性注意力(SwiLA),一种新颖的序列层,通过增强表示能力同时保留线性注意力的固定大小循环状态来弥合这一差距。我们从测试时回归框架推导出 SwiLA 的循环更新规则,将状态更新规则视为在线期望最大化在混合线性回归模型中的应用。在测试时,每个输出维度根据输入动态地在多个线性注意力组件之间进行选择。在关联回忆、上下文语言学习和语言建模基准测试中,SwiLA 表现出强大的性能,缩小了与 softmax 注意力的差距,甚至在多种设置下超越了它。
英文摘要
Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.