发表机构
The Hong Kong University of Science and Technology (Guangzhou); Tencent(香港科技大学(广州); 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对线性注意力固定大小记忆管理难题,提出门控槽注意力-2(GSA2),结合门控Oja规则-2与门控Delta规则-2进行双侧记忆修正,在保持线性时间和恒定内存下提升基准性能。
AI 中文摘要
线性注意力模型已成为标准注意力的高效替代方案,但有效管理其固定大小的循环记忆仍具挑战性。为改进记忆,近期工作探索了两个不同方向:用于精确修正与键关联值的delta规则变体,以及诸如门控槽注意力等基于槽的架构,用于分两阶段建模键和值记忆。我们观察到这些方向是互补的——delta规则提供有效的记忆修正,而两阶段结构提供了一种自然的方式在关联的两侧进行操作。基于这一见解,我们引入了一种新的门控Oja规则用于键侧修正,并通过解耦的擦除和写入控制对其进行扩展,得到门控Oja规则-2。随后,我们提出了门控槽注意力-2(GSA2),它通过共享潜在槽将用于键侧修正的门控Oja规则-2与用于值侧修正的门控Delta规则-2相结合。我们进一步推导了一种硬件高效的块状算法用于并行训练。实验表明,GSA2在多个基准上持续优于强线性注意力基线,同时保持线性时间序列建模和恒定内存的循环解码。
英文摘要
Linear attention models have emerged as efficient alternatives to standard attention, but effectively managing their fixed-size recurrent memory remains challenging. To improve memory, recent work has explored two distinct directions: delta-rule variants for precise correction of values associated with keys, and slot-based architectures such as Gated Slot Attention for modeling key and value memories in two stages. We observe that these directions are complementary--the delta rule provides effective memory correction, while the two-stage structure provides a natural way to operate on both sides of an association. Building on this insight, we introduce a new Gated Oja Rule for key-side correction and extend it with decoupled erase and write control to obtain Gated Oja Rule-2. We then introduce Gated Slot Attention-2 (GSA2), which combines Gated Oja Rule-2 for key-side correction with Gated Delta Rule-2 for value-side correction through shared latent slots. We further derive a hardware-efficient chunkwise algorithm for parallel training. Experiments demonstrate that GSA2 consistently improves over strong linear-attention baselines across benchmarks while retaining linear-time sequence modeling and constant-memory recurrent decoding.