arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

门控槽注意力-2:线性注意力中的双侧联想记忆修正

Gated Slot Attention-2: Two-Sided Associative Memory Correction in Linear Attention

Ruijie Li, Shengnan Ding, Weimin Zhang, Derick Tang, Zhanpeng Zeng, Qinsong Zeng, Ming Chen, Jiaxi Hu, Yuxuan Liang

arXiv 2610.02816首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Tencent(香港科技大学(广州); 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对线性注意力固定大小记忆管理难题,提出门控槽注意力-2(GSA2),结合门控Oja规则-2与门控Delta规则-2进行双侧记忆修正,在保持线性时间和恒定内存下提升基准性能。

AI 中文摘要

线性注意力模型已成为标准注意力的高效替代方案,但有效管理其固定大小的循环记忆仍具挑战性。为改进记忆,近期工作探索了两个不同方向:用于精确修正与键关联值的delta规则变体,以及诸如门控槽注意力等基于槽的架构,用于分两阶段建模键和值记忆。我们观察到这些方向是互补的——delta规则提供有效的记忆修正,而两阶段结构提供了一种自然的方式在关联的两侧进行操作。基于这一见解,我们引入了一种新的门控Oja规则用于键侧修正,并通过解耦的擦除和写入控制对其进行扩展,得到门控Oja规则-2。随后,我们提出了门控槽注意力-2(GSA2),它通过共享潜在槽将用于键侧修正的门控Oja规则-2与用于值侧修正的门控Delta规则-2相结合。我们进一步推导了一种硬件高效的块状算法用于并行训练。实验表明,GSA2在多个基准上持续优于强线性注意力基线,同时保持线性时间序列建模和恒定内存的循环解码。

英文摘要

Linear attention models have emerged as efficient alternatives to standard attention, but effectively managing their fixed-size recurrent memory remains challenging. To improve memory, recent work has explored two distinct directions: delta-rule variants for precise correction of values associated with keys, and slot-based architectures such as Gated Slot Attention for modeling key and value memories in two stages. We observe that these directions are complementary--the delta rule provides effective memory correction, while the two-stage structure provides a natural way to operate on both sides of an association. Building on this insight, we introduce a new Gated Oja Rule for key-side correction and extend it with decoupled erase and write control to obtain Gated Oja Rule-2. We then introduce Gated Slot Attention-2 (GSA2), which combines Gated Oja Rule-2 for key-side correction with Gated Delta Rule-2 for value-side correction through shared latent slots. We further derive a hardware-efficient chunkwise algorithm for parallel training. Experiments demonstrate that GSA2 consistently improves over strong linear-attention baselines across benchmarks while retaining linear-time sequence modeling and constant-memory recurrent decoding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑