三元线性注意力:用于长上下文序列建模的三维循环状态
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
浏览论文内容
中文总结 AI 辅助
本文提出三元线性注意力,通过将键、第二键和值的三元外积写入三维张量状态,以参数高效方式扩大RNN记忆容量,显著提升长上下文建模与召回性能。
中文摘要 AI 辅助
循环神经网络(RNN)将历史上下文压缩为固定大小的记忆状态,从而实现恒定时间推理。记忆状态的大小是影响其性能的关键因素,线性注意力的强大表现和复兴便是一个例证,它将普通RNN的向量值隐藏状态扩展为矩阵值隐藏状态。关键在于,线性注意力以参数高效的方式实现了这一点,特别是通过使用键和值向量的外积来写入矩阵值隐藏状态。我们推广了这一构造,提出了三元线性注意力,它将一个键、第二个键和一个值的三元外积写入三阶(即3D)张量状态,并通过将两个键轴与两个查询进行收缩来读取该状态。一个$E$维的第二个键因此使状态大小增加$E$倍,同时仅增加两个投影。三元线性注意力与数据依赖的遗忘、增量规则和分块并行训练兼容。应用于门控DeltaNet和标量门控线性注意力时,三元线性注意力显著改善了长上下文语言建模和召回,优于扩大状态大小的替代方法。
英文摘要
Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An $E$-dimensional second key thus yields an $E$-fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- MIT-IBM Computing Research Lab(MIT-IBM计算研究实验室)
机构由 AI 辅助整理,请以论文原文为准。