arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12814cs.LGcs.AI

RunningTensor:将线性注意力推广到高阶循环状态

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

Luca Herranz-Celotti, Vincent Guigue

首次发表
浏览论文内容

中文总结 AI 辅助

RunningTensor将线性注意力的矩阵记忆推广到高阶张量,通过秩1外积更新和向量查询收缩,在保持线性时间的同时提升记忆容量,并在多查询关联回忆及语言理解任务上优于基线。

中文摘要 AI 辅助

线性注意力和状态空间模型提供了线性时间的序列建模,但其循环记忆仍然是一个二阶张量(即矩阵),限制了状态中可表示的交互阶数。我们引入了RunningTensor,它将这种记忆推广到阶数为$o$的张量,通过秩为1的外积进行更新,并通过与$o-1$个向量查询进行收缩来读取。阶数$2$恢复为线性注意力;我们以阶数$3$作为概念验证进行研究,同时保留循环形式和并行形式,并在序列长度$T$上保持线性,将工作记忆容量从$\mathcal{O}(W^2)$提升到$\mathcal{O}(W^o)$。在合成多查询关联回忆任务上,RunningTensor优于线性注意力和状态空间模型基线。在预训练后,它还在语言理解和非合成检索任务上提升了性能,这表明高阶循环状态可以提供超越矩阵值状态的有用额外记忆容量。

英文摘要

Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order-$o$ tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries. Order $2$ recovers linear attention; we study order $3$ as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length $T$ and improving working memory capacity from $\mathcal{O}(W^2)$ to $\mathcal{O}(W^o)$. On synthetic multi-query associative recall, RunningTensor outperforms linear-attention and SSM baselines. After pretraining, it also improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.

发表机构

  • ISIR Sorbonne(索邦大学ISIR)
  • Berkeley Lab(伯克利实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑