arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Maglev:滑动循环记忆

Maglev: Sliding Recurrent Memory

Bo Liu, Qiang Liu

arXiv 2608.02870首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Maglev,一种带固定记忆的循环Transformer架构,通过耦合模型与记忆一致性损失实现训练可并行性,在验证损失及下游预训练基准上优于基线,参数共享可减少内存。

AI 中文摘要

我们提出Maglev,这是一种具有固定大小记忆的循环Transformer架构,它在泛化滑动窗口注意力的同时,仍保持训练时的可并行性。Maglev由两个耦合模型组成:预填充器Q,它利用全注意力(实践中,我们对Q采用交错的全注意力与滑动窗口注意力,因为这能带来更强的性能;核心要求是Q比P更具表达能力,且能访问完整历史)生成记忆目标m'_t;解码器P,它仅使用滑动窗口注意力和循环键值(K/V)注入,为下一个词元预测生成解码器记忆m_t。我们通过记忆一致性损失训练Maglev,该损失使m_t与m'_t对齐,从而允许推理时仅使用P。实验表明,Maglev在验证损失和下游预训练基准上均优于滑动窗口和潜在循环Transformer基线;此外,P与Q共享参数可减少参数内存,同时保留大部分性能增益。

英文摘要

We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller $Q$, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for $Q$, as this yields stronger performance. The essential requirement is that $Q$ be more expressive than $P$, with access to the full history.} to produce memory targets $m'_t$, and a decoder $P$, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories $m_t$ for next-token prediction. We train \ours{} with a memory consistency loss that aligns $m_t$ with $m'_t$, allowing inference to use $P$ alone. Empirically, \ours{} improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines. Moreover, sharing parameters between $P$ and $Q$ reduces parameter memory while preserving most of the gains.

CommentsNeural Architecture Research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑