arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27763cs.LGcs.CLstat.ML

用于持续学习的快速权重注意力

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对持续学习,推导了Falcon系列快速权重更新规则,结合多种计算形式,在语言建模等任务中保持竞争力并提升可变数字加法的长度外推能力,分离了序列模型的时间对齐等关键机制。

中文摘要 AI 辅助

循环快速权重记忆与选择性状态空间模型将不断扩展的上下文压缩为固定大小的循环状态,使状态转换成为在线学习规则。我们在写后读自回归语义下研究该规则,对于此处考虑的前缀预测目标,步骤t处的局部快速记忆示例为前缀对齐对(𝐱_t,𝐲_t)=(φ(𝐤_{t-1}),𝐯_t)。常见的同步骤关联(φ(𝐤_t),𝐯_t)仍为因果关系,但优化不同的内部目标。我们推导了平方误差回归和负内积目标的归一化一阶更新,回归系列包含Falcon-1(标量NLMS更新)、Falcon-2(其逐列扩展)、Falcon-3(滑动窗口小批量更新),Falcon-1A/Falcon-2A/Falcon-3A为对应的内积变体。我们提供循环、掩码并行和分块并行形式,以及数值稳定的正衰减重归一化,代表性变体在语言建模中保持竞争力,并提升可变数字加法的长度外推能力,该框架在循环序列模型中分离了时间对齐、可塑性、遗忘和有界重放。

英文摘要

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step $t$ is the prefix-aligned pair $(\mathbf{x}_t,\mathbf{y}_t)=(ϕ(\mathbf{k}_{t-1}),\mathbf{v}_t)$. The common same-step association $(ϕ(\mathbf{k}_t),\mathbf{v}_t)$ remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative inner-product objectives. The regression family comprises Falcon-1 (a scalar NLMS update), Falcon-2 (its per-column extension), and Falcon-3 (a sliding-window mini-batch update); Falcon-1A/Falcon-2A/Falcon-3A are the corresponding inner-product variants. We provide recurrent, masked-parallel, and chunk-parallel forms, together with numerically stable positive-decay renormalization. Representative variants remain competitive in language modeling and improve length extrapolation on variable-digit addition. This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models.

补充信息

↑