arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习长度可外推的循环模型

Learning Length-Extrapolatable Recurrent Models

Hanwen Jiang

arXiv 2609.09157首次发表:更新:

发表机构

Adobe Research(奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对循环模型在长序列上训练失效的问题,提出时间信用稳定化(CST)方法,通过稳定状态信用信号提升模型在超出训练长度达128倍时的性能。

AI 中文摘要

循环模型为长上下文建模提供了一条自然路径,然而通过时间反向传播(BPTT)训练的模型往往在超出其训练长度范围时失效。经典分析强调梯度沿时间路径会消失或爆炸。然而,密集的逐词元损失仍然可以在严重衰减的情况下训练出共享的循环规则,这表明衰减本身并不能决定学习是否失败。我们转而研究状态信用:即未来损失在贡献于参数更新之前,通过该信号到达早期循环状态的路径。据此,我们直接对状态信用进行干预,并提出时间信用稳定化(CST)方法。在反向传播过程中,CST局部地重新缩放状态信用信号,以稳定其范数,而不旋转被校正的分量,同时保持前向计算不变。由于受控的合成任务和真实数据表现出不同的信用动态,我们将CST针对每种情况进行了专门化处理。在两种设置下,CST都提升了超出训练长度的性能,在高达训练长度128倍的范围内观察到了增益。

英文摘要

Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑