arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长度泛化需要适当的正则化

Length Generalization Needs Proper Regularization

Pavlo Vasylenko, Matthias Lindemann, André F. T. Martins, Marcos Treviso

arXiv 2610.04518首次发表:更新:

发表机构

Instituto Superior Técnico, Universidade de Lisboa; Instituto de Telecomunicações; ELLIS Unit Lisbon; Gandara AI(里斯本大学高等技术学院; 电信研究所; ELLIS 里斯本分部; Gandara AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示长度泛化不仅依赖架构,还受训练正则化影响;权重衰减阻碍外推,而调整dropout位置可显著改善,并提出VPAD策略进一步提升外推性能。

AI 中文摘要

长度泛化是序列模型在训练期间未见过的上下文长度上表现良好的能力。在这项工作中,我们表明,实现长度泛化的挑战不仅与架构选择(如位置编码和注意力机制)有关,还与训练过程本身有关。我们研究了正则化如何影响长度泛化,并发现权重衰减会阻碍外推。相比之下,当重新考虑dropout在架构中的位置时,dropout能改善外推。在这方面,我们表明,在层归一化之前放置dropout的标准做法会引入系统性的分布不匹配,而将dropout恰好放在线性投影之前则解决了这个问题。例如,一个带有滑动窗口注意力的改进型SmolLM3,使用dropout进行持续预训练,可以在Needle-in-a-Haystack上完美外推到64倍,并在RULER和HELMET上远远超出预训练上下文大小。Mamba2也受益于dropout,这表明该问题具有架构无关性。我们进一步提出了方差保持仿射dropout(VPAD),这是一种新的dropout策略,能大幅减少由此产生的预激活方差不匹配,从而在变压器中带来进一步的外推改进。

英文摘要

Length generalization is the ability of sequential models to perform well on context lengths unseen during training. In this work, we show that the challenge of achieving length generalization is related not only to architectural choices such as positional encoding and the attention mechanism but also to the training procedure itself. We study how regularization affects length generalization and find that weight decay can hinder extrapolation. In contrast, dropout improves extrapolation when its placement within the architecture is reconsidered. In that regard, we show that the standard placement of dropout before layer normalization introduces a systematic distributional mismatch, and that applying dropout just before the linear projection resolves this issue. For example, a modified SmolLM3 with sliding window attention, continually pre-trained with dropout, can extrapolate perfectly to 64$\times$ on Needle-in-a-Haystack and far beyond the pre-training context size on RULER and HELMET. Mamba2 also benefits from dropout, suggesting an architecture-agnostic nature of the problem. We further propose Variance-Preserving Affine Dropout (VPAD), a new dropout strategy that substantially reduces the resulting pre-activation variance mismatch, leading to further extrapolation improvements in transformers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑