发表机构
Institute for Theoretical Physics, Goethe University Frankfurt(法兰克福大学理论物理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究Transformer在固定上下文长度下对未见词元间距离的泛化能力,通过合成延迟复制任务,探讨位置编码、数据多样性及迁移学习正负性,强调理解底层机制的重要性。
AI 中文摘要
针对Transformer的分布外长度泛化,即从短上下文外推到长上下文的任务,已被广泛研究。本文聚焦于距离泛化,即在保持上下文长度固定的情况下,改变训练与推理时词元间距离,以检验模型性能。我们构建了两个合成延迟复制任务,两者均涉及源与召回之间的有限距离,词元被完全或选择性复制,并在训练中未见过的延迟上测试模型。我们探讨三个问题:(A)位置编码方案如RoPE和ALiBi相对于无位置编码(NoPE)是否提升了距离分辨率?(B)数据多样性,即训练中看到的词元间距离数量,如何影响性能?(C)距离迁移学习何时为正或负?我们进行了深入调查,发现提升对底层机制的理解至关重要。
英文摘要
Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distances between source and recall, where tokens are copied either fully or selectively, and test models on delays unseen during training. We address three questions: (A) Do positional encoding schemes such as RoPE and ALiBi improve distance resolution relative to no positional encoding (NoPE)? (B) How does data diversity, the number of inter-token distances seen in training, affect performance? (C) When is distance transfer learning positive or negative? We present a thorough investigation, finding that it is paramount to improve our understanding of the underlying mechanisms.
Comments15 pages, 7 figures