发表机构
University of Michigan; Eastern Michigan University(密歇根大学; 东密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探讨LLMs学习k-非局部语言的能力,发现其在不同k值非局部语言上交叉熵损失相当,但非局部性越强收敛越慢,为非局部依赖更难学习的观点提供了学习速度层面的证据。
AI 中文摘要
我们研究基于Transformer的语言模型(LLMs)学习所谓k-非局部语言的能力,这类语言在任意长度为k的连续符号区间之间均无互信息。我们构造了k值递增的此类语言,发现在其上训练的LLMs无论非局部性如何,交叉熵损失均相当,但在非局部性更强的语言上收敛速度更慢。我们的发现支持非局部依赖更难学习的观点,但这种偏差的证据来自学习速度而非学习成功与否。
英文摘要
We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.