发表机构
University of Göttingen(哥廷根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究通过非语言数据(如音乐、概率文法、元胞自动机)进行结构迁移作为语言建模的权重初始化,发现其可降低损失但下游表现不稳定,仅部分替代语言数据。
AI 中文摘要
高效的语言学习需要减少对大量数据和计算资源的依赖的方法。我们研究了结构迁移:首先在非语言数据上训练模型,以诱导对自然语言有用的先验。这种方法是一种用于多语言语言建模的权重初始化形式。我们通过下一个词元预测损失、模型中的权重偏移以及下游语言基准来评估迁移效果。几种符号数据类型——尤其是音乐、概率文法和元胞自动机——产生的语言建模损失低于随机初始化。这些收益与后续语言训练期间较小的权重偏移相一致,表明结构迁移将模型定位在参数空间中更有利的区域。然而,较低的损失并不一致地转化为更好的下游语言表现,并且来自非语言数据的迁移不如额外的语言数据高效。我们得出结论,对于下一个词元预测的训练目标,非语言数据可以作为语言数据的部分替代品,但不能可靠地支持更广泛的语言泛化。
英文摘要
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorable region of the parameter space. However, a lower loss does not translate consistently into better downstream linguistic performance, and transfer from non-language data is less efficient than additional language data. We conclude that non-language data can serve as a partial substitute for language data for the training objective of next-token prediction but does not reliably support broader linguistic generalization.
CommentsEMNLP 2026, BabyLM Challenge; 18 pages, 11 figures