发表机构
Technion – Israel Institute of Technology; Kempner Institute; Harvard University(以色列理工学院; 肯普纳研究所; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过在代码切换文本上预训练小型Transformer模型,发现代码切换课程学习能有效对齐跨语言表征,提升多语言预训练效果,并在BabyLM评估中优于基线。
AI 中文摘要
多语言社区的儿童经常进行代码切换,即在一次话语中使用多种语言。我们能否通过在代码切换文本上训练语言模型来诱导跨语言对齐?我们在两个1亿词的多语言语料库上预训练了小型仅解码器Transformer模型:一个基础语料库,由英语、荷兰语和中文的BabyBabelLM数据集混合而成;另一个语料库,通过使用LLM插入词级和句级代码切换从基础语料库生成。我们发现,在代码切换数据上训练能够对齐平行文本的表征,尤其是在不同文字之间,并且这种对齐在随后对单语文档的训练中得以保持。在从词级代码切换到句级代码切换再到单语文档的学习课程下,在代码切换数据上训练的模型在BabyLM评估套件上优于未使用代码切换数据训练的基线模型。我们的工作将代码切换课程学习定性为多语言预训练的一种有效数据增强方法。我们在以下网址发布代码、数据和模型:此https URL。
英文摘要
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.
Comments17 pages, 8 figures. Accepted to the BabyLM Workshop at EMNLP 2026