arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将语码转换语境引入认知启发的双语模型训练

Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training

Zhuojing Huang, Luise Pohlmann, Lisa Beinborn

arXiv 2610.06161首次发表:更新:

发表机构

University of Göttingen(哥廷根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过控制语码转换的结构位置和动态切换率,在两种类型不同的语言对上训练双语模型,发现合成语码转换数据能改善类型相近语言的跨语言对齐。

AI 中文摘要

在语言习得过程中,双语儿童经常接触到语码转换的输入,并将其作为认知支架,以加速词汇增长和跨语言句法映射。相比之下,计算双语模型通常是在交错排列的单语语料库上进行预训练的。虽然在预训练期间引入合成语码转换已成为一种有前景的策略,用以增强跨语言对齐和下流任务性能,但控制其成功的结构和发育参数仍然鲜为人知。在本工作中,我们通过控制两个关键变量,即语码转换的结构位置和跨训练阶段的动态切换率,研究了在两种类型上不同的语言对中使用合成语码转换数据进行训练的效率。我们的结果表明,使用语码转换数据进行训练能改善类型上相近语言的跨语言对齐。

英文摘要

During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.

CommentsEMNLP 2026, BabyLM Challenge; 18 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑