发表机构
University of Cambridge; EleutherAI; MIT(剑桥大学; EleutherAI; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Beetle是一个受控的双语语言模型预训练框架,通过独立操控训练条件,训练并发布了330个模型,发现分阶段课程优于平衡双语训练,为计算心理语言学提供了双语加工研究工具。
AI 中文摘要
双语语言模型(LMs)为研究训练条件如何塑造第二语言(L2)行为提供了一个受控环境,但先前的工作通常同时改变暴露结构、规模和架构,使得难以将效应归因于任何单一因素。我们引入了Beetle,一个受控的语言模型预训练框架,其中分词器、目标语言、训练预算和暴露结构均可独立操控,从而能够对训练条件进行系统且可比较的实验。利用Beetle,我们训练并发布了285个双语和45个单语开源语言模型,这些模型带有丰富的检查点,覆盖了多种暴露计划、数据规模和第一语言(L1),以研究多语言预训练以及双语和第二语言学习的计算建模。在人类双语和第二语言阅读时间预测及语法判断任务上评估模型时,我们发现,与平衡双语训练相比,分阶段和具有时间结构的课程始终能更好地与语言学习者的阅读时间对齐,且在较小数据规模和类型学上更接近的语言对中收益最大。Beetle模型是合适的工具,有助于将计算心理语言学从其主流的单语、以英语为中心的关注点转向人类双语加工的模型,以研究跨语言学习动态,同时支持基于社区开发的受控模型家族。
英文摘要
Bilingual language models (LMs) offer a controlled setting for studying how training conditions shape second-language (L2) behaviour, but prior work typically varies exposure structure, scale, and architecture at once, making it difficult to attribute effects to any single factor. We introduce Beetle, a controlled language model pretraining framework in which tokeniser, target language, training budget, and exposure structure are each independently manipulable, enabling systematic and comparable experimentation of training conditions. Using Beetle, we train and release 285 bilingual and 45 monolingual open-source LMs with rich checkpoints across a range of exposure schedules, data scales and first languages (L1s) to study multilingual pretraining and computational modelling of bilingualism and second language learning. Evaluating models on human bilingual and second language reading-time prediction and grammaticality judgement tasks, we find that staged and temporally structured curricula consistently improve alignment with language learner reading time compared to balanced bilingual training, with the largest gains at smaller data scales and for typologically closer language pairs. The Beetle models are well suited tools to help move computational psycholinguistics beyond its prevailing monolingual, English-centric focus toward models of human bilingual processing, to study cross-lingual learning dynamics, while supporting community-based development of controlled model families.
CommentsAccepted EMNLP Main Conference 2026