发表机构
College of Foreign Languages and Literature, Fudan University(复旦大学外国语言文学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出循环GPT-BERT,通过深度参数共享和循环计算,在数据有限时以较少参数达到相当性能,并在BabyLM 2026任务中验证了其有效性。
AI 中文摘要
当训练数据有限时,增加参数数量并非提升语言模型性能的唯一途径。一个较小的参数集,若被反复应用,也能带来相当的性能。我们在BabyLM 2026 Strict-small设置下研究循环GPT-BERT,将GPT-BERT的掩码下一词预测与因果语言建模目标与深度维参数共享相结合。我们在预处理后的748万词英文语料上训练,并比较目标比率、非循环与循环架构以及循环次数。我们最终的$4\ imes12$模型使用四个物理层进行十二次循环遍历,包含1218万参数。BabyLM 2026排行榜报告总体平均分为35.42,NLP平均分为48.48。与公开的BabyLM 10M Strict-small GPT-2和GPT-BERT基线相比,该模型在选定的语言学和下游指标(包括BLiMP和GLUE)上以更少的参数取得了相当的性能。循环消融实验表明,额外的循环计算可以改善训练并在选定的语言任务上保持强劲性能,而在其他任务上较差的性能可能揭示了循环设计的固有局限:仅使用少数物理层限制了模型的表示空间。
英文摘要
When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare objective ratios, non-looped and looped architectures, and loop counts. Our final $4\times12$ model uses four physical layers for twelve recurrent traversals and contains 12.18M parameters. The BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. Compared with public BabyLM 10M Strict-small GPT-2 and GPT-BERT baselines, it achieves comparable performance on selected linguistic and downstream metrics, including BLiMP and GLUE, with fewer parameters. The loop ablations show that additional recurrent computation can improve training and preserve strong performance on selected linguistic tasks, whereas poorer performance on other tasks may reveal an inherent limitation of the looped design: using only a few physical layers restricts the model's representational space.
Comments10 pages, 2 figures