子词分段婴儿语言模型:学习分词以实现样本高效的预训练
Subword Segmental BabyLMs: Learning to Tokenise for Sample-Efficient Pretraining
浏览论文内容
中文总结 AI 辅助
本文针对2026年BabyLM挑战赛开发了SubSegGPT与SubSegDeBERTa两种可学习子词分词的子词分段语言模型,在两个赛道均取得优于基线的结果,验证了该方法可提升BabyLM预训练的样本效率。
中文摘要 AI 辅助
在标准的语言模型(LM)训练流程中,子词分词被用作预处理步骤。子词分段语言建模是一种替代范式,其中分词在训练过程中学习,使模型能够发现优化其训练目标的子词单元。本文中,我们提交了2026年婴儿语言模型(BabyLM)挑战赛的参赛作品,为此开发了两种新的子词分段语言模型:SubSegGPT和SubSegDeBERTa。SubSegGPT是一种仅解码器模型,在自回归预训练过程中学习分词;SubSegDeBERTa是一种基于编码器的模型,联合学习生成和分词掩码词。我们针对Strict和Strict-small两个赛道对这两种模型进行了训练。我们在Strict赛道的最佳参赛作品是SubSegDeBERTa,它在零样本评估中取得了显著提升;在Strict-small赛道的最佳参赛作品是SubSegGPT,它优于基于分词的基线模型。我们的结果表明,可学习的子词分词能够提高BabyLM预训练的样本效率。我们分析了模型的子词学习动态,发现分词会逐渐收敛到平衡形态对齐与细粒度分词的子词单元。
英文摘要
In the standard LM training pipeline, subword tokenisation is applied as a preprocessing step. Subword segmental language modelling is an alternative paradigm in which tokenisation is learned during training, allowing the model to discover subword units that optimise its training objective. In this paper, we present our submission to the 2026 BabyLM Challenge, for which we develop two new subword segmental LMs: SubSegGPT and SubSegDeBERTa. SubSegGPT is a decoder-only model that learns tokenisation during autoregressive pretraining. SubSegDeBERTa is an encoder-based model that jointly learns to generate and tokenise masked words. We train both for the Strict and Strict-small tracks. Our top submission to Strict is SubSegDeBERTa, which achieves notable gains in zero-shot evaluation. Our top submission to Strict-small is SubSegGPT, which outperforms tokenisation-based baselines. Our results show that learnable subword tokenisation can improve sample-efficiency for BabyLM pretraining. We analyse the subword learning dynamics of our models and find that tokenisation gradually converges on subword units that balance morphological alignment and fine-grained segmentation.
发表机构
- University of Cape Town(开普敦大学)
机构由 AI 辅助整理,请以论文原文为准。