发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对有限数据预训练场景,提出递归Transformer结合因式分解词嵌入的方法,在10M和100M词预算下性能优于标准Transformer,且与2025年BabyLM挑战赛优胜者表现相当。
AI 中文摘要
在有限数据下的预训练需要与网页级语言建模不同的缩放视角。在固定数据预算但计算资源相对充足的情况下,增加参数数量仅在达到最优规模前有效;超过该临界点后,模型会过拟合,泛化能力恶化。我们在1000万至1亿词的预训练预算、两个语料库及多个下游评估任务中研究该行为,发现最优规模强烈依赖于数据预算和下游目标。我们认为标准Transformer在该场景下缩放效果差,因为词嵌入占用了参数预算的很大一部分,且每词元计算与表征能力绑定。为解决该耦合问题,我们研究递归Transformer,其通过在深度维度复用共享模块来缩放计算,同时结合因式分解词嵌入以减少词汇映射参数。我们训练了三个递归模型,发现它们在1000万和1亿词规模下均优于标准Transformer,且与2025年BabyLM挑战赛优胜者表现相当。
英文摘要
Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scale; beyond that point, models overfit and generalization worsens. We study this behavior across 10M-100M word pre-training budgets, two corpora, and multiple downstream evaluations, and find that optimal size depends strongly on both the data budget and the downstream target. We argue that standard Transformers scale down poorly to this setting, because embeddings consume a large fraction of the parameter budget and per-token computation is tied to representational capacity. To address this coupling, we study recursive Transformers, reusing a shared block across depth to scale compute, together with factorized embeddings to reduce vocabulary-map parameters. We train three recursive models and find that they outperform standard Transformers at 10M and 100M words, while remaining competitive with BabyLM Challenge 2025 winners.