arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合成预预训练在规模上依然有效,但并非作为语法先验

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal Štefánik, Aline Villavicencio, Nikolaos Aletras

arXiv 2609.39827首次发表:更新:

发表机构

University of Sheffield; Mohamed bin Zayed University of Artificial Intelligence; Hugging Face; National Institute of Informatics; University of Exeter; Federal University of Rio Grande do Norte(谢菲尔德大学; 穆罕默德·本·扎耶德人工智能大学; Hugging Face; 国立情报学研究所; 埃克塞特大学; 北里奥格兰德联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统验证了合成预预训练(PPT)在更大规模(至70亿参数、1000亿令牌)下仍能提升下游性能与令牌效率,但收益并非源于语法先验,而是来自改善长程检索的PPT任务,且对数据混合组成具有鲁棒性。

AI 中文摘要

预预训练(PPT)在合成非自然语言数据上提高了语言模型预训练(PT)期间的令牌效率。先前的工作将此收益归因于一种语法先验,即在PPT期间学习到的结构归纳偏置,该偏置可迁移至自然语言语法。然而,PPT仅在最多10亿参数的模型和低于20亿令牌的PT预算下,且主要基于网页文本进行了测试。尚不清楚PPT在更大规模以及结合多种来源(如代码和数学)的更现实的PT数据混合下是否有效。因此,我们对PPT进行了一项全面研究,涵盖五个PPT任务、四种PT数据混合、四个参数规模(5亿至70亿)以及高达1000亿令牌的PT预算。我们的结果表明,PPT的下游性能和令牌效率收益在规模上持续存在,例如在30亿规模下至少节省210亿PT令牌。然而,与先前工作相反,我们没有发现一致证据表明这些收益源于语法先验。下游性能在不同模型规模下并不一致地与语法可接受性对齐。相反,我们发现下游收益来自改善长程检索的PPT任务。最后,PPT的性能收益对PT数据混合的组成方式具有鲁棒性,仅当网页文本缺失时才会减弱。总体而言,PPT是PT的一种低成本补充,未来的PPT任务设计应针对长程检索而非自然语言语法。

英文摘要

Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.

CommentsPreprint. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑