发表机构
RND NLP, DAIMLD(RND NLP,DAIMLD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出循环语言模型的实用训练方案,通过减少训练预算、提升推理性能及最小化稠密模型转换方法,实现低成本高效训练并验证循环机制的有效性。
AI 中文摘要
循环语言模型通过重复应用共享层块来增加有效深度,但现有的大规模方案需要在数万亿个令牌上进行多阶段训练,而循环带来的收益难以与数据和训练差异区分开来。在本工作中,我们为循环语言模型建立了实用的训练方案,并取得三项主要成果。(1) 我们开发了一条计算高效的从头训练流程,将训练预算从Ouro中的7.7万亿令牌减少到3100亿令牌,同时保持强大的推理性能。预训练后接高质量的中期训练,结合学习率预热和更强的出口门控正则化,使得无需先前的多阶段调度即可实现稳定的循环训练。(2) 在受控比较下,我们的1.4B LoopLM在相同数据和令牌预算上训练的参数量匹配的稠密模型在所有12个评估基准上均表现更优,包括在GSM8K上+14分、MATH上+10分、DROP上+22分。在匹配推理计算量下,它在数学推理和阅读理解方面接近3.9B稠密模型,而仅使用其36%的参数。(3) 我们引入了一种将预训练稠密模型转换为循环模型的最小方案:一个单一的学习输入混合标量和一个平滑的出口损失,无需特定于步骤的参数。应用于Qwen3-1.7B-Base时,Looped Qwen在两种数据机制下的每个评估基准上均优于同等继续训练的稠密基线,在精选混合数据上GSM8K、MATH和MMLU-Pro上具有统计显著的提升。综合来看,这些结果使得循环语言模型从头训练的成本大幅降低,并且易于引入现有预训练检查点,同时隔离了循环本身带来的收益。
英文摘要
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36\% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.