arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37169cs.LGcs.AIcs.CL

轨迹汤:通过多样化轨迹推动LLM中期训练的算力扩展前沿

Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

  • Ant Group(蚂蚁集团)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou

AI总结:

轨迹汤通过将中期训练预算分配到多个独立分支并平均合并最强检查点,超越了串行饱和,扩展了算力前沿,提升了聚合下游性能。

AI中文摘要:

中期训练为预训练的大语言模型赋予了专业化和推理能力,但这一阶段的收益受到限制,因为额外的串行计算带来的下游改进甚微,甚至可能降低某些能力,这为中期训练所能吸收的算力设置了实际上限。我们重新审视了如何将这一算力分配给单次运行或多次相似优化。我们发现,从共享检查点在不同受控配方下分叉出的分支会到达参数空间中可测量的不同区域,并建立了一种扩展单次运行无法提供的兼容多样性形式。因此,我们引入了轨迹汤,它将中期训练预算分配到多个独立分支,并通过轨迹内和轨迹间的平均将验证集上选择的最强检查点整合到单个模型中。局部偏差和方差分析区分了两种平均层次,表明轨迹间平均消除了轨迹内平均无法触及的残余误差,而检查点选择带来的偏差限制了值得合并的检查点数量。在不同模型规模、学习率调度、令牌预算和轨迹数量下,轨迹汤在匹配预算下优于最强单轨迹平均的聚合下游性能,并随着预算扩大而持续改进,且在相同的后训练流程后优势得以保持。这些结果将轨迹分配和合并定位为一种超越串行饱和、扩展中期训练算力前沿的实用方法。

英文摘要:

Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.

↑