arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08966cs.AIcs.CL

好的预训练,差的SFT:训练栈中的检查点质量

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

  • Aleph Alpha(阿莱夫阿尔法)

机构由 AI 辅助整理,请以论文原文为准。

Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges

AI总结:

本文发现,在30B混合专家训练中,预训练损失或基准分数最高的检查点并非后续训练的最佳起点,而解密度更高的检查点在下游训练后表现更优。

AI中文摘要:

语言模型检查点通常根据预训练损失或基准分数进行选择,假设得分最高的检查点将始终是后续训练的最佳起点。我们证明,在完整的30B混合专家训练流程中,这一假设可能失效。在完整下游训练栈后表现更好的检查点也具有更高的解密度,即在局部权重扰动下保留下游性能。

英文摘要:

Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.

↑