顺序预训练偏好大型模型
Sequential Pretraining Favors Large Models
- Purdue University(普渡大学)
- University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示小型模型在顺序预训练中因首因偏差而学习低效,提出暴露疗法正则化以提升后期数据分布的学习,表明大型模型优势部分源于鲁棒性而非特征代表性。
AI中文摘要:
大型神经网络通常能获得小型模型无法学习的能力。这是源于大型模型学习了更具代表性的特征,还是因为它们对训练过程中未预见到的不利影响更具鲁棒性?我们定义并量化了这样一种不利影响,即首因偏差,它指的是早期数据分布的暴露对后期学习造成的损害程度。我们表明,小型模型可能将学习能力低效地分配给早期分布,而充分过参数化的模型则对此影响具有鲁棒性。这种低效在预训练中尤其重要,因为基础模型经常顺序而非联合地遇到异质数据分布。因此,小型基础模型可能难以学习训练后期遇到的数据分布,当后期数据强调诸如代码、数学和推理等期望能力时,这尤其有害。受这些发现的启发,我们引入了暴露疗法(ET),一种简单的正则化方法,它促进了顺序预训练期间学习能力的更高效分配。我们证明,ET 改善了基础模型在后期数据分布上的性能,以及在高达十亿参数规模的模型中的整体能力。总的来说,我们的结果表明,大型基础模型的一些优势可能源于对不利训练影响的更大鲁棒性,而非学习了更具代表性的特征,并且改进的训练算法可以在较小的模型中恢复其中的一些优势。
英文摘要:
Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models' performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.