AI 中文总结
本研究提出相对泛化不变性(RGI),发现优化器和架构变化仅引起token级损失的均匀偏移,而数据流改变会显著影响相对泛化,为区分LLM预训练各组件作用提供新视角。
AI 中文摘要
大语言模型(LLM)的预训练性能由训练三元组的三个组成部分共同塑造:优化器、模型架构和训练数据流。然而,这些组件如何以不同方式影响性能仍不清楚。我们通过研究相对泛化迈出了隔离其影响的第一步。我们引入了相对泛化不变性(RGI),即任意两个token之间的验证损失差异在不同模型间的不变性。我们表明,RGI在广泛的优化器和中等架构变化下近似成立,这表明这些选择会在token级损失上引起近似均匀的偏移。相反,改变训练数据流可以显著改变相对泛化。我们进一步表明,RGI不能仅由神经正切核或平均场机制解释,并证明它可以在过参数化的二次模型中涌现。总体而言,我们的工作将RGI识别为LLM预训练中的一个新现象,有助于区分优化器和架构与训练数据的影响。
英文摘要
Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.