发表机构
LMU Munich; Munich Center for Machine Learning (MCML); Mila – Quebec AI Institute; University of Montreal; Saarland University(慕尼黑大学; 慕尼黑机器学习中心; 魁北克人工智能研究所; 蒙特利尔大学; 萨尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分解记忆训练动态,发现序列级梯度对齐为关键,低层模型层参与最多,通过参数干预可消除记忆,为预测和干预记忆提供新方法。
AI 中文摘要
记忆已被提出作为一种机制来解释语言模型如何拟合其训练分布的尾部,但其训练动态尚未被很好地理解。在这项工作中,我们通过将记忆序列的损失轨迹在训练和模型参数上进行分解,对记忆进行了细粒度的观察。在Pythia模型家族中,我们研究了重复训练序列(背诵)和稀有序列(回忆)的记忆情况。我们发现,在这两种情况下,记忆都以序列级别的梯度对齐为特征,尽管背诵会受到与其他训练影响的不对齐的困扰,这会导致遗忘,从而解释了这些示例需要更高重复度的必要性。我们进一步表明,较低的模型层最参与记忆和遗忘。在预测记忆方面,我们的分解优于交叉熵基线,尤其是在较大的模型和训练早期。通过干预一小部分高度有影响力的参数,我们能够在最终模型中消除记忆。综合来看,这些发现增进了我们对记忆在训练过程中如何发展的理解,并为预测和干预记忆提供了见解。
英文摘要
Memorization has been proposed as a mechanism to explain how language models fit the tail of their training distributions, but its training dynamics are not understood well. In this work, we take a fine-grained look at memorization by decomposing the loss trajectory of memorized sequences over training and model parameters. Across the Pythia family, we study memorization of duplicated training sequences (recitation) and rare ones (recollection). We find that memorization in both cases is characterized by sequence-level gradient alignment, though recitation suffers from misalignment with other training influences which causes forgetting, explaining the necessity for higher duplication of these examples. We further show that the lower model layers are the most involved in memorization and forgetting. Predicting memorization, our decomposition improves over a cross-entropy baseline, especially in larger models and early in training. Intervening on a small set of highly influential parameters we are able to ablate memorization in the final model. Together, these findings advance our understanding of how memorization develops during training and offer insights for predicting and intervening on it.