发表机构
Intuition Machines Inc(直觉机器公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨微调中优化器历史导致语言模型遗忘的机制,通过分解损失和干预实验,发现历史梯度主要造成答案质量损失,而当前梯度起保护作用。
AI 中文摘要
在微调过程中,即使当前梯度作用于保持先前学习的答案的概率,语言模型也可能降低对这些答案的概率分配。借助动量,每次更新还会携带在较早模型状态下计算的梯度,这些存储的贡献可能将模型推向相反方向。我们通过将旧任务损失分解为答案间的混淆和答案集外的概率泄漏,来研究这种优化器记忆如何导致遗忘。在三个语言模型家族中,答案质量持续下降,而旧答案间的区分通常改善:模型变得更不可能产生它仍能正确区分的答案。分解Adam更新揭示了这种答案质量损失的相反贡献。在训练过程中,累积历史偏向泄漏,而当前梯度则反对泄漏。按年龄解析历史表明,有害贡献主要来自新任务的较旧梯度,而近期梯度倾向于保护旧答案。历史效应的变化主要由其相对于旧任务梯度的方向主导。在匹配初始更新范数的同时重置动量的干预措施确立了存储历史影响保留,具有状态依赖的即时效应和较长Adam延续中较低的最终旧任务损失,主要通过恢复答案质量实现。最后,沿有限更新的积分表明,大多数采样的大损失增加由局部投影捕获,而沿历史的曲率放大了一些事件。总之,这些发现揭示了优化器的记忆如何侵蚀已学习的行为,即使其当前梯度旨在保持它。
英文摘要
During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimiser memory contributes to forgetting by separating old-task loss into confusion among its answers and leakage of probability outside the answer set. Across three language-model families, answer mass consistently declines while discrimination among old answers usually improves: the model becomes less likely to produce answers that it can still distinguish correctly. Decomposing Adam updates reveals opposing contributions to this loss of answer mass. Over training, accumulated history favours leakage, while the current gradient opposes it. Resolving history by age shows that the harmful contributions come mainly from older gradients of the new task, whereas recent gradients tend to protect the old answers. Changes in history's effect are dominated by its orientation relative to the old-task gradient. Interventions that reset momentum while matching the initial update norm establish that stored history affects retention, with state-dependent immediate effects and lower final old-task loss over longer Adam continuations, mainly through recovered answer mass. Finally, integration along finite updates shows that most sampled large loss increases are captured by local projections, while curvature along history amplifies some events. Together, these findings reveal how an optimiser's memory can erode learned behaviour even as its current gradient acts to preserve it.