发表机构
MPI for Intelligent Systems, Tübingen; Jinesis Lab, University of Toronto & Vector Institute; EuroSafeAI; Hector Foundation; Stanford University; ELLIS Institute Tübingen(马克斯·普朗克智能系统研究所(蒂宾根); Jinesis实验室,多伦多大学与向量研究所; EuroSafeAI; 赫克托基金会; 斯坦福大学; ELLIS研究所(蒂宾根))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究微调中语言模型的虚假遗忘机制,发现遗忘包含可逆的共享访问丧失与不可逆的个体事实侵蚀,后者才是灾难性的,其主导性取决于新数据对旧记忆的移动方式。
AI 中文摘要
语言模型在微调过程中看似遗忘的知识往往仍然被存储且可以恢复,这一现象被称为虚假遗忘。在新事实上的微调甚至可能产生自我抵消的遗忘:旧事实的召回能力崩溃,随着仅在新事实上的训练继续而恢复,然后才最终永久消退。我们试图理解这种遗忘何时并非灾难性的。一个最小化的联想记忆模型用三个要素再现了这些动态:具有共享结构的键、集中的新值和网络中的归一化。微调将所有旧表征沿一个共同方向移动,隐藏旧事实的同时保留其相对几何结构;一旦新事实被学习,归一化会撤回这一偏移,而事实特定的变化则累积并导致侵蚀。此外,在合成数据上训练的Transformer中减去共同偏移可消除崩溃,而在预训练语言模型中从每次权重更新中移除单一方向可恢复旧事实。因此,遗忘结合了共享的、可逆的访问丧失与个体事实的缓慢侵蚀,只有后者才是灾难性的。哪一者占主导取决于新数据是将旧记忆一起移动还是分开移动。
英文摘要
Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.