学习如何在记忆-泛化谱系中支配遗忘
How Learning Governs Unlearning across the Memorization-Generalization Spectrum
浏览论文内容
中文总结 AI 辅助
本文研究模型学习方式(记忆与泛化)对遗忘效果的影响,发现泛化型模型遗忘时保留损伤更大,并验证了该趋势在大型语言模型中的适用性,为改进遗忘方法提供了见解。
中文摘要 AI 辅助
虽然遗忘旨在消除通过学习获得的不良能力,但很少有研究探讨模型的学习方式如何影响其后续的遗忘。在本文中,我们从记忆和泛化这两个模型在训练过程中采用的最具代表性但相互竞争的策略的角度来研究这种联系。我们首先使用模加中的顿悟现象对记忆型和泛化型模型进行分类,并比较它们对遗忘的反应,结果表明后者遭受更大的保留损伤,即在保留集上的性能下降更大。此外,我们通过引入分桶模加来进行更细粒度的分析,在该分析中,可以在记忆-泛化谱系中显式控制这两种策略的各自贡献。在这种设置下,我们再次确认相同的趋势持续存在且几乎单调。我们进一步证明,这种关系也适用于大型语言模型在逐字和事实回忆设置中的遗忘。最后,我们为开发更好的遗忘方法提供了两个实用见解,强调了在遗忘中考虑学习动态的重要性。
英文摘要
While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.
发表机构
- Seoul National University(首尔大学)
- Hanyang University(汉阳大学)
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。