发表机构
University of Electronic Science and Technology of China; Southwest Jiaotong University(电子科技大学; 西南交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对大语言模型遗忘的附带变更问题,提出PTP-U框架,结合目标遗忘与非目标能力恢复,在三个基准测试中实现了最优的遗忘-保留权衡。
AI 中文摘要
大语言模型的机器遗忘旨在移除不需要的知识,同时保留模型的其余能力。尽管现有方法采用保留目标或限制编辑发生的位置,但达到期望的遗忘水平仍可能留下损害非目标行为的附带变更。我们的恢复比较表明,其中一些变更可以被逆转,同时保留观察到的遗忘性能。在这项工作中,我们提出了Propose-Then-Project Unlearning(PTP-U),这是一个将目标遗忘与非目标能力恢复相结合的框架。PTP-U首先应用局部分析编辑来削弱目标知识关联,然后将非目标输出分布与原始模型的输出分布对齐,以恢复能力,同时保持固定的遗忘约束。两个阶段都服务于共同目标:满足遗忘要求,同时保留流畅的生成和非目标任务上的性能。在三个基准测试中,PTP-U在评估的方法中实现了最强的遗忘-保留权衡,达到了81.22%-91.03%的遗忘率,同时平均保留了94.20%的非目标效用。在匹配的遗忘水平下,PTP-U始终保留更高的非目标效用。
英文摘要
Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.