发表机构
The University of Tokyo; KAIST(东京大学; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出FRAG预测器,基于权重选择性而非距离,结合FRP方法提升LLM再学习鲁棒性,揭示权重选择性对鲁棒性的解释力优于单独距离。
AI 中文摘要
机器遗忘旨在让模型遗忘特定数据,但被遗忘的大语言模型(LLMs)往往无法保持遗忘状态:简短的微调就能恢复被移除的知识。现有的鲁棒性预测器依赖全局权重空间位移,但当随机或破坏性更新导致性能崩溃时,仅靠距离可能产生误导。我们认为再学习鲁棒性取决于更新结构:鲁棒的遗忘应影响遗忘关键权重,同时保留保留关键权重。我们提出遗忘-保留对齐间隙(FRAG),这是一种无需训练的预测器,可在不运行再学习攻击的情况下对更新的遗忘-保留对齐进行评分,且比全局距离更可靠地将选择性更新与密集更新区分开。基于“遗忘关键、保留 sparing”原则,遗忘剪枝(FRP)提升了再学习鲁棒性。我们的结果表明,权重选择性比单独的距离更能解释鲁棒性。代码可在该 https URL 获取。
英文摘要
Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget-critical weights while sparing retain-critical ones. We introduce the Forget-Retain Alignment Gap (FRAG), a training-free predictor that scores an update's forget-retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance. Building on the forget-critical, retain-sparing principle, Forget-Retain Pruning (FRP) improves relearning robustness. Our results suggest that weight selectivity better explains robustness than distance alone. Code is available at https://github.com/Yi1-Chen/FRAG.
CommentsAccepted to EMNLP 2026 Main Conference