arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

距离并不足够:遗忘-保留对齐间隙可预测大语言模型的再学习鲁棒性

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

Yi Chen, Hanna Hsieh, Shuhong Liu, Chuanbo Hua, Zihan Ma, Kun Wang, Joo-Young Kim

arXiv 2608.25429首次发表:更新:

发表机构

The University of Tokyo; KAIST(东京大学; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出FRAG预测器,基于权重选择性而非距离,结合FRP方法提升LLM再学习鲁棒性,揭示权重选择性对鲁棒性的解释力优于单独距离。

AI 中文摘要

机器遗忘旨在让模型遗忘特定数据,但被遗忘的大语言模型(LLMs)往往无法保持遗忘状态:简短的微调就能恢复被移除的知识。现有的鲁棒性预测器依赖全局权重空间位移,但当随机或破坏性更新导致性能崩溃时,仅靠距离可能产生误导。我们认为再学习鲁棒性取决于更新结构:鲁棒的遗忘应影响遗忘关键权重,同时保留保留关键权重。我们提出遗忘-保留对齐间隙(FRAG),这是一种无需训练的预测器,可在不运行再学习攻击的情况下对更新的遗忘-保留对齐进行评分,且比全局距离更可靠地将选择性更新与密集更新区分开。基于“遗忘关键、保留 sparing”原则,遗忘剪枝(FRP)提升了再学习鲁棒性。我们的结果表明,权重选择性比单独的距离更能解释鲁棒性。代码可在该 https URL 获取。

英文摘要

Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget-critical weights while sparing retain-critical ones. We introduce the Forget-Retain Alignment Gap (FRAG), a training-free predictor that scores an update's forget-retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance. Building on the forget-critical, retain-sparing principle, Forget-Retain Pruning (FRP) improves relearning robustness. Our results suggest that weight selectivity better explains robustness than distance alone. Code is available at https://github.com/Yi1-Chen/FRAG.

CommentsAccepted to EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑