arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

存储并非策略:面向LLM遗忘的状态条件支持控制

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang

arXiv 2609.37858首次发表:更新:

发表机构

Southeast University; University of Sydney; Northwestern University; Together AI; The Chinese University of Hong Kong Shenzhen(东南大学; 悉尼大学; 西北大学; Together AI; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM遗忘中固定参数子集并非最优的问题,提出干预分数与动态重排序方法,在多个基准上显著优于现有基线,支持定位、干预选择与支持修正的分离。

AI 中文摘要

许多局部化的大语言模型(LLM)遗忘方法会从定位信号中选取一小部分参数子集,并在优化过程中保持该子集固定不变。然而,与目标最相关的参数未必是最适合更新的参数,且候选干预措施的价值会随优化进程而改变。在一项受控实验中,存储定位分数的受试者工作特征曲线下面积(AUROC)达到0.981,但存储身份与更优干预措施仅在36个目标中的17个上一致,而低秩适配(LoRA)在36个目标中赢得35个。我们引入了干预分数(Intervention Score),它根据实际遗忘更新的预测效果对可编辑组进行排序,同时考虑附带损害,并据此构建静态干预值基线(Static-IV)。随后,我们引入了选择性动态干预重排序(DIR-R),仅在校准探针证明比较合理时才重新审视该子集。在Natural-TOFU数据集上,我们的方法在19/20的方法与目标比较中具有正的描述性边际,尽管其中若干边际接近零。在LACUNA定位精度基准上,我们的平均终端效用在所有六项负偏好优化(NPO)和SimNPO比较中均更高:NPO边际范围为+0.431至+0.848,SimNPO边际范围为+0.503至+0.571。梯度差(GradDiff)目标揭示了显著的场依赖性。相对于Static-IV,主要的四场GradDiff评估取得六胜、六平、零负,平均和成对中位数增益分别为+0.165和+0.0025。这些证据支持将定位、初始干预选择和基于检查点的支持修正分离开来。

英文摘要

Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.

Comments18 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑