arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11408cs.CL

测量,而非优化:预测大语言模型遗忘中的恢复情况

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

  • AMAP, Alibaba Group(阿里巴巴集团 AMAP)

机构由 AI 辅助整理,请以论文原文为准。

Zirui Song, Huaxing Liu, Xiang Wang, Shuai Li, Xinye Li, Lang Gao, Jinghui Zhang, Zheng Lu, Fengxian Ji, Xiaojun Chang, Xiuying Chen

AI总结:

该研究提出推理时审计方法J-Access,对398个采用8种遗忘方法的公开模型审计后发现,J-Access可评估模型残留知识恢复易感性,直接最小化J-Access无法促进真正删除,内部审计不应随意转为优化目标。

AI中文摘要:

现有白盒研究表明,大语言模型在遗忘后会保留目标知识的潜在痕迹,即便这些知识不再出现在模型的输出中。然而,现有的审计仅局限于一次性诊断:目前尚不清楚这些残留信号能否预测持续训练下的未来恢复情况,或能否作为可靠的优化目标。解决这一差距对于确定内部审计能否从事后评估转向主动风险监测及更安全的遗忘至关重要。我们提出J-Access,一种推理时审计方法,该方法使用雅可比视角将中间表示映射到词汇空间,并测量目标概念在模型输出路径中保持可访问的频率。我们假设残留可访问性反映了恢复易感性:更靠近输出路径的知识只需更少的微调即可恢复,从而恢复速度更快。我们审计了涵盖8种遗忘方法的398个公开遗忘模型,发现:(1)大多数遗忘模型的可访问性高于仅保留数据的基准水平;(2)攻击前的可访问性在模型层面可预测恢复速度和程度,但无法识别哪些特定事实会被恢复;(3)直接最小化J-Access并不能促进真正的删除,相反,模型会学会向审计隐藏知识,从而产生更低的审计分数但更大的攻击后恢复。这些发现将J-Access定位为评估遗忘模型残留易感性的模型级诊断工具。我们认为内部审计应作为遗忘评估中的独立诊断维度,未经验证不得转化为优化目标。

英文摘要:

Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outputs. However, existing audits remain limited to one-off diagnostics: it is unclear whether these residual signals can predict future recovery under continued training or serve as reliable optimization targets. Resolving this gap is essential to determine whether internal auditing can move beyond post-hoc evaluation toward proactive risk monitoring and safer unlearning. We propose J-Access, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway. We hypothesize that residual accessibility reflects recovery susceptibility: knowledge that remains closer to the output pathway requires less fine-tuning to restore, leading to faster recovery. We audit 398 public unlearned models spanning eight unlearning methods. We find that: (1) most unlearned models retain access above the retain-only gold level; (2) pre-attack accessibility predicts recovery speed and extent at the model level, but cannot identify which specific facts will be recovered; and (3) directly minimizing J-Access does not promote genuine deletion. Instead, the model learns to hide knowledge from the audit, producing lower audit scores but greater post-attack recovery. These findings position J-Access as a model-level diagnostic for assessing residual susceptibility in unlearned models. We argue internal audits should serve as an independent diagnostic dimension in unlearning evaluation, and should not be converted into optimization targets without validation.

补充信息

↑