发表机构
University of Massachusetts Amherst; Microsoft Corporation(马萨诸塞大学阿默斯特分校; 微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对TOFU数据集和Llama-3.2-3B-Instruct模型,引入ASR指标评估遗忘方法,发现标准指标表现优异的遗忘方法仍存在高对抗恢复风险,提出需补充对抗性压力测试作为遗忘评估的必要部分。
AI 中文摘要
机器遗忘旨在从模型中移除目标训练数据的影响,同时保留其剩余能力,但评估此类信息是否真正无法获取仍具挑战性。现有基准主要在干净、非对抗性查询下评估遗忘,而未解决看似已遗忘的信息是否仍可通过策略性提示恢复的问题。我们针对TOFU数据集,基于Llama-3.2-3B-Instruct模型,对基于提示和基于微调的遗忘方法进行统一评估,随后对在标准指标下表现优异的方法开展对抗性鲁棒性评估。我们引入攻击成功率(Attack Success Rate, ASR)这一基于大语言模型的评判指标,用于衡量对抗性响应中泄漏分数超过0.2的比例,并通过8个攻击套件评估恢复情况。结果显示,干净查询遗忘与对抗性鲁棒性之间存在显著差距:尽管若干基于微调的方法达到了0.91以上的遗忘质量,但目标信息仍可被恢复,其ASR介于72.8%至84.3%之间,接近未受保护的基础模型87.5%的ASR;相比之下,干净的多语言改写仅产生2.95%的测得泄漏。人工审核进一步发现,二元ASR决策与人类事实评估在10个案例中有7个一致,表明ASR提供了有用但不完美的行为可恢复性信号。这些发现表明,仅靠标准指标下的优异表现不足以确立遗忘后的鲁棒性,推动了对抗性压力测试作为遗忘评估的补充组成部分。
英文摘要
Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whether such information has truly become inaccessible remains challenging. Existing benchmarks primarily assess unlearning under clean, non-adversarial queries, leaving open whether information that appears forgotten can still be recovered through strategic prompting. We address this gap through a unified evaluation of prompt-based and fine-tuning-based unlearning methods on TOFU using Llama-3.2-3B-Instruct, followed by an adversarial robustness evaluation of methods that perform strongly under standard metrics. We introduce Attack Success Rate (ASR), an LLM-as-judge metric that measures the fraction of adversarial responses whose leakage score exceeds $0.2$, and evaluate recovery across eight attack suites. Our results reveal a substantial gap between clean-query forgetting and adversarial robustness. Although several fine-tuning-based methods achieve Forget Quality above $0.91$, targeted information remains recoverable with ASRs between $72.8\%$ and $84.3\%$, close to the $87.5\%$ ASR of the unprotected base model. In contrast, clean multilingual reformulations yield only $2.95\%$ measured leakage. A manual audit further finds agreement between binary ASR decisions and human factual assessments in seven of ten cases, indicating that ASR provides a useful, though imperfect, signal of behavioral recoverability. These findings show that strong standard-metric performance alone is insufficient to establish robustness after unlearning and motivate adversarial stress-testing as a complementary component of unlearning evaluation.
Comments19 pages, 5 figures