发表机构
University of North Carolina at Chapel Hill; Honda Research Institute; Michigan State University(北卡罗来纳大学教堂山分校; 本田研究所; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ARIA,一种基于稀疏自编码器的测试时遗忘方法,通过轻量线性检测器门控遗忘状态,在保持模型权重不变的同时,显著改善遗忘-保留权衡,并抵抗多种遗忘后攻击。
AI 中文摘要
机器遗忘旨在从已训练的大语言模型(LLM)中移除特定知识,而无需从头重新训练。现有方法通过梯度上升及其改进来修改模型权重。虽然这些基于权重的方法在某些基准上有效,但它们表现出尖锐的遗忘-效用权衡,即对目标知识的更强遗忘可能降低模型效用,且被遗忘的知识可能在遗忘后的微调或提示攻击下重新出现。我们提出ARIA(自编码器门控的推理时遗忘),一种测试时遗忘方法,它保持模型权重不变,并仅在生成进入与遗忘相关的状态时门控对不需要知识的访问。ARIA使用稀疏自编码器(SAE)潜变量来训练一个轻量级线性检测器,然后对触发状态应用可解释的干预,测试时开销可忽略不计。在TOFU、R-TOFU和WMDP上的实证评估表明,ARIA在思考模型(DeepSeek-R1-Distilled-Qwen-1.5B)和指令模型(Gemma-3-1B-it)上均改善了遗忘-保留权衡,优于基于权重的基线,例如,显著降低WMDP-cyber遗忘集的准确率,同时将MMLU保持在未遗忘模型的1%以内。我们进一步引入了三种针对权重空间和解码空间恢复的遗忘后对抗攻击,并发现ARIA在所有三种攻击下均保持稳健,攻击下遗忘变化小于1%。利用ARIA可解释性的特征级案例研究表明,某些保留退化可能反映了遗忘数据背后的响应风格,而非目标知识本身的泄漏,这突显了遗忘任务构建中潜在的偏差来源。
英文摘要
Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.