arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型真的会遗忘吗?模型遗忘中的隐藏状态泄漏及其修复方法

Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it

Hadi Reisizadeh, Jiajun Ruan, Sijia Liu, Mingyi Hong

arXiv 2609.36612首次发表:更新:

发表机构

University of Minnesota; Michigan State University; IBM Research(明尼苏达大学; 密歇根州立大学; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM遗忘评估仅关注输出层面的局限,提出探针对抗表示抑制(PARS)方法,直接最小化隐藏表示中的信息泄漏,以实现在对抗性攻击下更强的真正遗忘保证。

AI 中文摘要

大型语言模型(LLM)中的遗忘通常在输出层面进行评估,此时模型似乎抑制了敏感或不受欢迎的内容。在本工作中,我们表明这种评估可能造成遗忘的错觉:即使输出层面的泄漏被消除,敏感信息仍可能编码在模型的隐藏表示中。我们首先提供理论分析,确立输出抑制与表示擦除之间的根本分离。具体来说,我们表明解码器可以被使得对敏感方向任意不敏感,将输出层面的泄漏降至零,而隐藏表示仍保留底层信息。为了实证验证这一现象,我们在跨变压器层的隐藏状态上训练生成式探针解码器,从而实现逐层的信息泄漏测量。在三个广泛使用的基准测试(TOFU、MUSE和WMDP)以及最先进的遗忘方法中,我们发现即使标准输出层面指标表明遗忘成功,大量敏感信息仍可从隐藏表示中恢复。为了解决这一差距,我们提出了探针对抗表示抑制(PARS),这是一种遗忘目标,对抗性地最小化从隐藏表示中可提取的信息。PARS直接针对表示泄漏,并在对抗性探针和重学习攻击下提供显著更强的擦除保证,优于所有评估的基线。我们的结果突显了现有遗忘范式的基本局限性,并表明在LLM中真正的遗忘不仅需要控制模型输出,还需要控制隐藏表示中编码的信息。代码可在以下网址获取:此https URL。

英文摘要

Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at https://github.com/OptimAI-Lab/HiddenStateUnlearning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑