让大语言模型真正遗忘:通过搜索、选择与切断知识路径实现深度遗忘
Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths
浏览论文内容
中文总结 AI 辅助
本文提出一种深度遗忘框架,通过搜索、选择并切断知识路径来打破关系结构,实现大语言模型的真正遗忘,同时保持模型效用。
中文摘要 AI 辅助
虽然被遗忘的语言模型可能不再直接回忆某个事实,但该事实往往仍可通过相关知识的多次推理得以恢复。大多数现有的遗忘技术忽视了这一漏洞,孤立地针对事实进行遗忘,却使其支撑知识保持完整。为实现真正的遗忘,我们提出了一种与现有遗忘算法兼容的通用深度遗忘框架。我们的方法自适应地探索显式响应和潜在内部表征,以发现有效的推理路径,将其编译为置信度感知的支撑子图,并应用图最小割来切断所有恢复路径,同时保留无关知识。为严格评估深度遗忘,我们引入了一个模型特定的流程,从原始文本中提取并完善知识图谱,并通过校准的模型置信度对其进行过滤,以反映模型真正保留的内容。综合实验表明,选择性地遗忘支撑知识比表面方法产生更深的遗忘效果,同时保持模型效用,突显了真正的遗忘需要打破使事实重建成为可能的关系结构。
英文摘要
While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact. To achieve true forgetting, we propose a general deep unlearning framework compatible with existing unlearning algorithms. Our approach adaptively explores both explicit responses and latent internal representations to discover valid reasoning paths, compiles them into a confidence-aware supporting subgraph, and we apply a graph minimum cut to sever all recovery paths while preserving unrelated knowledge. To rigorously evaluate deep unlearning, we introduce a model-specific pipeline that extracts and completes knowledge graphs from raw text, filtering them by calibrated model confidence to reflect what the model genuinely retains. Comprehensive experiments demonstrate that selectively unlearning supporting knowledge yields substantially deeper forgetting than superficial methods while preserving model utility, highlighting that genuine unlearning requires breaking the relational structures that enable factual reconstruction.
发表机构
- Tongji University(同济大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Purdue University(普渡大学)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。