AI 中文总结
提出GRAPHSU图引导控制器,扩展语言模型选择性遗忘的删除范围,在TOFU和PISTOL基准上,相较仅种子基准降低最多49.5个百分点的软泄漏,验证控制支持路径的重要性。
AI 中文摘要
企业会在专有数据上微调语言模型,但这些数据可能因隐私、合同或合规义务后续需要移除。选择性遗忘可在移除请求知识的同时保留模型效用,是完全重新训练的实用替代方案,但现有方法将明确识别的遗忘样本视为完整删除范围,当目标知识可通过释义、别名或相邻训练样本恢复时,该方法不足。我们提出GRAPHSU,一种图引导控制器,通过构建加权支持路径图、在图中传播删除压力,并对高风险邻居应用分级遗忘强度,将删除范围扩展至遗忘种子之外。在虚构遗忘任务(TOFU,合成作者资料问答基准)和PISTOL(围绕相互关联事实样本构建的结构遗忘基准)上,使用GPT-2 Medium和Llama-3.2-3B-Instruct模型,GRAPHSU在所有删除设置中实现了最低的效用可行软泄漏,相较于匹配的仅种子基准,泄漏最多降低49.5个百分点,证明有效的企业遗忘需要控制支持路径,而非仅遗忘种子。
英文摘要
Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligations. Selective unlearning removes requested knowledge while preserving model utility, offering a practical alternative to full retraining, but existing methods treat the explicitly identified forget examples as the complete deletion scope. This is insufficient when target knowledge remains recoverable through paraphrases, aliases, or neighboring training examples. We propose GRAPHSU, a graph-guided controller that expands the deletion scope beyond forget seeds by constructing a weighted support-route graph, propagating deletion pressure through it, and applying graded forgetting strengths to high-risk neighbors. On the Task of Fictitious Unlearning (TOFU), a synthetic author-profile question-answering benchmark, and PISTOL, a structural-unlearning benchmark built around interconnected factual samples, with GPT-2 Medium and Llama-3.2-3B-Instruct, GRAPHSU achieves the lowest utility-feasible soft leakage across all deletion settings, reducing leakage by up to 49.5 percentage points over a matched seed-only baseline, demonstrating that effective enterprise unlearning requires controlling support routes, not just forget seeds.