坦白你所知道的:大语言模型遗忘中的遗忘集与模型知识的错位
Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning
- Dongguk University-Seoul(东国大学首尔校区)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM遗忘中遗忘集与模型知识错位的问题,提出数据盲框架CONFS,通过构建与模型对齐的遗忘集,实现了良好的遗忘-效用平衡,优于其他数据盲遗忘集构建方法。
AI中文摘要:
大语言模型(LLM)的机器遗忘通常假设预定义的遗忘集与模型已记忆的内容匹配,但在无法获取原始训练数据的现实隐私场景中,这种假设常不成立。我们将此差距称为遗忘集错位,并识别出两种情况:在欠遗忘中,遗忘集遗漏了已记忆的信息,导致泄露持续存在;在知识外遗忘中,算法被驱动“遗忘”模型从未学习过的知识,从而扰动参数并降低效用。通过梯度层面分析,我们表明这些行为源于遗忘目标的错位,而非特定的优化选择。随后,我们提出CONfession-to-Forget-Set(CONFS),这是一种数据盲框架,通过引出并形式化模型的已记忆知识,构建与模型对齐的遗忘集。在合成、多模态和真实世界基准测试中,CONFS在多项指标上接近黄金标准性能,实现了具有竞争力的遗忘-效用平衡,同时比其他数据盲遗忘集构建方法更好地保留了效用。
英文摘要:
Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. We term this gap forget-set misalignment and identify two cases. In Under Unlearning, the forget set omits memorized information and leakage persists. In Out-of-Knowledge Unlearning, the algorithm is driven to "forget" knowledge the model never learned, perturbing parameters and degrading utility. Using gradient-level analysis, we show these behaviors arise from misaligned unlearning targets rather than specific optimization choices. We then propose CONfession-to-Forget-Set (CONFS), a data-blind framework that constructs model-aligned forget sets by eliciting and formalizing the model's memorized knowledge. Across synthetic, multimodal, and real-world benchmarks, CONFS approaches Gold-standard performance on several metrics and achieves a competitive forgetting-utility balance, while preserving utility better than other data-blind forget-set constructions.