发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语言模型遗忘中遗忘集筛选缺失的问题,提出CleanSlate基准,发现不同遗忘集筛选器存在抑制效果弱或附带损伤的缺陷,表明遗忘集选择会影响遗忘效果与附带损伤。
AI 中文摘要
机器遗忘旨在从已训练好的模型中移除目标数据或行为,而无需从头开始重新训练。然而,大多数评估都假设要遗忘的示例是已知的。在实际的语言模型部署中,请求者可能会要求模型停止复制某首歌曲或某本书,却不知道万亿token语料库中哪些片段、文档、引用或近重复内容支撑了该行为。我们研究这一上游缺失问题,即遗忘集筛选:将抑制请求映射到传递给遗忘算法的数据。我们引入CleanSlate,一个针对歌曲和书籍逐字输出抑制的基准,包含特定于模型的提取轮廓、基于内容的问答以及能力保留评估。CleanSlate揭示了两种失败模式:自然词汇和精确子串筛选器通常会产生导致抑制效果较弱的遗忘集;评估感知筛选器几乎完全抑制了请求的延续,但会对非请求内容造成附带退化以及依赖于模型的能力损失。这些结果表明,一旦给定遗忘集,实际的遗忘不仅是一个优化问题:为遗忘选择的数据既决定了可以遗忘的内容,也决定了会被损坏的其他内容。
英文摘要
Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a requester may ask a model to stop reproducing a song or book without knowing which spans, documents, quotations, or near-duplicates in a trillion-token corpus support that behavior. We study this missing upstream problem, forget set curation: mapping a suppression request to the data passed to an unlearning algorithm. We introduce CleanSlate, a benchmark for verbatim output suppression over songs and books, with model-specific extraction profiles, content-grounded QA, and capability-retention evaluations. CleanSlate exposes two failure modes. Natural lexical and exact-substring curators often yield forget sets that lead to weak suppression. An evaluation-aware curator suppresses requested continuations almost completely, but causes collateral regression on non-requested content and model-dependent capability loss. These results show that practical unlearning is not only an optimization problem once a forget set is given: the data chosen for forgetting determines both what can be unlearnt and what else is damaged.
CommentsPresented at MemFM @ ICML 2026 and FoGen @ ICML 2026