狭义遗忘,广义保留:作为不对称泛化问题的模型遗忘
Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem
查看机构详情
- Tübingen AI Center, University of Tübingen(图宾根大学图宾根人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究大语言模型中模型遗忘的不对称泛化问题,提出SUITE评估协议和训练语料库,解决现有基准测试的不足,在此基础上引入JensUn++算法,实现更好的遗忘 - 保留效用权衡。
中文摘要 AI 辅助
大语言模型中的模型遗忘是指在保留所有其他能力的同时有针对性地去除特定知识,这对隐私和安全至关重要。然而,现有的基准测试对其测量并不可靠。它们存在“遗忘不足”的问题,即遗漏了在释义或间接查询下重新出现的知识;还存在“遗忘过度”的问题,即缺乏验证不相关知识是否被保留所需的语义、句法和词汇探测。这两种失败都反映了一个不对称泛化问题。遗忘评估必须涵盖同一目标事实的不同查询形式,测试遗忘是否超出精确的训练提示。保留评估必须探测一个大得多的隐式定义集,即与遗忘目标不相交的每个事实。保留集因此定义了有效的遗忘集,但当前数据集没有对这个遗忘 - 保留边界进行细粒度注释。我们通过SUITE来解决这个问题,SUITE是一种评估协议和训练语料库,用于捕获现实世界事实领域的遗忘 - 保留结构。在SUITE上训练的方法有显著改进,表明训练数据与算法设计同样重要。基于获得的见解,我们引入了JensUn++,这是一种遗忘算法,在顺序和联合遗忘设置中,在三个大语言模型中实现了最佳的遗忘 - 保留效用权衡。代码和数据集可在这个https网址获取。
英文摘要
Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities, critical for privacy and safety. Yet existing benchmarks measure it unreliably. They miss knowledge that resurfaces under paraphrased or indirect queries, a failure we call under-forgetting, and lack the semantic, syntactic, and lexical probes needed to verify that unrelated knowledge is preserved, a failure we call over-forgetting. Both failures reflect an asymmetric generalization problem. Forget evaluation must cover diverse query formulations of the same target facts, testing whether forgetting holds beyond exact training prompts. Retain evaluation must probe a far larger and implicitly defined set, namely every fact disjoint from the forget target. The retain set thus defines the effective forget set, yet current datasets provide no fine-grained annotation of this forget-retain boundary. We address this with SUITE, an evaluation protocol and training corpus that captures forget-retain structure for real-world factual domains. Methods trained on SUITE improve substantially, showing that training data is as important as algorithmic design. Building on the obtained insights, we introduce JensUn++, an unlearning algorithm that achieves the best forget-retain utility trade-off across three LLMs, in both sequential and joint unlearning settings. Code and datasets are available at https://amitpeleg.github.io/forget-narrowly-retain-broadly