arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRACE:面向大语言模型遗忘的梯度引导核心集选择方法

GRACE:Gradient-guided Coreset Selection for LLM Unlearning

Praveen Bushipaka, Andrea D'Angelo, Lucia Passaro, Tommaso Cucinotta

arXiv 2608.28361首次发表:更新:

发表机构

University of Pisa; Aarhus University(比萨大学; 奥胡斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型遗忘中需从异构语料推断遗忘集和保留集的问题,提出GRACE梯度引导核心集选择方法,通过梯度计算与匹配追踪选择核心集,在多领域、多模型、多算法实验中提升模型效用且保持遗忘质量。

AI 中文摘要

大语言模型的机器遗忘方法通常假设存在预先指定的遗忘集和保留集。但在实际场景中,用户请求可能仅提供少量不良行为示例,需要从异构语料库中推断出遗忘集和保留集。我们研究这一数据选择问题,提出GRACE(Gradient-guided Coreset Selection,梯度引导核心集选择)方法,用于构建大语言模型遗忘的遗忘集和保留集。GRACE首先从引发不良行为的种子示例中计算遗忘方向,再使用非负正交匹配追踪选择紧凑的遗忘核心集,其梯度可近似该方向。为保留模型效用,GRACE在投影去除遗忘方向后的梯度空间中,使用聚类正交匹配追踪选择保留示例。在两个目标领域、两个模型家族和四种遗忘算法的实验中,GRACE在保持相当遗忘质量的同时提升了模型效用,相较于现有基于梯度的选择方法表现出更稳定的性能提升。

英文摘要

Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests may provide only a few examples of undesired behavior, requiring forget and retain sets to be inferred from heterogeneous corpora. We study this data-selection problem and propose GRACE , a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning. GRACE first computes a forget direction from seed examples that elicit the undesired behavior, then selects a compact forget coreset whose gradients approximate this direction using non-negative orthogonal matching pursuit. To preserve model utility, it selects retain examples after projecting out the forget direction and applying clustered orthogonal matching pursuit in the remaining gradient space. Across two target domains, two model families, and four unlearning algorithms, GRACE improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.

Comments20 pages, 16 tables, 5 figures, accepted to EMNLP Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑