arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习遗忘什么:LLM表示空间的分布遗忘

Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

Pinaki Mohanty, Haoran Tang, Maggie Makar, Rajiv Khanna

arXiv 2609.38929首次发表:更新:

发表机构

Purdue University; University of Michigan(普渡大学; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Mamushi框架,用于非参数分布遗忘,通过概率分类器排序遗忘示例,实现最优选取,并在有毒语言和主题域移除任务中优于基线,减少下游遗忘所需示例数。

AI 中文摘要

机器学习系统日益面临需要移除整个数据领域(如有毒语言、有害行为或主题内容)的影响,而不仅仅是孤立记录的需求。近期工作将此问题形式化为“分布遗忘”:选择遗忘域的一个子集,其移除使训练分布远离不想要的群体,同时保持与期望群体的接近性。然而,现有分析常常施加参数化假设以获得易处理的选取规则。这些假设可能不适用于高维语言模型表示。我们引入Mamushi,一个用于非参数分布遗忘的框架,它使用概率分类器对遗忘示例进行排序,该分类器的贝叶斯最优logit等于遗忘与保留对数密度比(加上一个类别先验常数)。我们证明,对总体对数密度比进行阈值化可产生我们移除-保留目标的最优固定预算选取规则,并建立了一个非渐近转移保证,将分数估计和阈值校准误差与总体最优选取规则的退化联系起来。我们的实证评估涵盖使用不同表示的有毒语言移除和主题域移除机制的真实世界数据集,Mamushi实现了比其他基线更有利的移除-保留权衡。我们的工作表明,Mamushi可作为下游机器遗忘程序的高效选取方法,减少达到固定遗忘目标所需的遗忘示例数量。

英文摘要

Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabilistic classifier whose Bayes-optimal logit equals the forget-to-retain log-density ratio (up to an additive class-prior constant). We show that thresholding the population log-density ratio yields the optimal fixed-budget selection rule for our removal--preservation objective and establish a non-asymptotic transfer guarantee relating score-estimation and threshold-calibration errors to degradation from the population-optimal selection rule. Our empirical evaluation spans real-world datasets on toxic-language removal and topical-domain removal regimes using different representations, with \textsc{Mamushi} achieving a more favorable removal--preservation trade-off than other baselines. Our work shows that \textsc{Mamushi} can serve as an efficient selection approach for downstream machine unlearning procedures, reducing the number of forget examples required to reach a fixed forgetting target.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑