arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GROM:无梯度快速一次性机器遗忘

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

Paweł Batorski, Przemysław Spurek, Paul Swoboda

arXiv 2608.05783首次发表:更新:

AI 中文总结

GROM是一种无梯度快速一次性机器遗忘方法,通过闭式加性更新实现高效知识移除,在多数据集上取得最优遗忘-效用权衡,且能抵御低比特量化攻击。

AI 中文摘要

机器遗忘已成为从大型语言模型(LLMs)中安全移除特定敏感知识的关键能力。当前最先进的方法主要依赖迭代的训练时遗忘,通过微调实现。然而,即使利用LoRA等参数高效降维技术,基于梯度的优化仍计算成本高昂且缺乏明确的解析公式,还可能仅隐藏而非移除目标知识,甚至对遗忘后的模型进行简单量化就能恢复大量本应擦除的内容。为解决这一问题,我们提出一种新型一次性遗忘方法,摒弃迭代优化,转而采用直接、精确的解析解。我们将遗忘过程构建为岭正则化最小二乘优化问题,推导得到目标权重矩阵的闭式加性更新规则。该更新规则迫使选定层抑制不需要的内容,同时严格保留其在保留数据上的行为。GROM仅通过无梯度前向传播计算,无需反向传播,也无需迭代至收敛,仅需数秒即可完成权重编辑,比传统微调快几个数量级。大量评估表明,GROM在TOFU-5%、TOFU-10%、MUSE-Books、MUSE-News和WMDP数据集上实现了最先进的遗忘-效用权衡,显著降低了计算开销且未牺牲整体模型性能。由于该更新从权重中移除目标内容而非对其进行掩码,GROM还能抵御低比特量化攻击,而基于梯度的基线方法看似遗忘的内容会被该攻击恢复。我们的代码公开于此https URL。

英文摘要

Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state-of-the-art approaches primarily rely on iterative, training-time unlearning via fine-tuning. However, even when utilizing parameter-efficient dimensionality reduction techniques like LoRA, gradient-based optimization remains computationally expensive and lacks explicit analytical formulations. It can also leave the targeted knowledge merely hidden rather than removed, to the point that simply quantizing the unlearned model restores much of what it was supposed to have erased. To resolve this, we propose a novel one-shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution. We frame the unlearning process as a ridge-regularized least-squares optimization problem, deriving a closed-form additive update for targeted weight matrices. This update forces the selected layer to suppress unwanted content while strictly preserving its behavior on retained data. Computed from gradient-free forward passes alone, with no backpropagation and no iteration to convergence, GROM applies the weight edit in mere seconds, which makes it orders of magnitude faster than traditional fine-tuning. Extensive evaluations demonstrate that GROM achieves state-of-the-art forgetting-utility trade-offs on TOFU-5%, TOFU-10%, MUSE-Books, MUSE-News and WMDP, significantly reducing computational overhead without sacrificing overall model performance. Because the update removes the targeted content from the weights instead of masking it, GROM also withstands the low-bit quantization attack that recovers much of the content a gradient-based baseline had appeared to forget. Our code is publicly available at https://github.com/Batorskq/GROM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑