用于机器遗忘的因果路由
Causal Routing for Unlearning
- Sapienza University of Rome(罗马大学)
- Aarhus University(奥胡斯大学)
- ISTC-CNR(意大利国家研究委员会认知科学与技术研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出因果路由遗忘(CRU),通过定位并仅抑制模型中表达目标概念的神经元,实现高效且不损害模型整体效用的机器遗忘,显著降低计算成本并提升遗忘效果。
AI中文摘要:
大语言模型无法像删除文件那样进行遗忘。奇怪的是,我们被要求移除某些从未被特别放置在某处的东西。模型从一段文本中获取的信息现在散布在数十亿个权重中。现有方法为了改变一件事而重写所有这些权重,且没有一种方法能指出是哪个部分产生了这种改变。为解决此问题,我们引入了用于机器遗忘的因果路由(CRU),通过询问概念在模型中的表达位置,仅抑制该部分。对遗忘集进行一次未训练的向前传播,根据神经元激活的变化程度对它们进行排序。然后,这些神经元上的小型路由模块对需要遗忘的概念进行门控和抑制。在CRU中,基础模型被冻结,任何行为变化仅由被门控的神经元引起;因此,路由是因果的。由于我们的参数效率(仅为基础模型参数的约0.01%),遗忘一个概念的成本为14 GiB,而基线方法需要71 GiB。在TOFU上,CRU与保留模型无法区分(p > 0.05,KS检验),且从未被帕累托支配,而每个比较的基线仅在较大的遗忘批次上通过牺牲实用性来匹配其遗忘效果。在RWKU上,它实现了0.052的对抗性探针召回率,而最强基线的召回率为0.250,这意味着知识已被移除,而不仅仅是更难访问。因此,在查询时决定干预,而非事先固定干预,是我们认为机器遗忘应沿此方向发展的关键轴。
英文摘要:
LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.