纠缠表征会放大遗忘中的附带损害
Entangled Representations Amplify Collateral Damage in Unlearning
浏览论文内容
中文总结 AI 辅助
该研究通过控制SGTM训练的不同解纠缠程度的254M参数语言模型,验证了表征纠缠会放大遗忘附带损害,解纠缠程度更高的模型保留-遗忘权衡更优。
中文摘要 AI 辅助
可解释性研究中一个长期存在的直觉是,表征纠缠(即神经网络中知识域之间的结构共享)会使遗忘变得更困难。尽管该直觉被广泛认可,但从未在受控实验中得到直接验证。我们提出了一种验证方法:通过重新利用选择性梯度掩码(Selective Gradient Masking,SGTM),我们在英文维基百科上训练了一组6个参数规模为254M的语言模型,这些模型在生物学与非生物学知识之间具有不同程度的解纠缠。对该组中的每个模型应用三种标准遗忘方法后,我们发现解纠缠程度更高的模型始终能实现更好的保留-遗忘权衡:在固定遗忘水平下,对于三种方法中的两种,解纠缠程度最高的模型产生的保留代价约低4倍;对于第三种方法,保留代价约低1.3倍。由于我们的干预仅改变模型,未改变数据或遗忘算法,这直接证明了表征纠缠是遗忘中附带损害的原因之一,正如可解释性研究人员长期以来所怀疑的那样。类似的设计可用于验证可解释性的其他结构主张。
英文摘要
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly tested in a controlled experiment. We present a way to do so: by repurposing Selective Gradient Masking (SGTM), we train a suite of six 254M-parameter language models on English Wikipedia with graded levels of disentanglement between biology and non-biology knowledge. Applying three standard unlearning methods to every model in the suite, we find that more disentangled models consistently achieve better retain-forget trade-offs: at a fixed level of forgetting, the most disentangled models incur roughly $4\times$ lower retain cost under two of the three methods, and $1.3\times$ lower under the third. Because our intervention changes only the model, not the data or the unlearning algorithm, this is direct evidence that representational entanglement is one of the causes of collateral damage in unlearning, as interpretability researchers have long suspected. A similar design could be used to test other structural claims from interpretability.
发表机构
- University of Oxford(牛津大学)
- University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。