发表机构
DIBRIS – Department of Informatics, Bioengineering, Robotics and Systems Engineering, University of Genova; Centre for Defense Higher Studies (CASD)(热那亚大学信息、生物工程、机器人与系统工程系(DIBRIS); 国防高等研究中心(CASD))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究机器遗忘中基于梯度显著性的权重选择对表示级遗忘的贡献,通过可控消融实验比较不同掩码策略,发现表示级遗忘主要受梯度集中和表示几何控制,有效遗忘需直接作用于潜在表示的目标。
AI 中文摘要
机器遗忘旨在消除特定训练数据的影响,同时保留模型效用。许多先进方法通过将遗忘更新限制在基于梯度显著性选择的参数子集来实现这一目标。然而,基于显著性的权重选择对表示级遗忘的实际贡献尚不清楚。本文首次对SalUn使用的显著性掩码机制进行了可控消融实验。在CIFAR-10和CIFAR-100数据集上使用ResNet-18进行匹配计算实验设计,比较基于显著性的掩码与等稀疏度的随机掩码和无约束更新,同时保持遗忘目标、优化计划和计算预算不变。在包括线性探测、原型恢复和逐层CKA在内的多个表示级评估中,三种配置在表示级恢复能力上具有统计等效性。研究发现,在应用任何掩码之前,遗忘梯度强烈集中在最终网络层(在CIFAR-10上约92%的平方梯度能量),导致所有掩码策略在相同的表示子空间内运行。此外,显著性掩码的类别特异性有限(特异性指数为0.09-0.11),在不同的遗忘类别中选择高度重叠的参数子集。研究结果表明,在研究的设置中,表示级遗忘主要受梯度集中和表示几何结构的控制,而不是由显著性选择的权重的特定身份决定。更广泛地说,结果支持了越来越多的证据,表明有效的表示级遗忘需要直接作用于潜在表示的目标,而不是越来越复杂的权重选择策略。
英文摘要
Machine unlearning aims to remove the influence of specific training data while preserving model utility. Many state-of-the-art approaches pursue this goal by restricting the forgetting update to a subset of parameters selected through gradient-based saliency. Although such methods are widely adopted, the actual contribution of saliency-based weight selection to representation-level forgetting remains unclear. In this work, we perform the first controlled ablation of the saliency masking mechanism used by SalUn. Using a matched-compute experimental design on CIFAR-10 and CIFAR-100 with ResNet-18, we compare saliency-based masking against random masks of equal sparsity and unconstrained updates, while keeping the unlearning objective, optimization schedule, and computational budget fixed. Across multiple representation-level evaluations, including linear probing, prototype recovery, and layer-wise CKA, the three configurations exhibit statistically equivalent representation-level recoverability. We find that forget gradients are strongly concentrated in the final network layers (approximately 92% of the squared gradient energy on CIFAR-10) before any mask is applied, causing all masking strategies to operate within the same representational subspace. Furthermore, saliency masks show limited class specificity (specificity index 0.09-0.11), selecting highly overlapping parameter subsets across different forget classes. Our findings suggest that, in the studied setting, representation-level forgetting is primarily governed by gradient concentration and representation geometry rather than by the specific identity of saliency-selected weights. More broadly, the results support a growing body of evidence indicating that effective representation-level unlearning requires objectives that act directly on latent representations rather than on increasingly sophisticated weight-selection strategies.
Comments51 pages, 7 figures. Submitted to Neural Networks