发表机构
ÉTS Montreal; Mila - Quebec AI Institute; CNRS, CentraleSupélec - Université Paris-Saclay; LIVIA; ILLS(蒙特利尔工程学院; 米拉-魁北克人工智能研究所; 法国国家科学研究中心、中央高等电力学院-巴黎萨克雷大学; LIVIA实验室; ILLS机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出无源重学习审计(SFRA)方法,在无源设置下诊断类遗忘中的遗忘问题,引入重学习得分(RS)量化遗忘类可恢复性,发现多种遗忘方法存在显著无源可恢复性。
AI 中文摘要
类遗忘旨在消除模型识别指定遗忘类别的能力,同时保留对保留类的性能。然而,遗忘后的低遗忘准确率并不一定意味着类结构已被擦除。近似遗忘方法可改变分类器决策边界,但会在表示中留下可恢复的结构。现有研究表明遗忘类可被恢复,但现有方法需要真实的遗忘或保留样本、辅助数据或参考检查点。我们在严格的无源设置中研究类重学习,仅使用遗忘后的模型,通过分类器头部更新判断遗忘类是否可被恢复。我们的方法基于理论分析,建立了充分对齐条件,在该条件下,对合成探测集进行单步梯度下降可增大遗忘类的预期 logit 间隔。在此基础上,我们提出白盒无源重学习审计(SFRA),该方法在表示空间中生成候选嵌入,使用模型引导的置信度过滤构建高置信度的保留类探测样本,以及低置信度的边界相邻探测样本,并将其重新标记为遗忘类。默认使用高斯采样和 Softmax 置信度,对替代提议分布和不确定性准则的消融实验表明,可恢复性并非特定于这些选择。为量化可恢复性,我们引入重学习得分(RS),该得分联合衡量遗忘类恢复和保留准确率的保持情况,并报告相对于重新训练参考的类匹配 ΔRS。在 CIFAR-10、CIFAR-100 和 TinyImageNet 数据集上,使用 ResNet-18、ViT-B/16 和 Swin-T 模型的实验表明,多种遗忘方法表现出显著的无源可恢复性,且对于部分方法,该可恢复性超过匹配的重新训练参考。
英文摘要
Class unlearning aims to remove a model's ability to recognize designated forget classes while preserving performance on retain classes. However, low forget accuracy after unlearning does not necessarily mean the class structure has been erased. Approximate unlearning methods can alter classifier decision boundaries while leaving recoverable structure in the representation. Prior work has shown that forget classes can be recovered, but existing approaches require real forget or retain samples, auxiliary data, or reference checkpoints. We study class relearning in a strictly source-free setting, asking whether a forget class can be recovered through a classifier-head update using only the unlearned model. Our approach rests on a theoretical analysis establishing a sufficient alignment condition under which a single gradient step on a synthetic probe set increases the expected logit margin of the forget class. Building on this, we propose a white-box Source-Free Relearning Audit (SFRA), which generates candidate embeddings in representation space and uses model-guided confidence filtering to construct high-confidence retain probes and low-confidence boundary-adjacent probes that are relabelled as the forget class. Gaussian sampling and Softmax confidence are used by default, while ablations with alternative proposal distributions and uncertainty criteria show that recoverability is not specific to these choices. To quantify recoverability, we introduce the Relearning Score (RS), which jointly measures forget-class recovery and retain-accuracy preservation, and report class-matched $Δ$RS relative to a retrained reference. Experiments on CIFAR-10, CIFAR-100, and TinyImageNet with ResNet-18, ViT-B/16, and Swin-T show that several unlearning methods exhibit substantial source-free recoverability, and that for a subset of methods this recoverability exceeds the matched retrained reference.