遗忘还是微调?针对噪声标签修正的机器遗忘策略对比研究
Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction
- Universidade Federal Rural de Pernambuco(伯南布哥联邦农村大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究对比五种机器遗忘策略在不同噪声类型数据集上的效果,发现策略适配噪声结构,为噪声标签修正的机器遗忘策略选择提供实用指南。
AI中文摘要:
噪声标签仍是训练深度神经网络的关键挑战,因为记忆错误标签会降低泛化能力。训练后识别出噪声样本后,标准解决方案是在清洗后的数据集上从头重新训练模型,但随着数据集和模型规模增长,该方案的成本越来越高。机器遗忘(MU)作为一种计算高效的替代方案最近出现,但不同MU策略用于噪声标签修正的相对有效性仍鲜为人知。本研究在CIFAR-10、CIFAR-100和真实世界噪声数据集Food-101N上,针对对称、非对称、实例依赖和开放集噪声,对五种MU方法(NegGrad、微调(FT)、随机标记(RL)、SalUn和MUNBa)进行了实证对比研究。核心发现是,合适的遗忘策略取决于噪声结构:简单的FT在大多数闭集场景中是强大的基准方法;RL和SalUn是最一致的鲁棒方法,在实例依赖噪声下,其准确率以仅为重新训练的一小部分计算成本接近重新训练的准确率;MUNBa主要在极端对称噪声下展现优势。相反,在开放集噪声下,我们发现对清洗后的子集重新训练会相对于噪声基线降低准确率,因此在该场景下近似重新训练后的模型并非合适目标。在Food-101N上,所有MU方法仍具有竞争力,且准确率接近重新训练,同时运行时间减少了一个数量级。这些发现为选择用于训练后噪声标签修正的MU策略提供了实用指南。
英文摘要:
Noisy labels remain a critical challenge for training deep neural networks, since memorizing incorrect labels degrades generalization. Once noisy samples are identified after training, the standard solution is to retrain the model from scratch on the cleaned dataset, which is increasingly expensive as datasets and models grow. Machine Unlearning (MU) has recently emerged as a computationally efficient alternative, but the relative effectiveness of different MU strategies for noisy-label correction remains poorly understood. In this work, we conduct a comparative empirical study of five MU methods (NegGrad, Fine-Tuning (FT), Random Labeling (RL), SalUn, and MUNBa) across symmetric, asymmetric, instance-dependent, and open-set noise on CIFAR-10, CIFAR-100, and the real-world noisy dataset Food-101N. Our central finding is that the appropriate unlearning strategy is conditioned on the noise structure. Simple FT is a strong baseline across most closed-set scenarios; RL and SalUn are the most consistently robust methods and, under instance-dependent noise, approach retraining accuracy at a fraction of the computational cost; MUNBa shows advantages mainly under extreme symmetric noise. Under open-set noise, in contrast, we show that retraining on the cleaned subset degrades accuracy relative to the noisy baseline, so approximating the retrained model is not an adequate objective in this regime. On Food-101N, all MU methods remain competitive and achieve accuracies close to retraining despite reducing runtime by an order of magnitude. These findings provide practical guidelines for selecting MU strategies for post-training noisy-label correction.