arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习遗忘:通过学习遗忘行为实现机器遗忘

Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors

Hang Zhang, Kaifeng Zhang, Yixiao Ma, Weijie Xu, Ye Zhu, Kai Ming Ting

arXiv 2608.16700首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology Nanjing University; Centre for Cyber Resilience and Trust Deakin University(南京大学计算机软件新技术国家重点实验室; 迪肯大学网络韧性与信任中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出模型不可知的学习遗忘方法L2UL,通过从分布视角学习遗忘行为替代手动设计遗忘函数,在数据密集场景及ResNet等大模型上,实现与重新训练相当的准确率且效率更高。

AI 中文摘要

针对隐私法规要求,已开发出多种机器遗忘技术,使个人能够行使合法权利,将其数据D_f从机器学习模型中移除。该过程通常通过使用称为U的遗忘函数来完成。现有方法侧重于设计复杂的U,以从先前的模型A(D)中遗忘子集D_f⊂D,使遗忘后的模型表现尽可能接近重新训练的模型A(D\backslash D_f)。然而,这些方法在处理海量训练数据时往往面临高计算成本,因为即使对于参数较少的模型,U的复杂结构也会成为瓶颈。受学习优化(Learning to Optimize)启发,我们提出了首个基于学习的模型不可知方法——学习遗忘(Learning-to-UnLearn,L2UL)。我们的核心见解是从手动设计U转变为从分布视角学习遗忘行为,从而通过学习获得简单高效的U。实验结果表明,L2UL的准确率与重新训练的结果相当,同时展现出令人印象深刻的效率,尤其在数据密集型场景中表现突出。此外,我们在更大的模型ResNet上验证了该方法的性能和可扩展性。

英文摘要

Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise their legal right to have their data $D_f$ removed from a machine learning model. This process is typically accomplished via the use of an unlearning function denoted as $U$. Existing methods focus on designing an intricate $U$ to unlearn $D_f \subset D$ from a previous model $A(D)$, so that the unlearned model performs as closely as possible to the retrained model $A(D \setminus D_f)$. However, these methods often suffer from high computational costs when dealing with massive training data, as the complex structures of $U$ become a bottleneck even for models with fewer parameters. Inspired by Learning to Optimize, we introduce the first learning-based model-agnostic approach, Learning-to-UnLearn (L2UL). Our core insight is to shift from manually designing $U$ to learning the unlearning behaviors from a distribution perspective, thereby acquiring a simple and efficient $U$ via learning. Our experimental results demonstrate that the accuracy achieved by L2UL is comparable to that of retraining while exhibiting impressive efficiency, particularly in data-intensive scenarios. Furthermore, we validate the performance and scalability of our method on larger models ResNet.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑