AI 中文总结
研究模型遗忘中认证遗忘的优化复杂性,利用一致凸正则化器证明相关距离新界,开发新二阶遗忘算法,在实现线性模型认证遗忘上有快速收敛率,涵盖逻辑与指数回归遗忘,显示二阶信息优势。
AI 中文摘要
我们研究模型遗忘,即从训练好的模型中移除记忆的训练数据。具体而言,我们从优化角度研究认证遗忘的算法复杂性。我们将遗忘算法的目标形式化为同时实现认证遗忘和优化精度。利用一致凸正则化器的概念,我们用泛化误差的新替代证明了初始模型与遗忘后模型之间距离的新界。理论上表明,如果遗忘模型能很好地预测移除的数据,相应的优化问题就简单。此外,我们开发了一种具有各向异性高斯机制和最先进全局收敛性的新二阶遗忘算法。对于具有拟自协调损失的线性模型,我们证明了该方法在实现认证遗忘方面的快速收敛率。作为直接应用,我们的理论涵盖逻辑回归和指数回归的遗忘,并显示了与一阶遗忘方法相比利用二阶信息的可证明优势。
英文摘要
We study machine unlearning: the removal of memorized training data from a trained model. Specifically, we investigate the algorithmic complexity of certified unlearning from an optimization perspective. We formalize the goal of an unlearning algorithm as simultaneously achieving certified unlearning and optimization accuracy. Utilizing the notion of uniformly convex regularizers, we prove new bounds on the distance between initial and unlearned models using a novel substitute for generalization error. Thus we theoretically demonstrate that if the removed data is well-predicted by the unlearned model, the corresponding optimization problem is simple. Furthermore, we develop a new second-order unlearning algorithm with an anisotropic Gaussian mechanism and state-of-the-art global convergence. We prove fast rates for our method in achieving certified unlearning for linear models with quasi-self-concordant losses. As a direct application, our theory covers unlearning for logistic and exponential regressions and shows a provable benefit of utilizing second-order information compared to first-order unlearning methods.