arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于持续机器遗忘的极小极大高斯机制

Minimax Gaussian Mechanisms for Continual Machine Unlearning

Qi Kuang, Yin Xia

arXiv 2610.11628首次发表:更新:

发表机构

Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对持续机器遗忘问题,开发了用于牛顿更新的高斯机制,通过极小极大优化校准噪声,在高斯差分隐私框架下实现了与精确重训练相近的模型更新效果,经模拟和信用违约数据验证有效。

AI 中文摘要

机器遗忘是指在记录被删除后更新已训练模型,目标是在不重复完整训练流程的情况下匹配精确重训练的效果。我们针对序列删除请求开发了用于牛顿更新的高斯机制。利用高斯差分隐私(GDP)及其自适应组合规则,我们证明了所发布的全部序列模型在统计上难以与匹配的精确重训练区分开。为了针对经验风险最小化校准这些机制,我们推导了牛顿近似相对于精确重训练的误差上界,以及该误差在每批删除后如何变化的上界。独立高斯噪声通过每次发布时的全残差上界进行校准,而高斯随机游走噪声则使用更小的残差增量上界。这些上界产生的分配方案可在所得GDP认证约束下最小化所有发布中的最坏情况最大噪声方差。基于计数的上界,随机游走在删除上限为M时渐近匹配单次发布的最坏情况方差,而独立噪声则会产生一个阶为M的额外因子。基于集合的上界可利用已删除记录的梯度和海森矩阵降低噪声方差。对于单条记录删除,我们进一步证明,基于计数的独立噪声、基于计数的随机游走噪声和基于集合的独立噪声在各自的残差或增量上界下,于固定高斯协方差中是极小极大的。在某些数据序列上,基于集合的上界允许方差适应已删除记录,相比固定协方差可提升阶为(log M)²的性能。残差和噪声上界还能在所有删除策略下,相对于精确重训练实现参数和预测的一致性。模拟实验和信用违约数据分析对这些上界、噪声方差和估计误差进行了评估。

英文摘要

Machine unlearning updates a trained model after records are deleted, aiming to match exact retraining without repeating the full training procedure. We develop Gaussian mechanisms for Newton updates under sequential deletion requests. Using Gaussian differential privacy (GDP) and its adaptive composition rule, we show that the full sequence of released models is statistically difficult to distinguish from matched exact retraining. To calibrate these mechanisms for empirical risk minimization, we derive upper bounds on the error of the Newton approximation relative to exact retraining and on how this error changes after each deletion batch. Independent Gaussian noise is calibrated using bounds on the full residual at each release, whereas Gaussian random walk noise uses smaller bounds on residual increments. These bounds yield allocations minimizing the worst-case maximum noise variance across releases under the resulting GDP certification constraints. With count-based bounds, the random walk asymptotically matches the worst-case variance of a single release at deletion cap $M$, while independent noise incurs an additional factor of order $M$. Set-based bounds can reduce the noise variances by using gradients and Hessians of the deleted records. For singleton deletion, we further show that count-based independent noise, count-based random walk noise, and set-based independent noise are minimax among fixed Gaussian covariances under their respective residual or increment bounds. With set-based bounds, allowing variances to adapt to deleted records can improve on every fixed covariance by a factor of order $(\log M)^2$ on some data sequences. The residual and noise bounds also yield parameter and predictive consistency relative to exact retraining, uniformly over deletion policies. Simulations and a credit default data analysis evaluate bounds, noise variances, and estimation errors.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑