批归一化幻觉:机器遗忘评估中归一化伪影的诊断
The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation
浏览论文内容
中文总结 AI 辅助
本文诊断了机器遗忘评估中批归一化引起的伪影,提出定点算子框架区分测量与编码器失败,并证明该伪影可显著逆转遗忘准确率,而组归一化可消除之。
中文摘要 AI 辅助
近似机器遗忘旨在无需从头重新训练的情况下,消除特定训练数据对已训练模型的影响。我们发现,在基于批归一化(BatchNorm)的架构上评估遗忘效果时,存在一个此前未被记录的混淆因素:对保留数据仅进行一次前向传播(该操作不修改任何权重)即可确定性地重写模型的归一化状态,并逆转表面指标所显示的遗忘效果。我们将此操作形式化为一个保持权重的定点算子,并证明其引起的任何前后差异均可明确归因于批归一化的运行统计量,而非遗忘方法对权重所做的任何修改。这一归因论断清晰地区分了测量失败(批归一化伪影)与编码器失败(残留在权重中的信息,近期并行工作已有所记录),且同一算子框架可对线性探测提升进行唯一分解,分为批归一化测量偏差和编码器几何成分。实验上,在标准基准上对九种评估方法,该伪影使头条遗忘准确率逆转高达78个百分点;攻击者仅需10张无标签图像即可恢复大部分被掩盖的准确率;而严格的组归一化(GroupNorm)对照在所有方法上将伪影降至零。所测试的成员推断攻击在校准后变化很小,表明所观察到的评估失败主要存在于遗忘准确率和线性探测中。
英文摘要
Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting. We formalize this operation as a weight-preserving fixed-point operator and prove that any pre-versus-post gap it induces is provably attributable to BN running statistics rather than to any modification the unlearning method made to the weights. This attribution claim cleanly separates measurement failure (BN artifact) from encoder failure (residual weight-encoded information, recently documented in concurrent work), and the same operator framework yields a unique decomposition of linear-probe elevation into BN-measurement-bias and encoder-geometry components. Empirically, the artifact reverses headline forget accuracy by up to 78 pp across nine evaluated methods on standard benchmarks; an attacker with as few as 10 unlabeled images recovers most of the masked accuracy; and a strict GroupNorm control reduces the artifact to zero across all methods. The tested membership-inference attacks change little under recalibration, locating the observed evaluation failure in forget accuracy and linear probing.
发表机构
- BITS Pilani(BITS Pilani(比拉理工学院 Pilani校区))
- KIIT(KIIT大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。