发表机构
Mila; ILLS; École normale supérieure Paris-Saclay; IRT Saint Exupéry(米拉研究所; ILLS; 巴黎-萨克雷高等师范学校; 圣埃克苏佩里研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出对比稀疏自编码器SCALPEL,通过选择性表示层干预实现细粒度机器遗忘,在TOFU基准上优于现有方法,兼顾目标移除与背景保持。
AI 中文摘要
机器遗忘旨在移除目标信息,同时保留模型的其他能力。在现实场景中,例如欧盟GDPR下的隐私请求,目标可能很窄,例如与单个人相关的信息。仅行为层面的遗忘可能不足,这促使直接对内部表示进行干预。然而,标准的机制可解释性提取器对此类目标的选择性较差。我们识别出基于重建的提取中存在能量偏差,该偏差偏向于主导背景结构,而非低能量的目标特定成分。我们引入SCALPEL,一种对比稀疏自编码器,旨在学习更具选择性的遗忘特征。我们从理论上证明,对比训练促进目标选择性特征,且我们的选择分数控制预期的背景知识扰动。我们在TOFU上对Qwen、Llama和Gemma进行了实验验证,SCALPEL显著优于NMF和标准SAE干预,并与Gradient Difference和RMU相当,从而弥合了机制可解释性与细粒度遗忘之间的差距。
英文摘要
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.