arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MUtE:概念擦除与反事实干预的双重框架

MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions

Antoine Saillenfest

arXiv 2609.11253首次发表:更新:

发表机构

onepoint(onepoint)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MUtE双重框架,通过推导新型擦除函数实现概念擦除与反事实生成的无缝切换,并施加平移偏差约束,实验证明其能改善算法公平性并生成反事实文本。

AI 中文摘要

从表示中擦除概念特定信息已被证明有助于减轻偏见或解释模型决策。其联合目标是转换原始表示,使得目标概念变得不可预测,同时最大限度地保留与概念无关的信息。在这项工作中,我们重新审视概念擦除的最优界限,以推导出一类新颖的擦除函数,该函数自然诱导出确定性的双重反事实映射。弥合理论最优性与实际表示学习之间的差距,我们设计了一种实现,该实现对反事实轨迹施加平移偏差——这一约束与许多概念在现代语言模型中的几何表现方式相一致。我们的框架能够在概念擦除与反事实生成之间无缝切换。我们通过实验证明了其在改善下游算法公平性和生成反事实文本方面的有效性。

英文摘要

Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models. Our framework enables seamless navigation between concept erasure and counterfactual generation. We empirically demonstrate its efficacy in improving downstream algorithmic fairness and generating counterfactual texts.

Comments21 pages, 3 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑