arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考后门修复评估:区分总体干净效用与良性性能保持

Rethinking Backdoor Repair Evaluation: Distinguishing Aggregate Clean Utility from Benign Performance Preservation

Baogang Song, Changtian Song, Jian Chen, Fan He, Junwei Zhou, Jianwen Xiang, Dongdong Zhao

arXiv 2609.25579首次发表:更新:

AI 中文总结

本文从保持视角重新审视后门修复评估,指出总体干净准确率会掩盖局部类别性能退化,提出最差类别和尾部类别保持损失以补充评估,实证表明攻击抑制与总体性能良好并不保证各类别性能均匀保持。

AI 中文摘要

后门修复旨在抑制被攻陷模型中的恶意行为,同时保持良性任务性能。现有研究通常使用攻击成功率(ASR)和总体干净准确率来评估这些目标,但总体干净准确率可能掩盖标签空间中一小部分集中出现的显著性能退化。我们从保持的视角重新审视良性性能评估,将总体干净效用与先前可用的逐类别性能的保持区分开来。我们通过比较修复前后的干净性能来定义逐类别保持损失,并表明聚合可能通过局部损失稀释和跨类别补偿来隐藏局部退化。为补充总体干净准确率,我们使用最差类别保持损失和尾部类别保持损失来刻画局部保持损失。我们针对具有代表性的后门攻击、修复方法、数据集、攻击目标和模型架构进行了系统性实证研究,并在干净标签攻击下进行了额外验证。结果表明,有效的攻击抑制和良好的总体干净性能并不必然意味着先前可用的良性性能在各类别间得到均匀保持。显著的局部保持损失可能仍然存在,其严重程度和逐类别结构随修复条件而变化。这些发现促使在ASR和总体干净准确率之外,采用面向保持的逐类别评估。

英文摘要

Backdoor repair aims to suppress malicious behavior in compromised models while preserving benign task performance. Existing studies typically evaluate these objectives using Attack Success Rate (ASR) and Overall Clean Accuracy, but aggregate clean accuracy can obscure substantial degradation concentrated in a small portion of the label space. We revisit benign-performance evaluation from a preservation perspective by distinguishing aggregate clean utility from the preservation of previously available class-wise performance. We define class-wise preservation loss by comparing clean performance before and after repair and show that aggregation can hide localized degradation through localized-loss dilution and cross-class compensation. To complement Overall Clean Accuracy, we characterize localized preservation loss using Worst-Class Preservation Loss and Tail Preservation Loss. We conduct a systematic empirical study across representative backdoor attacks, repair methods, datasets, attack targets, and model architectures, with additional validation under clean-label attacks. Results show that effective attack suppression and favorable aggregate clean performance do not necessarily imply uniform preservation of previously available benign performance across classes. Substantial localized preservation losses can remain, and their severity and class-wise structure vary across repair conditions. These findings motivate preservation-oriented class-wise evaluation alongside ASR and Overall Clean Accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑