发表机构
School of Mathematics; University of Edinburgh; Beijing Jinhe Technology Co., Ltd.; Inno Asset Management(数学学院; 爱丁堡大学; 北京金和科技有限公司; Inno资产管理公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对后门防御成本评估偏差,提出TARE方法,通过在从未投毒的孪生模型上运行防御来分离投毒移除成本,揭示并修正了基准测试中的负成本与零成本误判。
AI 中文摘要
后门防御排行榜会打印出干净准确率的下降,并将其解读为移除成本。仅在受投毒的受害者模型上测量时,这种下降无法将移除效果与防御对任何模型所造成的影响区分开来,并且继承了受害者的初始状态。在BackdoorBench的16种攻击中,有3种攻击的受害者初始状态来自配置文件:WaNet、BPP和Input-Aware附带了一个永远不会触发的MultiStepLR学习率调度器,因此它们的受害者模型永远不会进行退火(学习率调整),并且在30/31个公开的CIFAR单元格中,其准确率最低,不超过5%。在PreAct-ResNet18上,微调系列防御方法将低初始状态恢复到自身水平,因此在那里公布的移除成本为负值,基准测试的评级将这种“增益”截断为零,并且在我们阅读的48篇引用防御论文中,有2篇基于这些单元格提出了零成本的声明;TSBD和CGD,使用其代码重新运行,在从未投毒的模型上也表现出“增益”。仅对该调度器行进行2×2的编辑就能隔离出原因,其交换的臂在运行前就自行注册:微调系列在干净模型上的成本符号双向反转,而其在退火受害者上公布的增益仅缩小至接近零,44/44个种子遵循该调度,并在BPP、FT-SAM、CIFAR-100和VGG19-BN上复现,且在第二个工具包中诱导产生。TARE在相同配方、调度和种子的从未投毒孪生模型上运行相同的防御(在BackdoorBench上,投毒图像不超过10张,仅在攻击成功率低于5%时允许);孪生模型所损失的即为“皮重”。在BadNets网格上,八个防御中有七个对孪生模型收费(Neural Cleanse仅在其检测器触发时),范围从+0.13(微调)到+5.70个百分点(I-BAU);第八个ABL则摧毁了它。在攻击内部,初始状态在排名中相互抵消,因此皮重在那里不重新排序;投毒在初始状态之外所增加的内容在两种估计器下打印,且未进行校正,其移除份额未被识别。我们发布了三键补丁、一个带符号的皮重列(7种攻击×8种防御)以及TARE-Z,一个用于种子稳定防御的无孪生估计器。
英文摘要
Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR that never fires, so their victims never anneal and are the least accurate in 30/31 public CIFAR cells at $\leq$5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to their own level, so there the published cost is negative, the benchmark's rating clips the "gain" to zero, and 2 of 48 citing defense papers we read rest a no-cost claim on those cells; TSBD and CGD, re-run with their code, "gain" on a never-poisoned model too. A $2\times2$ editing only that scheduler line isolates the cause, its swapped arms self-registered before they ran: the sign of the fine-tuning family's clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. TARE runs the same defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, $\leq$10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On the BadNets grid seven of eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 (fine-tuning) to +5.70 points (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share unidentified. We ship the three-key patch, a signed tare column (7 attacks $\times$ 8 defenses) and TARE-Z, a twin-free estimator for seed-stable defenses.
Comments84 pages (9-page main text)