AI 中文总结
该研究提出基于双生成器学习的干净标签遗忘激活后门框架,通过建立持久潜在关联与可移除抑制,在CIFAR-10和ImageNet-10上实现低遗忘前攻击成功率与强遗忘后激活,提升后门攻击的隐蔽性。
AI 中文摘要
现有后门攻击通常在植入后门后立即生效,可能在利用前被暴露。激活的潜伏后门(通过机器遗忘缓解此类行为暴露)在训练后保持非活跃状态,仅在选定训练记录被遗忘后才生效。然而,现有方法难以在干净标签约束和现实遗忘需求下同时实现低遗忘前攻击成功率和强遗忘后激活。实现该转变需同时建立持久的潜在关联和可移除的抑制影响。为应对此挑战,我们提出一种基于双生成器学习的干净标签遗忘激活后门框架,并将其公式化为双层优化问题:通过模拟潜在后门建立和机器遗忘,该框架交替学习建立潜在触发器-目标关联的样本特定触发器,以及提供可移除抑制的标签一致伪装样本。一旦小部分伪装样本被遗忘,抑制解除,潜伏后门激活。在CIFAR-10和ImageNet-10上的实验表明,与代表性后门基线相比,我们的方法在多种遗忘算法下,遗忘前保持更低的攻击成功率,同时实现更强的遗忘后激活。这些结果表明,在干净标签和现实删除约束下,通过协调持久潜在关联与可移除抑制,可实现可靠的潜伏转激活转变。
英文摘要
Existing backdoor attacks often become effective immediately after backdoor implantation and may therefore be exposed before exploitation. Machine unlearning activated dormant backdoors mitigate such behavioral exposure by remaining inactive after training and becoming effective only after selected training records are unlearned. However, existing methods struggle to simultaneously achieve a low pre-unlearning attack success rate and strong post-unlearning activation under clean-label constraints and realistic unlearning requests. Achieving this transition requires jointly establishing a persistent latent association and a removable suppressive influence. To address this challenge, we propose a clean-label unlearning-activated backdoor framework based on dual-generator learning and formulate it as a bilevel optimization problem: By simulating latent backdoor establishment and machine unlearning, the framework alternately learns sample-specific triggers that establish a latent trigger-to-target association and label-consistent camouflage samples that provide removable suppression. Once a small subset of camouflage samples is unlearned, the suppression is lifted and the dormant backdoor is activated. Experiments on CIFAR-10 and ImageNet-10 show that our method maintains lower pre-unlearning attack success rates while achieving stronger post-unlearning activation across multiple unlearning algorithms than representative backdoor baselines. These results demonstrate that reliable dormancy-to-activation transitions can be achieved by coordinating a persistent latent association with removable suppression under clean-label and realistic deletion constraints.
Comments12 pages, 7 figures, 4 tables;