通过诱饵捷径与知识解耦缓解后门攻击
Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling
浏览论文内容
中文总结 AI 辅助
本研究提出TR防御方法,通过蜜罐捷径捕获后门知识,结合知识解耦策略与自动捷径生成,在四个基准数据集和五个模型架构上有效缓解多种后门攻击,同时保留良性性能。
中文摘要 AI 辅助
后门攻击对深度神经网络构成严重威胁,尤其当训练依赖第三方数据时,攻击者可通过数据投毒注入恶意行为。本研究发现,后门行为在与主网络联合训练时,倾向被更简单的并行分支吸收。基于该洞察,我们提出训练时防御方法TR(Trapping and Removing),引入轻量捷径分支作为“蜜罐”捕获后门知识,训练后丢弃该捷径即可移除后门,无需额外数据。为在保持良性性能的同时增强后门隔离,我们设计基于熵权重分配的知识解耦策略,引导投毒样本流经蜜罐,使主网络专注良性学习。此外,我们引入自动捷径生成策略提升跨模型架构的泛化性。在四个基准数据集和五个模型架构上的大量实验表明,该方法可有效缓解多种后门攻击,同时保留良性数据的性能。代码:this https URL 与 this http URL。
英文摘要
Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third-party data, allowing adversaries to inject malicious behaviors through data poisoning. In this work, we reveal that backdoor behaviors tend to be absorbed by a simpler parallel branch when jointly trained with the main network. Motivated by this insight, we propose Trapping and Removing (TR), a simple yet effective training-time defense that introduces a lightweight shortcut branch as a "honeypot" to trap backdoor knowledge. After training, backdoors can be removed by discarding the shortcut, without requiring any additional data. To further enhance backdoor isolation while maintaining benign performance, we design a knowledge decoupling strategy with entropy-based weight assignment, encouraging poisoned samples to flow through the honeypot while guiding the main network to focus on benign learning. In addition, we introduce an automatic shortcut generation strategy to improve generalization across model architectures. Extensive experiments on four benchmark datasets and five model architectures demonstrate that our approach effectively mitigates a wide range of backdoor attacks while preserving performance on benign data. Code: https://github.com/Zixuan-Zhu/TR}{github.com/Zixuan-Zhu/TR.
发表机构
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。