超越小补丁:黑盒环境下多样后门触发器的检测与净化
Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers
- School of Engineering and Applied Sciences, Washington State University(华盛顿州立大学工程与应用科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出面向部署的黑盒后门防御方法TRIM,通过区域分割、自适应触发器发现与选择性净化,在推理时移除后门触发器,性能优于现有方法,可低至1.16%的攻击成功率并保持高干净准确率。
AI中文摘要:
深度神经网络(DNNs)越来越多地部署在现实视觉系统中,但其预测可能被后门攻击秘密操纵,后门攻击中恶意触发器会引发目标误分类,同时保持较高的干净准确率。现有防御方法常依赖模型内部信息、训练数据或干净验证样本,仅能在可获取训练后模型黑盒访问权限时难以部署。我们提出TRIM(Trigger Removal by Identifying Manipulated Regions,即通过识别受操纵区域移除触发器),这是一种面向部署的黑盒防御方法,可在推理时检测并选择性移除后门触发器,无需模型内部信息、训练数据或干净样本。TRIM的核心思路是识别导致模型异常行为的图像区域,仅净化这些区域同时保留良性内容。TRIM通过三个关键组件实现创新:(i)基于深度特征表示的区域分割;(ii)通过修复和基于扩散的重建进行自适应触发器发现,以隔离导致误分类的区域——无需对触发器类型、形状或位置做假设;(iii)选择性区域净化,在保留良性内容的同时清理中毒区域。为支持实际部署,TRIM还会缓存先前识别触发器的特征嵌入,实现高效识别并避免冗余的检测与净化。在多样数据集和后门类型(包括混合、稀疏、不同尺寸及多个触发器)上的大量实验表明,TRIM的性能始终优于现有黑盒防御,可将攻击成功率(ASR)降至低至1.16%,同时保持高达87.87%的干净准确率。这些结果证明,即便防御者无法获取任何辅助数据,在推理时实现有效的后门缓解也是可行的。
英文摘要:
Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.