发表机构
Changsha University of Science & Technology; Xiangtan University; Yunnan University; Huazhong University of Science and Technology; City University of Hong Kong; Griffith University(长沙理工大学; 湘潭大学; 云南大学; 华中科技大学; 香港城市大学; 格里菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对目标检测后门攻击,提出输入阶段黑盒防御方法ODPure,通过腐败-重建-选择范式净化输入,在保持精度的同时有效防御多种后门攻击。
AI 中文摘要
随着自动驾驶等应用的发展,目标检测已获得显著关注,同时也暴露出严重破坏模型完整性的关键漏洞,如后门攻击。具体而言,此类攻击涉及当输入中存在预定义触发器时,改变对象的类别(即对象误分类)、移除边界框(即对象消失),或为非存在对象生成边界框提议(即对象生成)。尽管图像分类的后门防御已相当成熟,但针对目标检测的研究仍相对不足。现有防御通过扫描输出或模型来应对这些威胁,但需要丢弃恶意数据或模型。这种补救措施无法为目标检测流程提供持续且准确的感知流。为解决这些局限性,我们提出ODPure,一种针对目标检测的新型输入阶段黑盒防御方法,其基于输入净化以确保稳定的感知流。针对目标检测器密集预测的特性,我们的腐败-重建-选择(CRS)范式通过利用多样化的腐败组合来中和触发器,生成大量冗余提议,然后通过生成先验恢复细粒度结构线索,最终采用投票对生成的检测结果达成共识。综合实验表明,我们的方法在保持基线准确性的同时,对多种后门攻击和触发器类型提供了稳健的防御。我们的代码可在https://github.com/Alex66366/ODPure获取。
英文摘要
With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulnerabilities like backdoor attacks that severely compromise model integrity. Specifically, such attacks involve altering the categories of objects (i.e., object misclassification), removing bounding boxes (i.e., object disappearance), or generating bounding box proposals for non-existent objects (i.e., object generation) when a predefined trigger is present in the input. Although backdoor defenses for image classification are well-established, the research for object detection remains comparatively underexplored. Existing defenses address these threats by scanning outputs or models for potential backdoors but require discarding either malicious data or models. This remedy fails to enable a continuous and accurate perceptual stream for the object detection pipeline. To address such limitations, we propose ODPure, a novel input-stage black-box defense for object detection, which is based on input purification that ensures stable perception flows. Tailored to the dense prediction nature of object detectors, our Corruption-Reconstruction-Selection (CRS) paradigm operates by neutralizing triggers through a diverse portfolio of corruptions to generate a massive pool of redundant proposals, then recovering fine-grained structural cues via generative priors, and finally employing voting to reach a consensus on the resulting detections. Comprehensive experiments demonstrate that our method provides robust defense against diverse backdoor attacks and trigger types while preserving baseline accuracy. Our code is available at https://github.com/Alex66366/ODPure.
Comments13 pages, 8 figures (including supplementary materials); Code available at https://github.com/Alex66366/ODPure