arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后门作为探针:CLIP的测试时对抗防御

Backdoor as Probe: Test-Time Adversarial Defense for CLIP

Zhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen, Shiguang Shan, Xilin Chen

arXiv 2609.34641首次发表:更新:

发表机构

Xuzhou University of Technology; University of Chinese Academy of Sciences(徐州工程学院; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出后门作为探针(BaP)的测试时对抗防御方法,利用后门机制将对抗性激活偏移转化为检测信号,通过闭式模型编辑构建探针,在16个基准上将鲁棒准确率从1.0%提升至52.3%,并保持干净准确率,推理加速5.7倍。

AI 中文摘要

测试时对抗防御在不重新训练的情况下提高了视觉-语言基础模型(如CLIP)的鲁棒性。然而,对抗性激活偏移通常被视为需要抑制的失真,而非可利用的信号。我们通过重新利用后门的触发器到目标机制,将这些偏移转化为防御信号。关键是将一个防御者控制的后门作为探针植入,该探针在干净输入下弱激活,但在对抗性偏移下强激活。基于这一见解,我们提出了“后门作为探针”(BaP),一种针对CLIP的测试时对抗防御。BaP通过对选定的MLP层进行闭式模型编辑来构建探针。它将平均对抗性激活偏移和防御者指定的语义方向分别投影到该层的低能量输入和输出激活子空间,以获得触发器和目标方向。在推理时,对抗性输入沿目标方向产生可测量的响应,用于检测。BaP随后通过优化一个小的扰动来选择性纠正检测到的输入,将其表示从对抗性偏移引导回干净子空间。在16个基准上的实验表明,BaP将平均鲁棒准确率从1.0%提高到52.3%,同时保持干净准确率,实现了与最先进方法相当的性能,并具有高达5.7倍的推理加速。BaP进一步展示了其对大型视觉-语言模型上对抗性攻击的泛化能力。项目页面:此https URL

英文摘要

Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose \emph{Backdoor as Probe} (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models. Project page: https://robin-wzq.github.io/Backdoor-as-Probe/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑