发表机构
Washington State University(华盛顿州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CLIP嵌入空间后门攻击,提出轻量级全黑盒防御CLIPGuard,通过分段扰动检测与语义修复净化恶意区域,实验将攻击成功率降至1.05%,保持准确率86.34%。
AI 中文摘要
对比语言-图像预训练(CLIP)已成为一种主流的视觉骨干模型,因其强大的迁移性和零样本能力。然而,近期研究揭示了一个关键漏洞:嵌入空间后门攻击。通过仅污染极小一部分图像-文本对,攻击者可以植入隐蔽触发器,诱导CLIP联合嵌入空间发生定向偏移。与操纵分类器logits的传统后门不同,这些攻击直接破坏表示,使其在极低污染率下高度有效且难以检测。现有防御方法需要访问模型参数、梯度、logits或干净验证数据——这些假设在现实黑盒部署中很少成立。此外,当前黑盒方法难以准确定位小规模或分布外触发器。我们提出CLIPGuard,一种轻量级且完全黑盒的防御方法,专门用于缓解CLIP编码器中的嵌入空间后门。CLIPGuard通过测量分段级嵌入扰动来识别恶意区域,并通过语义修复选择性地仅净化可疑分段,从而保留良性视觉内容和对齐质量。在STL-10、ImageNet及多种触发器家族(包括BadCLIP、BadNets、混合、基于补丁和排版攻击)上的大量实验表明,CLIPGuard将攻击成功率降至低至1.05%,同时保持干净准确率高达86.34%,持续优于现有黑盒防御方法,包括CleanCLIP和CleanerCLIP。我们的代码可在该https URL获取。
英文摘要
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git