DEFUSE:面向带生成先验的自监督编码器的可泛化后门防御方法
DEFUSE: Generalizable Backdoor Defense for Self-Supervised Encoders with Generative Priors
- Ant Group(蚂蚁集团)
- Purple Mountain Laboratories(紫金山实验室)
- City University of Hong Kong(香港城市大学)
- Beijing Institute of Technology(北京理工大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出DEFUSE,一种基于条件扩散生成模型的可泛化后门检测框架,用于防御自监督学习编码器的后门攻击,实验表明其性能优于现有方法且泛化性更强。
AI中文摘要:
自监督学习(SSL)编码器易受后门攻击,对视觉SSL编码器及视觉-语言编码器均构成威胁。现有防御方法通常仅针对其中一种范式设计,且依赖难以在实践中满足的限制性假设,例如可获取未受感染的分布内数据或预计算的伪标签。为解决这些局限,本文提出DEFUSE,一种面向SSL编码器的可泛化后门检测框架。受贝叶斯后验推理启发,本文将后门检测重新表述为以条件扩散生成模型为参数的、基于表示的图像似然估计问题。未受感染的表示倾向于生成语义一致的重构结果,而带后门的表示更可能被映射到攻击者的目标类别或语义无意义的图像,偏离原始语义,从而暴露后门。然而,精确似然计算难以实现,因为高度抽象的表示会丢失像素级忠实重构所需的低级信息。因此,本文将目标放宽为语义重构,并在参考编码器提供的良好分离的表示空间中进行评估。本文并非从零开始训练,而是微调预训练的扩散模型,利用其生成先验将数据映射到自然图像流形,同时保留语义内容。大量实验表明,DEFUSE在各类攻击设置下的性能显著优于现有检测器,可泛化到视觉SSL编码器和视觉-语言编码器。值得注意的是,本文方法大幅降低了对受害者编码器或攻击策略的先验知识依赖。源代码可在指定URL获取。
英文摘要:
Self-supervised learning (SSL) encoders are vulnerable to backdoor attacks, posing threats to both visual SSL encoders and vision-language encoders. Existing defenses are typically designed for only one of these paradigms and rely on restrictive assumptions such as access to uninfected in-distribution data or precomputed pseudo-labels, which are difficult to satisfy in practice. To address these limitations, we propose DEFUSE, a generalizable backdoor detection framework for SSL encoders. Inspired by Bayesian posterior inference, we reformulate backdoor detection as a representation-conditioned image likelihood estimation problem parameterized by a conditional diffusion generative model. Uninfected representations tend to yield semantically consistent reconstructions, whereas backdoored ones are more likely to be mapped to the attacker's target class or semantically meaningless images, deviating from the original semantics and thereby exposing the backdoor. However, we find that the exact likelihood is intractable, because highly abstracted representations discard the low-level information necessary for pixel-faithful reconstruction. We therefore relax the objective to semantic reconstruction and evaluate it in a well-separated representation space provided by a reference encoder. Rather than training from scratch, we fine-tune a pretrained diffusion model, leveraging its generative prior to map data onto the natural image manifold while preserving semantic content. Extensive experiments demonstrate that DEFUSE substantially outperforms existing detectors across diverse attack settings, generalizing to both visual SSL and vision-language encoders. Notably, our method greatly reduces the reliance on prior knowledge about the victim encoder or the attack strategy. The source code is available at https://github.com/jsrdcht/DEFUSE .