当模态无法“共舞”:多模态对比学习中的共形后门检测
When Modalities Fail to Tango: Conformal Backdoor Detection in Multimodal Contrastive Learning
浏览论文内容
中文总结 AI 辅助
针对多模态对比学习后门检测方法的缺陷,提出CASCADE两阶段共形框架,在CC3M数据集上实现高TPR下低FPR与高AUROC,可应对自适应攻击。
中文摘要 AI 辅助
近年来,多模态对比学习(Multimodal Contrastive Learning, MCL)中的后门攻击受到越来越多的关注,因为许多下游任务严重依赖预训练的MCL模型。现有的基于检测的防御方法主要依赖CLIPScore指标,其假设是被投毒的图像-文本对的语义相似度更低。然而,我们发现现有方法存在两个关键缺陷:(1)良性对和被投毒对的CLIPScore分布存在大量重叠,削弱了该指标的可靠性;(2)固定阈值检测无法为重叠区域内的模糊样本提供统计保证。为克服这些局限,我们提出整合共形预测(Conformal Prediction, CP)这一通过非一致性分数(Nonconformity Scores, NCSs)量化不确定性的统计框架,以建立可证明的置信区间来检测被投毒的图像-文本对。基于CP,我们提出了一种新颖的两阶段“由粗到细”共形后门检测框架CASCADE。粗粒度阶段利用跨模态一致性识别高置信度的良性对和被投毒对;细粒度阶段从高置信度的被投毒对构建参考集,并为未识别子集中的每个样本计算基于文本空间相似度的实例级NCSs,这些NCSs用于衡量对投毒分布的一致性,从而能够精确识别未识别子集中潜在的被投毒对。在大规模CC3M数据集上进行的大量实验表明,CASCADE在100%真阳性率(TPR)下的平均假阳性率(FPR)为5.79%,在各类攻击下的平均受试者工作特征曲线下面积(AUROC)为0.9867,同时对自适应攻击仍有效。
英文摘要
Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend on pre-trained MCL models. Existing detection-based defenses predominantly rely on the CLIPScore metric, under the assumption that poisoned pairs exhibit lower semantic similarity between the image and the caption. However, we identify two critical flaws remaining in existing methods: (1) the substantial overlap between CLIPScore distributions of benign and poisoned pairs undermines the reliability of this metric, and (2) fixed-threshold detection cannot provide statistical guarantees for ambiguous samples within overlapping regions. To overcome these limitations, we propose integrating conformal prediction (CP), a statistical framework that quantifies uncertainty through nonconformity scores (NCSs), to establish provable confidence bounds for detecting poisoned image-caption pairs. Building on CP, we introduce CASCADE, a novel two-stage Coarse-to-Fine Conformal Backdoor Detection framework. The coarse-grained stage uses cross-modality consistency to identify high-confidence benign and poisoned pairs. In the fine-grained stage, a reference set is constructed from high-confidence poisoned pairs, and instance-level NCSs based on text-space similarity are computed for each sample in the unidentified subset. These NCSs measure conformity to the poisoning distribution and enable precise identification of latent poisoned pairs within the unidentified subset. Extensive experiments on the large-scale CC3M dataset demonstrate that CASCADE achieves an average FPR of 5.79% at 100% TPR and an average AUROC of 0.9867 across diverse attacks, while remaining effective against adaptive attacks.
发表机构
- University of Macau(澳门大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。