AI 中文总结
研究视觉语言预训练模型易受对抗攻击问题,基于其嵌入空间几何结构特点,提出GeoDetect方法,利用几何分数识别对抗样本,经综合评估,该方法能可靠检测多种VLP架构及威胁设置下的对抗样本,提升模型安全性与可靠性。
AI 中文摘要
视觉语言预训练模型(VLP)在实际应用中广泛使用,但易受对抗攻击。尽管对抗检测方法在单模态(视觉或语言)设置中取得成功,但其在VLP等多模态模型中的有效性和可靠性仍未充分探索。本文研究VLP嵌入空间的几何结构,发现其与单模态视觉模型不同的结构各向异性。理论分析表明,在此结构下,对抗攻击会增加干净样本与对抗样本之间的预期几何距离。基于此,提出GeoDetect,通过几何分数利用这些偏离流形区域的偏差来识别对抗样本。综合评估表明,该方法能可靠检测多种VLP架构和威胁设置下的对抗样本,为提高模型安全性和可靠性提供了强大且实用的方法。
英文摘要
Vision-language pre-trained models (VLPs) are widely used in real-world applications. However, they remain vulnerable to adversarial attacks. Although adversarial detection methods have demonstrated success in single-modality settings (either vision or language), their effectiveness and reliability in multimodal models such as VLPs remain largely unexplored. In this work, we study the geometry of VLP embedding spaces and observe structured anisotropy that differs from unimodal vision models. Our theoretical analysis shows that under this anisotropic structure, adversarial attacks increase the expected geometric separation between clean and adversarial examples (AEs). Specifically, we demonstrate that AEs consistently exhibit greater expected distances to randomly sampled points than their clean counterparts, indicating that AEs tend to push representations out of manifold regions. Building on these insights, we propose GeoDetect, which leverages these off-manifold deviations via geometric scores to identify AEs. Through comprehensive evaluations, we show that our approach reliably detects AEs across diverse VLP architectures and threat settings, covering unimodal and multimodal attacks as well as adaptive attacks, thereby providing a robust and practical approach to improving the safety and reliability of these models.
CommentsECCV 2026