arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过视觉-语言模型的响应轮廓检测对抗性图像

Detecting Adversarial Images through Response Profiles of Vision-Language Models

Arash Vashagh, Roozbeh Razavi-Far

arXiv 2610.10436首次发表:更新:

发表机构

University of New Brunswick(新不伦瑞克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用视觉-语言模型对多种语义提示的响应轮廓来检测对抗性图像,通过轻量级分类器实现强判别力,并优于嵌入几何基线。

AI 中文摘要

对抗性扰动可以改变冻结视觉-语言模型(VLM)的预测,同时使其置信度和图像-文本相似性模式看似合理。我们研究是否可以根据图像与一组通用语义提示交互的更广泛方式来识别对抗性输入。我们的检测器使用类别级统计、提示之间的关系、与干净参考分布的偏差以及弱图像变换下的稳定性来总结这些响应,生成一个紧凑的响应轮廓,由轻量级模型进行分类,而VLM保持固定。我们在多个公共图像数据集、多种CLIP风格视觉骨干网络以及一系列基于梯度、基于优化、自动化和空间攻击上评估了该方法。检测器在攻击特定设置中实现了强判别力,并在评估训练期间未见过的攻击时保持了显著性能。在受控的检测器特定协议下,响应轮廓表示优于评估的嵌入几何基线。进一步分析表明,特征组提供互补信息,且该方法在提示配置变化下仍然有效。我们还检查了推理成本和针对检测器感知的自适应攻击的性能。总体而言,结果表明跨语义提示的响应模式为冻结VLM中的对抗性图像检测提供了有用的互补信号。

英文摘要

Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑