arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07110cs.CVcs.AI

发现黑盒视觉模型中的自然变换漏洞

Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models

Dongsu Song, DaeYun GO, Jay Hoon Jung

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于查询的黑盒对抗场景攻击(ASA),利用多模态语言模型和文本引导生成器搜索自然编辑场景,以高成功率、低查询数发现视觉模型可复用的自然变换漏洞。

中文摘要 AI 辅助

自然对抗样本(NAEs)揭示了视觉模型在超出范数有界扰动的现实语义变化下可能失效。然而,在黑盒设置下生成NAEs仍然具有挑战性,因为现有的生成式攻击通常依赖代理模型、学习到的攻击先验或代价高昂的基于查询的优化,而暴露模型漏洞的自然变换事先是未知的。我们提出了对抗场景攻击(ASA),这是一种基于查询的黑盒框架,利用多模态语言模型和现代文本引导的生成编辑器,在自然语言编辑场景中进行搜索。ASA通过赢家-输家反馈联合探索背景、天气和材质/颜色变换,并使用贪婪探索器仅组合能提升攻击效果的场景。在多种ImageNet分类器上,ASA实现了比先前基于查询的生成式攻击显著更高的攻击成功率,同时需要更少的受害模型查询,并保持了具有竞争力的感知质量。此外,ASA展现出图像级和提示级可迁移性:其对抗图像在受害模型架构间保持有效,而其发现的编辑场景可复用于同类图像,在某些情况下还可跨架构复用。这些发现表明,视觉模型对自然变换模式具有可复用的漏洞,ASA可以在黑盒设置下高效识别这些漏洞。

英文摘要

Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond norm-bounded perturbations. However, generating NAEs in a black-box setting remains challenging because existing generative attacks often rely on surrogate models, learned attack priors, or costly query-based optimization, whereas the natural transformations that expose model vulnerabilities are unknown a priori. We propose \textbf{Adversarial Scenario Attack (ASA)}, a query-based black-box framework that searches over natural-language editing scenarios using a multimodal language model and a modern text-guided generative editor. ASA jointly explores background, weather, and material/color transformations through winner--loser feedback, and uses a greedy explorer to compose only attack-improving scenarios. Across diverse ImageNet classifiers, ASA achieves substantially higher attack success rates than prior query-based generative attacks while requiring fewer victim-model queries and preserving competitive perceptual quality. Moreover, ASA exhibits both image-level and prompt-level transferability: its adversarial images remain effective across victim-model architectures, while its discovered editing scenarios can be reused across same-class images and, in some cases, across architectures. These findings suggest that vision models possess reusable vulnerabilities to natural transformation patterns, which ASA can efficiently identify in a black-box setting.

补充信息

↑