AI 中文总结
该研究针对VLMs的排版攻击问题,提出模型无关无训练的黑盒防御方法QuISE,通过文本定位、语义替换及答案一致性判断提升防御性能,在多基准和模型上取得良好效果。
AI 中文摘要
排版攻击对视觉语言模型(VLMs)构成严重威胁,攻击者向图像中注入误导性文本,使模型依赖对抗性文本线索而非视觉证据。现有防御方法常需针对特定模型修改、额外训练或访问模型内部组件,限制了其在现代闭源VLMs中的适用性。本文提出QuISE,一种基于查询无关语义编辑的模型无关、无训练的黑盒防御方法。QuISE首先通过感知影响的文本定位识别可能影响当前查询的文本区域,随后用两个与查询和图像均无关的语义截然不同的替换文本替换这些区域,最终答案由编辑后图像的答案一致性决定。在三个排版攻击基准、四种攻击设置及四个VLMs上的大量实验表明,QuISE可稳定提升防御准确率,其恢复率达67.9%-75.0%,危害率为0.5%-1.1%。
英文摘要
Typographic attacks pose a critical threat to vision-language models (VLMs) by injecting misleading text into images and causing models to rely on adversarial textual cues rather than visual evidence. Existing defenses often require model-specific modifications, additional training, or access to internal model components, limiting their applicability to modern closed-source VLMs. In this paper, we propose QuISE, a model-agnostic, training-free black-box defense based on query-irrelevant semantic editing. QuISE first identifies text regions likely to affect the current query through influence-aware text localization. QuISE then replaces these regions with two semantically distinct replacement texts that are irrelevant to both the query and the image. The final answer is determined by answer consistency across the edited images. Extensive experiments on three typographic-attack benchmarks, four attack settings, and four VLMs show that QuISE consistently improves defended accuracy. QuISE achieves a recovery rate of 67.9-75.0% with a harm rate of 0.5-1.1%.
Comments10 pages, 7 figures; includes supplementary material