针对基于VLM的AI生成图像检测的排版攻击
Typographic Attack Against VLM-based AI-generated Image Detection
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究系统评估了针对视觉语言模型AI生成图像检测的排版攻击,发现推理模式更脆弱且攻击存在方向不对称,较大模型虽检测更准但更易受攻击。
中文摘要 AI 辅助
视觉语言模型(VLMs)越来越多地被用于AI生成图像(AIGI)检测,为真实性判断提供自然语言解释。然而,它们解释图像内文本的能力也可能使这些判断暴露于误导性的语义线索中。我们系统地评估了针对检测导向型、开放权重型和商业型VLM的排版攻击策略,考虑了真实到虚假和虚假到真实的攻击。我们的结果表明,推理模式通常比直接模式表现出更大的脆弱性,且攻击有效性表现出显著的方向不对称性。此外,较大的模型往往具有较高的干净检测准确率,但也具有较高的攻击成功率。我们进一步研究了在图像和文本变换下的攻击鲁棒性,并调查了指示正确类别的叠加层是否有助于错误纠正。综合这些分析,我们刻画了排版攻击如何影响真实性判断,并揭示了当前基于VLM的AIGI检测系统的局限性。
英文摘要
Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.