arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

被破坏但正确:为什么视觉语言模型在内部对自己撒谎

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama

arXiv 2610.03445首次发表:更新:

发表机构

University College London; Holistic AI(伦敦大学学院; Holistic AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示了视觉语言模型在对抗扰动下训练与推理行为的分离现象,定位到语言解码器的先验作用,并指出对抗鲁棒性主要源于解码器而非视觉编码器。

AI 中文摘要

一次有针对性的对抗性扰动可以将视觉语言模型(VLM)对固定目标标题的教师强制训练损失降至接近零,然而,同一个模型在允许自由生成时,却产生原始的正确描述,且没有任何目标痕迹。我们将这种分离称为训练/推理差距,并在Qwen2.5-VL-7B-Instruct上使用受控的两阶段PGD攻击对200张保留的COCO图像进行了精确的机制解释。首先,我们表明图像级像素统计,包括从CNN鲁棒性文献中正确重新实现的基于纹理的可攻击性度量,对哪些图像被破坏几乎没有预测能力(最佳预测器r=-0.050,p=0.484;岭回归R^2=0.069)。其次,使用logit透镜,我们将差距定位到单个自回归步骤:在已经生成正确第一个标记的条件下,目标标记的排名在每张图像和每个条件下都固定在152,064个词汇条目中的第3,488位,方差为零。第三,追踪所有28个LLM解码器层中的目标标记排名,揭示视觉编码器无论最终结果如何,都以相当的幅度破坏每张图像的表示,但语言模型解码器随后进行差异仲裁:对易受影响图像放大被破坏的信号,对抵抗图像则积极抑制其超过干净图像基线(p<0.001,秩双列相关r=0.579)。在合并隐藏状态上的线性探针以AUC=0.858区分这两种结果,尽管我们指出该估计中存在循环性担忧。这些结果共同表明,自回归VLM中的对抗鲁棒性在很大程度上是语言解码器先验的属性,而非视觉编码器的属性,这对部署的VLM系统的忠实性评估和防御应针对何处具有直接意义。

英文摘要

A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p<0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.

CommentsAccepted at the VLM4RWD Workshop (Grounded and Faithful Vision-Language Models for Real-World Deployment), NeurIPS 2026. 8 pages, 2 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑