arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非图像内容所致:无关语境在不告知VLM裁判的情况下破坏其稳定性

It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, Lotem Peled-Cohen

arXiv 2609.37863首次发表:更新:

发表机构

Technion(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出MIST压力测试,发现VLM裁判的标签受图像存在而非内容影响,揭示可替代性测试反映配置而非模型本身。

AI 中文摘要

视觉语言模型(VLM)正越来越多地被用于替代人类标注者,因此,可替代性测试应反映模型本身而非偶然的评估条件,这一点变得至关重要。我们引入了MIST,即误导性图像压力测试:包含200个英文句子,每个句子围绕一个既可作比喻义也可作字面义理解的短语构建,并配以描绘其含义的对齐图像、描绘相反含义的误导性图像或不配图像。指导规则要求仅依据句子本身决定标签,因此任何图像都不应改变任何答案。我们预期每幅图像都会将裁判的标签拉向其描绘的含义方向,但两种图像均未产生此效果。在十三个VLM裁判中,对齐图像改变了20.5%的标签,误导性图像改变了19.4%,每个裁判的数值都相近,且均高于在保留图像但删除忽略图像指令时产生的11.6%。然而,在两幅图像之间不同的标签中,仅有37%的标签移向了所示含义,并且无论图像缺失、对齐还是误导,与人类标注者的一致性均保持不变。在通过alt-test的七个裁判中,该效应小于未通过的六个裁判,但在所有裁判中都存在:影响裁判的是图像的存在,而非两幅图像中的哪一幅,因此可替代性判定描述的是配置,同样也描述了模型。

英文摘要

Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.

CommentsAccepted at TAE (Trust-AI-Eval) @ NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑