arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于合作原则评估视觉语言模型中的视觉问答

Evaluating VQA in Vision Language Models using Cooperative Principles

Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal

arXiv 2610.02878首次发表:更新:

发表机构

University of Memphis; Adobe Research; University of Maryland Baltimore County(孟菲斯大学; Adobe研究院; 马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估视觉语言模型在视觉问答中面对违反格赖斯准则的问题时的表现,发现模型性能下降,并揭示人类与模型在语用推理及解决违规上的差异。

AI 中文摘要

我们评估了视觉语言模型在视觉问答(VQA)中,当问题违反格赖斯准则时的表现。为此,我们使用视觉语言模型生成问题修饰符,这些修饰符添加了非必要、模糊或虚假的信息,并表明在存在此类违反的情况下,我们所评估的视觉语言模型(ChatGPT、Claude、Gemini和Llava)表现出性能下降。此外,我们实证展示了人类与视觉语言模型在语用推理上的差异,以及视觉语言模型在解决人类引发的违规与人工智能生成的违规时推理的差异。最后,我们表明,人类在解决视觉语言模型引发的违规时认知努力(通过实验中的任务时间衡量)较低,但视觉语言模型本身在此类情况下表现准确性较低。

英文摘要

We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑