arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17205cs.CLcs.CV

哪个来源更胜一筹?视觉-语言模型中依赖关系的任务依赖性

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

Rodela Ghosh, Aviral Gupta, Guangjing Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究视觉-语言模型的模态依赖,通过构建算术与图表冲突基准,发现其依赖关系随任务、证据结构等变化,不同模型在对应基准中呈现相反的模态依赖模式。

中文摘要 AI 辅助

视觉-语言模型(VLMs)结合图像与文本,但当二者存在冲突且其中一方更难识别时,模型如何调整对二者的依赖尚不明确。我们通过受控设置研究这种模态重新分配:在保持另一方清晰的前提下,将图像或文本的可读性降至四个等级,并追踪模型偏好的变化。我们通过将一个算术问题的渲染图像与另一个的文本配对,从GSM8K和SVAMP中构建冲突,使两个来源支持不同答案。我们还引入ChartQA-Conflict,这是一个人工审核的基准,包含229个图表-报告冲突,具有匹配的图表和表格图像表示。我们使用生成的答案和长度归一化的条件对数似然边际评估六个开放权重VLMs。在GSM8K和SVAMP上,六个模型中的五个更强烈地转向远离退化的文本而非退化的图像。在ChartQA-Conflict上,所有六个似然评分模型表现出相反的模式,更强烈地转向远离退化的视觉来源。在校准单模态准确性损失以及将图表替换为普通表格图像后,这种反转仍然存在。两个前沿API模型GPT-5.6-Luna和Gemini-3.5-Flash在行为上复制了ChartQA-Conflict的反转,其中GPT-5.6-Luna还匹配算术方向。这些结果表明,VLMs中的模态依赖并非固定,而是随任务、证据结构、模型和评估设置而变化。源代码可在this https URL获取。

英文摘要

Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model's preference changes. We build conflicts from GSM8K and SVAMP by pairing the rendered image of one arithmetic problem with the text of another, so the two sources support different answers. We also introduce ChartQA-Conflict, a manually reviewed benchmark of 229 chart-report conflicts with matched chart and table-image representations. We evaluate six open-weight VLMs using both generated answers and a length-normalized conditional log-likelihood margin. On GSM8K and SVAMP, five of six models shift more strongly away from degraded text than from degraded images. On ChartQA-Conflict, all six likelihood-scored models exhibit the opposite pattern, shifting more strongly away from the degraded visual source. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images. Two frontier API models, GPT-5.6-Luna and Gemini-3.5-Flash, behaviorally replicate the ChartQA-Conflict reversal, with GPT-5.6-Luna also matching the arithmetic direction. These results show that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings. The source code is available at https://github.com/Ro-netizen004/multimodal-arbitration-artifact.

发表机构

  • University of South Florida(南佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑