arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoGround:视觉-语言模型中模态干扰的测量与缓解

MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models

Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà, Roberto Dessì

arXiv 2609.33431首次发表:更新:

发表机构

Sapienza University of Rome; Harvard University; UC San Diego; Paradigma Not Diamond(罗马萨皮恩扎大学; 哈佛大学; 加州大学圣地亚哥分校; Paradigma Not Diamond)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MoGround数据集,通过单模态可回答性保证测量视觉-语言模型中的模态干扰,发现干扰与模态基础强度相关,并利用权重空间鲁棒性向量在七个模型上有效缓解干扰,同时保持多模态任务性能。

AI 中文摘要

我们发布了MoGround,一个涵盖四个视觉领域的视觉-语言数据集,其中每个问题的答案保证仅从单一模态中可获得。这一保证使我们能够测量模态干扰,即模型仅从一种模态正确回答问题,但一旦添加来自另一种模态的不相关内容后,便转而给出错误答案的失败现象。现有的探测方法很少以这种方式建立单模态可回答性,因此难以首先隔离干扰。在七个开源视觉-语言模型中,我们发现模态干扰并非普遍存在,而是依赖于具体模型。基础较弱的模态更容易受到干扰(相关系数r = +0.86),且干扰程度与基础强度成反比(相关系数r = -0.90)。单模态保证还使得一种需要区分相关与不相关上下文的缓解方法成为可能。仅基于MoGround的一个分割进行训练,一个权重空间鲁棒性向量将七个模型上的干扰降低了9%至51%,而在标准多模态任务上的平均准确率仅损失0.1个百分点。

英文摘要

We release MoGround, a vision-language dataset spanning four visual domains in which the answer to every question is guaranteed to be available from exactly one modality. This guarantee enables us to measure modality distraction, the failure in which a model answers a question correctly from one modality alone and then flips to a wrong answer once irrelevant content from the other modality is added. Existing probes rarely establish single-modality answerability this way, making it hard to isolate distraction in the first place. Across seven open-source VLMs, we find that modality distraction is not universal but model-dependent. The weaker-grounded modality is the more distracted one (r = +0.86), and distraction scales inversely with grounding strength (r = -0.90). The single-modality guarantee also enables a mitigation method that needs to distinguish between relevant and irrelevant context. Trained on one split of MoGround alone, a weight-space robustness vector reduces distraction on all seven models by 9% to 51%, at a cost of only 0.1 average points of accuracy on standard multimodal tasks.

Comments9 pages of the main paper with 3 tables and 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑