arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00067cs.CLcs.AI

多模态大语言模型在阅读前是否能“看见”?诊断语境谄媚性

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

Yi-Cheng Lai, Hen-Hsen Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究诊断多模态大语言模型的语境谄媚性,通过含998个案例的测试探究信息边界对该现象的影响,发现System-2视觉仲裁(S2VA)可显著提升模型表现,且该现象受文本引入时机等因素影响。

中文摘要 AI 辅助

外部文本可覆盖多模态大语言模型中相互冲突的图像证据,这一失效现象我们称之为多模态语境谄媚性。我们引入了一个含998个案例的诊断测试,该测试独立变化视觉证据、常识先验和外部文本,通过调整与语境无关的视觉见证者周围的信息边界,探究这种失效现象何时出现。在异常图像与Gemini生成的虚假文本配对的情况下,GPT-5.1在联合条件下的得分为7.9%,直接对与语境无关的见证者报告评分时得分为49.7%,在让见证者接触文本的匹配双调用见证者-仲裁者流程下得分为63.7%,在将文本对见证者隐藏的System-2视觉仲裁(S2VA)下得分为84.2%。在六个模型中,S2VA相比直接见证者报告提升了19.7至44.1个百分点,所有配对的95%置信区间均不包含零。最佳信息边界并非统一:文本语境对部分模型有辅助作用,且GPT-4o重新生成的子集改变了联合条件、仅见证者和S2VA的相对排序。因此,语境谄媚性对文本引入的时机、模型及语境来源均敏感。

英文摘要

External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.

发表机构

  • Institute of Information Science, Academia Sinica(中央研究院资讯科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑