arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同证据,不同判断:视觉/语音-文本冲突中的证据不可交换性

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong

arXiv 2609.26986首次发表:更新:

发表机构

University of Liverpool(利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过固定内容、交换位置的配对比较,揭示多模态模型中视觉/语音证据与文本冲突时存在跨模态证据不可交换性,即证据顺序影响模型判断,后置感知证据增强其依赖。

AI 中文摘要

对于多模态大语言模型,当图像或语音与伴随的文本发生冲突时,所测量的文本依赖度可能将模态偏好与证据位置纠缠在一起。早期关于文本偏差的研究通常使用固定的证据顺序,或随证据移动任务指令,导致顺序的贡献不明确。在本文中,我们采用配对比较方法,保持指令和证据内容固定,仅交换两个来源的位置,以量化这种潜在影响。在视觉和语音模型中,将图像或录音置于冲突文本之后,始终使答案向其内容偏移。我们还重新审视了先前的研究,并分析了为何其实验设置可能导致误导性结论。这些发现揭示了跨模态证据的不可交换性:当证据顺序改变时,相同证据可能导致不同判断,且将感知证据置于较后位置可增加模型对其内容的依赖。

英文摘要

For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence noncommutativity: the same evidence can lead to different judgments when its order changes, and placing perceptual evidence later can increase the model's reliance on its content.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑