发表机构
National Taiwan University; NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(国立台湾大学; 台湾大学人工智能卓越研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tri-PvP通过8,000样本三模态冲突基准,揭示全模态大语言模型存在稳健视觉偏差及视觉感知、音频命题的不对称证据形式偏差,且该偏差早期可解码、难以彻底缓解。
AI 中文摘要
全模态大语言模型(OLLMs)联合处理视觉、音频和文本,然而它们在跨模态冲突下的模态偏差仍未得到充分探索。现有基准将同一模态内的两种不同形式的证据混为一谈:感知信号(例如,一张狗的照片或录音)和命题信号(例如,陈述性声明“这是一只狗”),因此任何测得的模态偏差都固有地与证据形式偏差混杂在一起,无法清晰地归因于任一来源。为解决这一问题,我们引入了Tri-PvP,一个包含8,000个样本的三模态冲突基准,跨越视觉、音频和文本,其中视觉和音频各采取感知或命题形式。通过评估五个OLLMs,我们发现大多数模型和证据类型条件下存在稳健的视觉偏差。至关重要的是,我们揭示了证据形式偏差中的系统性不对称:模型在视觉中表现出对感知信号的更强偏差,而在音频中则对命题信号表现出更强偏差。通过逐层线性探测和对比解码的进一步分析表明,模态偏差已经可以从早期表示层线性解码,并且只能被部分缓解,这呼吁需要超越表面干预的缓解策略。
英文摘要
Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim "this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to either source. To address this, we introduce Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robust visual bias across most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses via layer-wise linear probing and contrastive decoding reveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.
CommentsEMNLP 2026 Findings