相同语义,不同结果:知识冲突下多模态大语言模型的模态鲁棒性研究
Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
浏览论文内容
中文总结 AI 辅助
该研究针对13个多模态大语言模型,发现其在知识冲突下存在模态鲁棒性不足的问题,图像形式的矛盾证据更易被接受,多模态RAG性能会受影响,监督微调可一定程度缓解该问题。
中文摘要 AI 辅助
多模态大语言模型(Multimodal large language models, MLLMs)越来越多地获得异构形式的上下文证据:文本段落、对应段落的渲染图像,或两者结合。然而,目前尚不清楚这些表面形式的处理一致性如何,尤其是当证据与模型的参数知识冲突时。我们在13个MLLMs和两个数据集上研究知识冲突下的模态鲁棒性,发现它们远非鲁棒。(1)与普遍认知相反,模型更倾向于以图像形式呈现的与参数知识矛盾的上下文,而非文本形式;(2)当同时呈现矛盾的文本和图像时,偏好的模态本质上是任意的,随输入顺序、模型和数据集变化。我们进一步证明这种不稳定性具有实际后果:它会降低多模态RAG的性能,且可被对抗性攻击利用。为缓解这种脆弱性,我们检验了几种简单技术——提示工程、引导、监督微调(supervised fine-tuning, SFT)和直接偏好优化;多数技术被证明无效,而SFT取得了中等成功。因此,我们呼吁更加关注这种不一致性,并认为它是根本性的,需要在多个训练阶段加以关注。
英文摘要
Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are processed, especially when the evidence conflicts with the model's parametric knowledge. We study modality robustness under knowledge conflict across 13 MLLMs and two datasets, and find them far from robust. (1) Contrary to common belief, models favor a context that contradicts parametric knowledge more readily in image form than in text form; (2) when a contradicting text and image are presented together, the preferred modality is essentially arbitrary, varying with input order, model, and dataset. We further demonstrate that this instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks. To alleviate this brittleness, we examine several simple techniques---prompting, steering, supervised fine-tuning (SFT), and direct preference optimization; the majority prove ineffective, whereas SFT achieves moderate success. We therefore call for greater awareness of this inconsistency and argue that it is fundamental, demanding attention at multiple training stages.
发表机构
- Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。