AI 中文总结
本文针对多模态大语言模型(MLLMs)面临的情境错觉问题,构建分类体系并推出MSIBench基准,发现27种模型配置均易受6类错觉影响,通过提示工程和监督微调最多提升性能20%。
AI 中文摘要
现实世界的情境表象可能与其 underlying 物理状态存在偏差,这对多模态大语言模型(MLLMs)在实际应用中的可靠性构成了挑战。本文将该现象命名为情境错觉,并研究两个问题:(1)MLLMs在这类错觉下的表现如何;(2)如何缓解这些局限性。我们首先构建了全面的“位置-对象-成因”分类体系,用于描述情境错觉发生的位置、涉及的目标以及产生的方式。基于该分类体系,我们推出了MSIBench——一个旨在评估MLLMs在情境错觉下的辨别、理解和推理能力的基准。对27种模型配置的评估显示,当前MLLMs极易受这类错觉影响,且呈现出与视觉观察、接地和推理相关的6种典型失败模式。为缓解这些局限性,我们基于系统检查和推理视觉证据以实现情境理解的核心思路,分别开发了针对闭源模型的提示工程方法和针对开源模型的监督微调方法。这两种简单却有效的方法最多可将模型性能提升20%,为在复杂现实环境中实现更可靠的多模态感知与推理提供了可行路径。
英文摘要
Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.
Comments9 pages, 5 figures