发表机构
Seoul National University; AIM Intelligence(首尔大学; AIM智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态大语言模型的物体幻觉问题,推出细粒度多图像物体幻觉基准测试MIOH,评估29个模型后发现顶尖模型仍存在明显失败模式,为开发可靠多模态AI提供关键评估工具。
AI 中文摘要
多模态大语言模型(MLLMs)正越来越多地部署在需要跨视觉上下文进行复杂推理的多图像场景中。然而,当前MLLMs仍存在根本性局限:物体幻觉,即生成看似合理但与事实不符的物体描述。现有基准测试主要针对单图像场景,或仅提供多图像的高级评估,无法系统诊断视觉复杂性和推理需求如何触发幻觉。为填补这一空白,我们推出MIOH——细粒度多图像物体幻觉基准测试,它通过三种多图像推理模式(综合、对比、选择性),在三种可控对抗压力(视觉上下文规模、感知难度、上下文偏差)下,系统评估四大基础任务(存在性、计数、属性、位置)中的物体幻觉。通过对29个模型的评估,我们发现即便是GPT-5和Gemini-2.5-Pro等最先进系统,在不同推理模式和任务中也存在明显的失败模式。我们的评估表明,幻觉不仅源于感知失败,还源于跨多图像维持物体表征时的整合阶段局限。MIOH为分析多图像物体幻觉提供了可控框架,是开发更可靠多模态AI系统的关键评估工具。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
CommentsAccepted at CVPR 2026
Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 18295-18305