发表机构
National University of Singapore; DAMO Academy, Alibaba Group; Renmin University of China; Zhejiang University; Hupan Lab; University of California, Berkeley; Rochester Institute of Technology; The Hong Kong University of Science and Technology(新加坡国立大学; 阿里巴巴达摩院; 中国人民大学; 浙江大学; 湖畔实验室; 加州大学伯克利分校; 罗切斯特理工学院; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型的幻觉问题,提出含UniHall基准与SAMF自演化模糊测试的评估框架,发现SOTA模型在模糊测试下性能显著下降,存在有用性-幻觉权衡。
AI 中文摘要
幻觉(hallucination)仍是多模态大语言模型(Multimodal Large Language Models, MLLMs)面临的持续挑战,严重限制了其在高风险应用中的可靠性。现有评估多基于静态基准,存在分类覆盖范围狭窄、性能快速饱和的问题,无法反映模型在不断变化的现实场景中的鲁棒性。为弥合这一差距,本文提出一种整合了全面基准与自演化压力测试的系统评估框架。首先,我们引入UniHall,这是一个基于涵盖物体、指令和知识维度的统一分类法的细粒度数据集。其次,为解决基准饱和问题,我们提出自适应性多模态模糊测试(Self-Adaptive Multimodal Fuzzing, SAMF),这是一种采用演化变异策略探索模型幻觉边界的自适应性框架。关键的是,为确保对动态输入的可靠评估,SAMF纳入了由多模态预言机集成驱动的结构化指标套件。我们的大量实验表明,与常规设置相比,最先进的MLLMs在模糊测试下表现出显著的性能下降,暴露出推理能力与事实依据之间的脱节。此外,我们发现了有用性-幻觉权衡,即强化学习对齐会无意中加剧指令跟随任务中的谄媚倾向。该框架、代码和基准可在此httpsURL获取。
英文摘要
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing. First, we introduce UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi-modal oracles. Our extensive experiments reveal that state-of-the-art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness-hallucination trade-off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction-following tasks. The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.
Comments47 pages, 17 figures