发表机构
George Mason University(乔治梅森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究评估6种LVLM幻觉缓解方法,发现其虽降低幻觉率但会削弱信息性,且无法迁移到多模态能力,提出应从忠实性-信息性-能力权衡评估幻觉缓解。
AI 中文摘要
近期针对大型视觉语言模型(LVLMs)的推理时幻觉缓解方法在幻觉基准上取得了显著提升,但目前仍不清楚更低的幻觉分数是反映了多模态 grounding 的提升,还是更保守的生成。我们评估了三种 LVLMs 和四个基准上的六种缓解方法,包括聚焦幻觉的评估和多样化能力基准 MMStar。我们的分析揭示了两个一致模式:第一,幻觉减少往往伴随信息性降低,即降低幻觉率的方法会同时减少对象召回、视觉覆盖或响应详细程度;第二,幻觉基准上的改进无法可靠迁移到更广泛的多模态能力,这些方法在细粒度感知和推理任务上表现不一致或性能下降。我们的发现表明,当前评估协议可能通过奖励保守生成而高估了进展,我们认为幻觉缓解应作为忠实性-信息性-能力的权衡来评估,而非仅通过幻觉分数。
英文摘要
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.