arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

幻觉揭示了多模态推理的哪些方面?通过对比解码探针诊断视觉接地失败

What Do Hallucinations Reveal About Multimodal Reasoning? Diagnosing Visual Grounding Failures via Contrastive Decoding Probes

Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang

arXiv 2609.16646首次发表:更新:

发表机构

Ocean University of China(中国海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出无需训练的SAFE对比解码框架,通过对比视觉接地与消融路径诊断LVLM幻觉,发现视觉依赖衰减、幻觉成簇及早期干预有效,并在MMHalBench上显著优于基线。

AI 中文摘要

当强大的多模态模型广泛可用时,进展需要超越基准分数的新科学方法论——使用模型作为理解行为的工具。我们通过提出以下问题来解决这一点:我们能否使用大型视觉-语言模型(LVLMs)作为实验工具来研究它们自身的失败动态?聚焦于视觉幻觉,我们引入了SAFE,一种无需训练的解码框架,它对比视觉接地和视觉消融的生成路径,以产生一个词元级别的对比接地分数,该分数识别模型何时偏向语言先验而非视觉证据。这一信号具有双重作用:作为检测非视觉接地词元的实用代理,以及作为解码时惩罚的基础。我们的分析得出了三个实证观察:视觉依赖在生成过程中衰减,幻觉在时间上成簇共现,早期干预减少聚类而不显著降低流畅性。在MMHalBench上,SAFE大幅优于所有比较的基线;在其他地方的结果则更为参差。我们认为设计对比探针体现了一个更广泛的使命:使用模型作为科学理解的工具。代码:此https URL。

英文摘要

When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by asking: can we use large vision-language models (LVLMs) as experimental instruments for studying their own failure dynamics? Focusing on visual hallucination, we introduce SAFE, a training-free decoding framework that contrasts visually-grounded and vision-ablated generation paths to produce a token-level contrastive grounding score that identifies when the model favors linguistic priors over visual evidence. This signal serves dual roles: as a practical proxy for detecting visually-ungrounded tokens, and as the basis for decoding-time penalties. Our analysis yields three empirical observations: visual dependency decays over generation, hallucinations co-occur in temporal clusters, and early intervention reduces clustering without substantially degrading fluency. On MMHalBench, SAFE substantially outperforms all compared baselines; results elsewhere are more mixed. We argue that designing contrastive probes exemplifies a broader mission: using models as instruments for scientific understanding. Code: https://github.com/zhaozhipeng1997/SAFE_public.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑