arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00231cs.CV

超越语言先验:诊断与修正多模态大语言模型中的视觉起源幻觉

Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM

Peiyang Xu, Xiaopei Zhu, Jun Zhu, Xiaolin Hu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出多模态大语言模型的幻觉存在视觉起源因素,提出 ACFT 方法,仅用 0.9% COCO 数据即可在多个基准上实现最优幻觉修正性能

中文摘要 AI 辅助

现有关于多模态大语言模型(MLLM)中对象幻觉的研究,主要将该问题归因于语言先验,例如过度依赖文本共现统计数据。我们通过定量证据挑战这一观点,指出了一个被忽视的互补原因:视觉起源幻觉,即幻觉源于错误的视觉特征提取以及图像与文本嵌入之间的错位。通过余弦相似度分析和 Smooth Grad-CAM 熵测量,我们发现幻觉样本表现出系统更低的图像-文本相似度(平均值为 0.158,而非幻觉样本为 -0.122),且注意力模式倒置:当目标对象存在时注意力分散,而当目标对象不存在时注意力错误集中。基于这一诊断,我们提出了对抗性对比微调(ACFT)。ACFT 采用对抗性幻觉属性翻转(AHAF)流程,该流程涉及最小的针对性对抗性扰动,可翻转图像的幻觉属性,从而构建完全对齐的正负样本对,用于对比微调。AHAF 同时作为诊断探针,揭示 MLLM 的视觉表示危险地接近幻觉决策边界。仅需 COCO 数据集的 0.9% 数据,且无额外推理开销,ACFT 在 POPE、MME 及四个描述级幻觉基准上,针对 LLaVA、MiniGPT-4 和 Qwen2.5-VL 均实现了最先进的性能。代码可在该 https URL 获取

英文摘要

Existing research on object hallucination in multimodal large language models (MLLMs) predominantly attributes the problem to language priors such as over-reliance on textual co-occurrence statistics. We challenge this view by presenting quantitative evidence for a complementary, under-explored cause: visual-origin hallucination, where hallucinations arise from incorrect visual feature extraction and misalignment between image and text embeddings. Through cosine similarity analysis and Smooth Grad-CAM entropy measurements, we show that hallucinated samples exhibit systematically lower image-text similarity (average 0.158 vs. -0.122) and inverted attention patterns, where attention is dispersed when the target object is present but wrongly concentrated when it is absent. Guided by this diagnosis, we propose Adversarial Contrastive Fine-Tuning (ACFT). ACFT uses an Adversarial Hallucination Attribute Flipping (AHAF) procedure, involving minimal, targeted adversarial perturbations that flip an image's hallucination attribute, to construct perfectly aligned positive-negative pairs, which are then used for contrastive fine-tuning. AHAF simultaneously serves as a diagnostic probe, revealing that MLLM visual representations lie dangerously close to hallucination decision boundaries. Requiring only 0.9% of the COCO dataset and adding zero inference overhead, ACFT achieves state-of-the-art performance on POPE, MME, and four description-level hallucination benchmarks across LLaVA, MiniGPT-4, and Qwen2.5-VL. Code is available at https://github.com/zxp555/ACFT_MM

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑