arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对LLaVA中物体幻觉背后的注意力头

Targeting the Attention Heads Behind Object Hallucination in LLaVA

Armaan Sandhu, Abhilasha Senapati, Hima Kammachi

arXiv 2608.24966首次发表:更新:

发表机构

University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLaVA的物体幻觉问题,通过注意力头诊断筛选出32个关键注意力头,结合LoRA适配器与推理控制器,在COCO数据集上显著降低了物体幻觉率。

AI 中文摘要

诸如LLaVA-1.5-7B这类视觉语言模型在生成描述时,常出现幻觉(生成图像中不存在的物体)。本文探究对该失效的可解释性诊断能否指导针对性修复,并测量该修复的实际改变。我们根据注意力头在幻觉物体词汇附近的图像注意力下降程度对其排序,再通过消融候选注意力头并测量幻觉标记对数概率的变化来筛选,最终得到32个注意力头的集合。我们对这些注意力头实施两种干预:头部切片LoRA适配器和推理时的 grounding 控制器。在400个保留的COCO图像上,该组合方法将CHAIRs(含幻觉物体的描述占比)从0.370降至0.230,CHAIRi(幻觉物体提及的占比)从0.156降至0.096(p<0.001,配对符号翻转检验)。两项对照实验强化了归因:随机头部LoRA对照(按层匹配且训练方式相同)在200个独立对照图像集上的表现未优于匹配基线,支持头部选择而非LoRA容量的作用;在固定解码预算下,CHAIR的降低效果持续存在且随预算增长(64个标记时为23%,128个标记时为58%),排除了纯最大标记或截断人为因素,尽管该方法仍更简短保守。该行为减少了无依据的物体提及,同时降低了物体召回率(从0.78降至0.70)。本文提出了从诊断到干预的物体幻觉流程,更重要的是,对基于诊断信号的干预实际效果给出了可控说明:其定位了具有真实非随机影响力的干预位点,以行为特征而非单一分数呈现。

英文摘要

Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure can guide a targeted fix, and we measure what that fix actually changes. We rank attention heads by how much their image attention drops around hallucinated object words, then screen the shortlist by ablating candidate heads and measuring the change in hallucination-token log probability, yielding a 32-head set. We restrict two interventions to these heads: a head-sliced LoRA adapter and an inference-time grounding controller. On 400 held-out COCO images, the combined method lowers CHAIRs (the fraction of captions with a hallucinated object) from 0.370 to 0.230 and CHAIRi (the fraction of hallucinated object mentions) from 0.156 to 0.096 (p < 0.001, paired sign-flip tests). Two controls sharpen attribution. A random-head LoRA control, matched layer-for-layer and trained identically, performs no better than the matched baseline on a separate 200-image control split, supporting the role of head selection rather than LoRA capacity. Under fixed decoding budgets, the CHAIR reduction persists and grows with budget (23% at 64 tokens to 58% at 128), arguing against a pure max-token or truncation artifact, although the method remains shorter and more conservative. The resulting behavior reduces unsupported object mentions while also lowering object recall (0.78 to 0.70). We present a diagnosis-to-intervention pipeline for object hallucination, and, more importantly, a controlled account of what acting on the diagnostic signal actually does: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.

Comments10 pages, 5 figures, 3 tables. Accepted at the Actionable Interpretability Workshop, COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑