arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于隐式特征稳定化的大视觉语言模型幻觉缓解方法

Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization

Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu, Prashant Khanduri

arXiv 2608.29924首次发表:更新:

发表机构

Wayne State University; Institute for AI and Data Science (AIDaS), Wayne State University(韦恩州立大学; 韦恩州立大学人工智能与数据科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大视觉语言模型的幻觉问题,提出隐式特征稳定化框架INFUSE,通过微调内置扰动不变性,在多基准模型上大幅降低幻觉率且无推理开销。

AI 中文摘要

大视觉语言模型(LVLMs)易产生幻觉:它们会流畅描述图像中不存在的物体、属性和场景。我们将这种失效的部分原因与模型表征的可测量属性——特征不稳定关联起来:输入的轻微语义保留扰动会导致学习到的嵌入发生大幅变化;幻觉率随这种变异性升高而上升。现有基于稳定性的补救措施是显式的,即它们在推理时通过潜在空间调控或约束解码进行干预,且每次查询都要付出代价。我们转而提出隐式稳定化方法:在微调阶段将扰动不变性内置到模型权重中,部署时无需额外运行任何组件。我们的框架INFUSE首先围绕扰动平均和真实锚点稳定视觉与文本表征,随后通过双向对比目标对齐跨模态的稳定表征。我们证明,锚点与扰动平均表征的均方根偏差随视图数量K以1/√K的速率缩小,且在利普希茨解码器下,这可限定任何扰动改变模型幻觉行为的幅度。在LLaVA-1.5、LLaVA-1.6和Qwen3-VL-8B-Instruct上,INFUSE使各基础模型的AMBER CHAIR相对降低46%-63%,同时提升了ObjHal、MMHal、HallusionBench和POPE的指标,且保留了VQA-v2和TextVQA的性能,全程无推理时开销。

英文摘要

Large Vision-Language Models (LVLMs) are prone to hallucinations: they fluently describe objects, attributes, and scenes that are not in the image. We connect part of this failure to a measurable property of their representations, feature instability, where mild semantics-preserving perturbations of the input cause large changes in the learned embeddings; hallucination rates rise together with this variability. Existing stability-motivated remedies are explicit, in the sense that they intervene at inference time through latent steering or constrained decoding, and pay for it on every query. We propose implicit stabilization instead: perturbation-invariance is built into the model weights during fine-tuning, and nothing extra runs at deployment. Our framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives. We prove that the anchor's root-mean-square deviation from the perturbation-mean representation shrinks at rate $1/\sqrt{K}$ in the number of views, and that under a Lipschitz decoder, this bounds how much any perturbation can change the model's hallucination behavior. On LLaVA-1.5, LLaVA-1.6, and Qwen3-VL-8B-Instruct, INFUSE reduces AMBER CHAIR by 46-63% relative to each base model, improves ObjHal, MMHal, HallusionBench, and POPE, and preserves VQA-v2 and TextVQA, all with no inference-time overhead.

Comments28 Pages, 12 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑