发表机构
Shanghai Jiao Tong University; Peking University; Beijing University of Chemical Technology; Nanjing University of Aeronautics and Astronautics; University of Science and Technology of China; Tsinghua University(上海交通大学; 北京大学; 北京化工大学; 南京航空航天大学; 中国科学技术大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LVLM的幻觉问题,提出无训练框架AIMS,通过自适应协调多上下文源引导,在缓解对象幻觉的同时保持通用多模态能力。
AI 中文摘要
幻觉问题仍是大型视觉语言模型(LVLM)面临的重大挑战。现有的无训练方法通常通过对比解码或视觉增强来缓解幻觉,往往会在生成过程中增加视觉证据的相对影响力。这引发了一个根本性问题:LVLM能否动态调节不同上下文源的贡献以抑制幻觉?在本研究中,我们调查并量化了LVLM在解码过程中如何协调多个上下文源,并研究这种内在行为如何用于指导幻觉缓解。我们发现,LVLM表现出内在的视觉关注倾向,可用于指导自适应视觉引导,而文本上下文也有助于缓解幻觉。基于这些发现,我们提出了AIMS(Adaptive Information Multi-source Steering,自适应信息多源引导),这是一种轻量级无训练框架,可在解码过程中自适应协调视觉、预填充文本和生成的上下文。具体而言,AIMS为三个上下文域构建紧凑原型,并估计它们与当前查询的亲和力以确定头级引导权重。生成的多源引导方向被应用于查询表示,实现自适应上下文集成,无需额外模型训练或辅助前向传播。在多个LVLM和解码策略上进行的大量实验表明,AIMS可有效缓解对象幻觉,同时保持具有竞争力的通用多模态能力。
英文摘要
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation. We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation. Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding. Specifically, AIMS constructs compact prototypes for the three context domains and estimates their affinities with the current query to determine head-wise steering weights. The resulting multi-source steering direction is applied to the query representation, enabling adaptive context integration without additional model training or auxiliary forward passes. Extensive experiments across multiple LVLMs and decoding strategies demonstrate that AIMS effectively mitigates object hallucination while maintaining competitive general-purpose multimodal capabilities.