发表机构
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大型视觉语言模型的幻觉问题,提出无需训练的EviAnchor框架,通过区域证据锚点和决策条件证据路由提升视觉证据利用率,在多个基准测试中改善了视觉接地性能。
AI 中文摘要
大型视觉语言模型(LVLMs)经常生成与视觉输入不符的内容。初步实验显示,视觉证据主要被整合到解码器中早期至中期的答案侧表示中,而其直接影响在后续层中逐渐减弱。这种衰减表明,早期获取的视觉证据在后续生成过程中可能未被充分利用。基于这一观察,我们提出了EviAnchor,这是一种无需训练的单分支推理框架,可在整个生成过程中保留并重新激活视觉证据。EviAnchor引入了区域证据锚点(REA)槽,用于将密集视觉令牌逐步聚合为空间结构化表示;随后通过决策条件证据路由,增强当前决策状态对这些视觉锚点的访问,减轻对文本上下文的过度依赖;最后,模型恢复原生Transformer计算,将检索到的视觉证据与问题语义及生成历史整合。在POPE、CHAIR和MMHal-Bench上的实验表明,该方法在视觉接地方面取得了一致的改进。
英文摘要
Large vision-language models (LVLMs) frequently generate content unsupported by visual inputs. Preliminary experiments show that visual evidence is primarily incorporated into answer-side representations in early-to-middle decoder layers, while its direct influence progressively weakens in later layers. This attenuation suggests that visual evidence acquired earlier may be insufficiently utilized during subsequent generation. Based on this observation, we propose EviAnchor, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation. EviAnchor introduces Regional Evidence Anchor (REA) slots to progressively aggregate dense visual tokens into spatially structured representations. It then strengthens the current decision state's access to these visual anchors through decision-conditioned evidence routing, mitigating excessive dependence on textual context. Finally, the model resumes its native Transformer computation to integrate the retrieved visual evidence with question semantics and generation history. Experiments across POPE, CHAIR, and MMHal-Bench demonstrate consistent improvements in visual grounding.