arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19075cs.CVcs.AIcs.CL

重新权衡证据:校准令牌级有序视觉证据以减轻大型视觉语言模型的幻觉

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

  • Pohang University of Science and Technology (POSTECH)(浦项科技大学(POSTECH))

机构由 AI 辅助整理,请以论文原文为准。

Jihae Jeong, Junha Choi, Hwanjo Yu

中文总结 AI 辅助

ReWEIGH是一种无需训练的解码干预方法,通过聚合视觉位置的令牌排名并施加有界惩罚,在多个7B至32B规模的LVLM上减少了最多21.3%的幻觉对象提及,且延迟增加极少。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)常产生幻觉,生成输入图像不支持的内容。解码时防止此类内容需要针对候选的、衡量图像对当前令牌支持强度的指标。模型的视觉令牌状态是该证据的自然来源,因为将每个状态通过输出头投影后,可得到该位置偏好的词汇项。这些位置级读数无法直接聚合,因为其概率幅度在不同视觉位置间不可比。词汇排名提供了与尺度无关的聚合基础,但令牌在基于排名的典型证据上仍存在系统性差异。我们提出ReWEIGH,一种无需训练的解码干预方法,它聚合视觉位置间的这些排名,并将每个候选令牌与从未标记图像估计出的令牌特定参考值进行比较。推理时,ReWEIGH在预填充阶段缓存图像证据,仅对低于其参考值的候选施加有界惩罚。在四个7B规模的主干模型上,ReWEIGH使幻觉对象提及减少最多21.3%,同时基本保留或提升描述性和通用性能。在缓存证据的情况下,平均每令牌的额外延迟为1.33%,且该减少效果可扩展至六个架构家族、32B参数规模的模型。

英文摘要

Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content during decoding calls for a candidate-specific measure of how strongly the image supports the token under consideration. The model's visual-token states offer a natural source of this evidence because projecting each state through the output head reveals which vocabulary items that position favors. These position-wise readouts cannot be pooled directly because their probability magnitudes are not comparable across visual positions. Vocabulary ranks provide a scale-invariant basis for pooling, but tokens still differ systematically in their typical rank-based evidence. We propose ReWEIGH, a training-free decoding intervention that aggregates these ranks across visual positions and compares each candidate with a token-specific reference estimated from unlabeled images. At inference, ReWEIGH caches the image evidence during prefill and applies a bounded penalty only to candidates that fall below their reference. On four 7B backbones, ReWEIGH reduces hallucinated object mentions by up to 21.3% while largely preserving or improving descriptive and general performance. With evidence cached, the average added latency is 1.33% per token, and the reductions extend across six architecture families to 32B parameters.

↑