arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LookBack:通过视觉参考使用情况对LVLM响应进行评分的位置与方法

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim

arXiv 2608.11847首次发表:更新:

发表机构

Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LVLM响应存在的图像相关幻觉问题,提出无需训练的LookBack评分方法,结合视觉回溯评分增强token似然,在四个基准和三个模型上提升了Best-of-$N$选择性能且开销极小。

AI 中文摘要

大型视觉语言模型(LVLMs)将视觉感知与语言生成相结合,能够生成涵盖图像理解与复杂推理的响应。然而,LVLMs不仅继承了文本层面的幻觉,还会针对图像产生幻觉,生成与所见内容不符的流畅响应,这使得LVLM响应评分本身更具挑战性。我们的诊断结果表明,从大型语言模型(LLMs)沿用的现有基于置信度的指标对LVLMs而言并不适用:移除输入图像几乎不会改变基于置信度的选择,这表明输出空间的置信度主要捕捉文本合理性,而非与图像的一致性。为解决这一差距,我们提出LookBack,一种无需训练的LVLM响应评分方法,该方法通过视觉回溯评分增强token似然,视觉回溯评分是衡量每个响应token对图像token的引用强度的轻量指标。在四个基准和三个模型上,LookBack相较于现有基线始终提升了Best-of-$N$选择的性能,且额外开销可忽略不计。

英文摘要

Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see. This makes LVLM response scoring inherently harder, and our diagnostics show that existing confidence-based metrics adopted from LLMs are insufficient for LVLMs. Specifically, removing the input image barely changes confidence-based selection, suggesting that output-space confidence primarily captures textual plausibility rather than agreement with the image. To address this gap, we propose LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens. Across four benchmarks and three models, LookBack consistently improves Best-of-$N$ selection over existing baselines with negligible additional overhead.

CommentsFindings of EMNLP 2026. 24 pages, 12 figures. Code: https://github.com/bscho333/LookBack

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑