发表机构
School of Computing and Information Systems; The University of Melbourne(计算与信息系统学院; 墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VisER 是一种无训练双向指标,通过视觉证据与视觉依赖两个互补视角检测 LVLMs 的物体幻觉,在 AUROC、AUPR 指标上优于多个基线方法。
AI 中文摘要
物体幻觉是大视觉语言模型(LVLMs)中持续存在的可靠性问题,模型生成的物体提及可能听起来合理但缺乏视觉依据。近期的无训练检测器利用 token 似然、注意力、视觉置信度或图文相似度等内部信号识别幻觉物体,这些信号虽有用,但常存在来源混淆问题,它们仅衡量物体在模型内部获得的支持强度,无法区分该支持来自物体特定的视觉证据还是生成的文本前缀。在困难案例中,幻觉物体仍可能因契合场景、关联附近视觉线索或符合生成文本前缀的逻辑而获得高内部支持。本文提出 VisER,一种用于物体级幻觉检测的无训练双向指标,VisER 从两个互补视角评估每个生成的物体提及:视觉证据衡量物体-上下文的兼容性是否有图像 token 提供的物体特定证据支持;视觉依赖衡量物体获得的支持更多来自图像还是生成前缀。结合这两个视角可得到更具来源感知的依据分数,同时避免额外的物体级验证生成。在多个 LVLMs 和基准测试中,VisER 在 AUROC 和 AUPR 指标上优于一系列基线方法。
英文摘要
Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal signals such as token likelihood, attention, visual confidence, or image-text similarity to identify hallucinated objects. These signals are useful, but they are often source-confounded. They measure how strongly an object is supported inside the model without distinguishing whether that support comes from object-specific visual evidence or the generated text prefix. In difficult cases, a hallucinated object can still receive high internal support because it fits the scene, is associated with nearby visual cues, or follows naturally from the generated text prefix. We propose VisER, a training-free two-sided metric for object-level hallucination detection. VisER evaluates each generated object mention from two complementary views. Visual Evidence measures whether object-context compatibility is backed by object-specific evidence from image tokens. Visual Reliance measures whether the object is supported more by the image than by the generated prefix. Combining these views gives a more source-aware grounding score, while avoiding additional object-level verification generations. Across multiple LVLMs and benchmarks, VisER improves AUROC and AUPR over a range of baselines.
CommentsAccepted at EMNLP 2026 Main Conference