发表机构
School of Computer Science and Engineering, Sun Yat-sen University; Peng Cheng Laboratory; University of Nottingham Malaysia; Hong Kong University of Science and Technology(中山大学计算机科学与工程学院; 鹏城实验室; 马来西亚诺丁汉大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TruthLens 是一种无需额外成本的 LVLM 对象幻觉检测框架,通过微调 LM 头部输出真实性分数,在 MS-COCO 数据集上的 Qwen2.5-VL-7B 上实现了超 17% 的 AUROC 提升,性能达当前最优。
AI 中文摘要
尽管大型视觉语言模型(LVLM)已取得显著进展,但对象幻觉仍是阻碍其可信部署的核心挑战。一项关键发现为本文研究提供了动机:真实对象与幻觉对象的 token 在隐藏表示中可清晰区分,但这种可分性在语言建模(LM)头部基本丧失。本文提出 TruthLens,这是一种无需辅助模型或额外推理成本的自评估框架,用于训练 LM 头部以暴露每个对象的真实性信号。具体而言,将一个极少使用的特殊 token 重新用作参考 token;对于每个对象-token 位置,提取 LM 头部分配给该特殊 token 的对数概率,并将其与预定义常数的差值定义为真实性分数。随后使用均方误差(MSE)目标对模型进行微调,该目标驱动真实对象的分数趋近于 1、幻觉对象的分数趋近于 0,同时通过发散约束保留模型原有的生成能力。尽管仅在有限的对象类别上进行训练,TruthLens 仍能有效泛化至标签空间大得多的基准数据集。在多个 LVLM 上开展的大量实验表明其达到了当前最优性能;值得注意的是,在 Qwen2.5-VL-7B 上,TruthLens 在 MS-COCO 数据集上的 AUROC 较此前最优方法提升了 17% 以上。本文代码可通过该 https URL 获取。
英文摘要
Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustworthy deployment. A key finding motivates our work: real and hallucinated object tokens are clearly separable in hidden representations, yet this separability is largely lost at the language-modeling (LM) head. We propose TruthLens, a self-evaluation framework that teaches the LM head to expose a per-object truthfulness signal without any auxiliary model or additional inference cost. Concretely, a rarely-used special token is repurposed as a reference token. For each object-token position, we extract the log-probability assigned to this special token by the LM head, and define its difference from a predefined constant as the truthfulness score. The model is then fine-tuned with an MSE objective that drives scores toward 1 for real objects and 0 for hallucinated ones, while a divergence constraint preserves the original generation capability. Despite being trained on only a limited set of object categories, TruthLens generalizes effectively to benchmarks with substantially larger label spaces. Extensive experiments across multiple LVLMs demonstrate state-of-the-art performance; notably, on Qwen2.5-VL-7B, TruthLens outperforms the previous best method on MS-COCO by over 17\% in AUROC. Our code is available at https://github.com/wyqstan/TruthLens.
CommentsAccepted by ECCV 2026