arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03817cs.CVcs.AI

UHP检测:大型视觉语言模型(LVLMs)在一致性空间中具有独特的幻觉模式

UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LVLMs的幻觉问题,提出UHP检测框架,通过图像与文本扰动模态、陈述与否定逻辑极性构建四组一致性特征,训练分类器,在AMBER和PhD数据集上显著优于现有基线。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)展现出强大的多模态推理能力,但仍易出现幻觉现象,即模型预测与视觉证据不相符。现有的黑盒幻觉检测方法通过单一一致性指标估计不确定性,隐含假设模型不确定性可由单一指标充分表征。然而,幻觉在不同行为探测中表现出多样化的不确定性形式,单一指标不足以表征其潜在行为。我们提出Unique Hallucination Pattern(UHP,独特幻觉模式)检测,这是一种全黑盒框架,将幻觉建模为由两个轴定义的结构化不确定性模式:扰动模态(图像与文本)和逻辑极性(陈述与其否定)。两者的交集产生四个互补的一致性组,捕捉模型不确定性的不同表现,从中提取组内和组间特征以训练轻量级分类器。在AMBER和PhD数据集上对三个LVLMs进行的综合实验表明,UHP检测始终优于现有的黑盒和白盒基线方法,与最强的黑盒方法相比,AUC-ROC提升最高达18.72%,AUC-PR提升最高达20.07%。广泛的 ablation研究表明,每个一致性组贡献互补信息,其组合形成结构化幻觉模式。此外,跨数据集评估显示,该学习到的模式可在不同基准间泛化,表明幻觉行为反映了模型特定的一致性模式。代码可在该httpsURL公开获取。

英文摘要

Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence. Existing black-box hallucination detection methods estimate uncertainty through a single consistency metric, implicitly assuming that model uncertainty can be adequately characterized by a single measure. However, hallucinations exhibit diverse manifestations of uncertainty across different behavioral probes, making a single measure insufficient to characterize their underlying behavior. We propose \emph{Unique Hallucination Pattern (UHP) Detection}, a fully black-box framework that models hallucination as a structured uncertainty pattern defined by two axes: perturbation modality (image vs.\ text) and logical polarity (a statement vs.\ its negation). Their intersection produces four complementary consistency groups that capture distinct manifestations of model uncertainty, from which both within-group and between-group features are extracted to train a lightweight classifier. Through comprehensive experiments on AMBER and PhD across three LVLMs, UHP Detection consistently outperforms prior black-box and white-box baselines, with improvements of up to $+18.72\%$ AUC-ROC and $+20.07\%$ AUC-PR over the strongest black-box methods. Extensive ablation studies demonstrate that each consistency group contributes complementary information and that their combination forms a structured hallucination pattern. Furthermore, cross-dataset evaluation shows that this learned pattern generalizes across benchmarks, indicating that hallucination behavior reflects a model-specific consistency pattern. \textbf{Code is publicly available at} https://github.com/amirezzati/uhpdet.

补充信息

↑