发表机构
University of Central Florida; Institute of Artificial Intelligence(中佛罗里达大学; 人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型视觉语言模型中的物体幻觉问题,提出无需训练的CORAL框架,利用不确定性感知的视觉数据分割和镜像统计量控制错误发现率,在多个基准上超越现有方法。
AI 中文摘要
多重物体幻觉是指大型视觉语言模型(LVLMs)生成视觉输入不支持的对象,这是由解码过程中的视觉不确定性引起的持续挑战。现有方法使用对比信号来减少幻觉,但它们依赖启发式方法,缺乏在图像级别对假阳性进行原则性控制。为解决这一问题,我们提出了幻觉的错误发现率控制(CORAL),这是一个无需训练的框架,采用不确定性感知的视觉数据分割策略来建模视觉不确定性,并利用镜像统计量在解码过程中量化视觉对比。通过从配对的、对称扰动的视觉输入中计算镜像统计量,CORAL估计虚假对象预测,并设置数据驱动的阈值来控制每张图像中错误发现的预期比例,从而抑制幻觉,同时保持对真实接地对象的高检测能力。该框架灵活,支持多种LVLMs,无需重新训练或监督即可缓解幻觉。在多个基准测试和多种评估指标上的大量实验表明,CORAL始终优于最先进的方法,提供了更可靠和稳健的幻觉控制。代码可在以下网址获取:this https URL
英文摘要
Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: https://changliu1993-cl.github.io/CORAL/
CommentsAccepted to NeurIPS 2026