arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29604cs.CVcs.MM

PROVE:基于证明引导的机制感知算子验证,用于医学视觉问答中的幻觉检测

PROVE: Proof-guided Regime-aware Operator Verification for Hallucination Detection in Medical Visual Question Answering

  • Fudan University(复旦大学)
  • Fudan University Shanghai Cancer Center(复旦大学附属肿瘤医院)

机构由 AI 辅助整理,请以论文原文为准。

Keyang Zhou, Siyi Li, Zhongnan Shi, Qichao Ying, Wei Tang, Zhenxing Qian

AI总结:

PROVE提出了一种黑盒幻觉检测方法,根据问题所需视觉证明类型动态调整验证策略,在医学VQA基准上以0.821 AUROC超越最强基线,显著提升检测准确性。

AI中文摘要:

在医学视觉问答(VQA)中,视觉语言模型(VLM)的幻觉可能导致自信但错误的回答,从而增加诊断错误的风险。现有的幻觉检测方法通常从响应一致性或视觉证据出发,统一估计VLM输出的可靠性。然而,这种跨问题的统一验证忽略了问题特定的特征,导致过度自信的错误被遗漏,以及因过度验证而产生的误报。我们提出了PROVE(基于证明引导的机制感知算子验证),一种黑盒检测器,它根据每个问题的证据结构调整验证策略。PROVE根据问题所要求的视觉证明类型,将问题分为三种验证机制,激活五个互补算子中的机制特定子集,并通过一个基于确定性问答特征的轻量级校准层,逐问题调整算子重要性。PROVE利用问题特定证据重新加权算子,并生成校准的风险评分。在三个医学VQA基准和四个前沿VLM的8048个测试样本上评估,PROVE实现了0.821的AUROC,比最强基线高出+0.159,在所有模型和基准上均取得一致改进。

英文摘要:

In medical visual question answering (VQA), hallucinations of vision-language models (VLMs) may lead to confident but incorrect responses, raising the risk of diagnostic errors. Existing hallucination detection methods uniformly estimate the reliability of VLM outputs from response consistency or visual evidence. However, such uniform verification across questions ignores question-specific characteristics, resulting in missed overconfident errors and false alarms from over-verification. We present PROVE (Proof-guided Regime-aware Operator Verification), a black-box detector that adapts verification strategy to the evidential structure of each question. PROVE classifies questions into three verification regimes based on what kind of visual proof they demand, activates a regime-specific subset of five complementary operators, and adjusts operator importance per question through a lightweight calibration layer conditioned on deterministic question-answer features. PROVE uses question-specific evidence to reweight operators and produce a calibrated risk score. Evaluated on 8048 test samples across three medical VQA benchmarks and four frontier VLMs, PROVE achieves 0.821 AUROC, outperforming the strongest baseline by +0.159, with consistent gains across all models and benchmarks.

↑