arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于医学视觉语言模型无注释幻觉缓解的反事实解剖学引导时空解码

Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

Yifan Lu, Adinath Dukre, Abhijit Das, Ziyun Zou, Haolin Yang, Yutong Xie, Imran Razzak

arXiv 2608.17427首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; MedOS(穆罕默德·本·扎耶德人工智能大学; MedOS)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出无需人工注释的CAST框架,通过自动选择解剖区域结合对比解码,在SLAKE和MIMIC-CXR数据集上针对三种Med-VLMs实现优于基线的幻觉缓解,提升了医学视觉语言模型的空间接地性。

AI 中文摘要

医学视觉语言模型(Med-VLMs)在医学视觉问答任务中表现出强大性能,但仍易产生幻觉,生成缺乏临床依据且未充分基于图像证据的陈述。解码阶段应用的缓解方法是实用解决方案,但通常缺乏解剖学意识或严重依赖真实标注,限制了其适用性。我们提出反事实解剖学引导时空解码(CAST)框架,完全在推理阶段运行,无需人工注释即可实现基于解剖学的幻觉缓解。CAST通过广泛的医学分割自动发现与给定查询相关的解剖区域,随后基于遮挡下答案似然的下降,通过反事实干预选择紧凑且具有因果信息的区域。在所选区域的引导下,CAST执行统一的对比解码过程,结合无分类器引导以校正空间注意力,以及逐步时间对比以调节生成动态。在SLAKE和MIMIC-CXR数据集上针对三种Med-VLMs的实验表明,CAST始终优于强基线,且超过依赖真实标注的解码策略。我们的结果表明,自动选择的紧凑区域无需专家注释即可提供高效的对比引导,为改善空间接地并减少幻觉提供了实用且可推广的解决方案。代码可在指定URL获取。

英文摘要

Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.

CommentsAccepted by MICCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑