发表机构
SAIL Lab, University of New Haven(萨伊勒实验室,纽黑文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究医学视觉语言模型中视觉解释的因果评估,通过多种方式审核注意力热图忠实度,发现没有评估的VLM满足忠实标准,热图虽视觉上安心但不忠实,临床解释需可控指标和因果扰动而非仅视觉检查。
AI 中文摘要
注意力和显著性热图被广泛用于解释胸部X光片上医学视觉语言模型(VLM)的输出,但它们是否真正突出驱动预测的图像证据尚未经过因果测试。我们通过与PadChest上放射科医生的边界框重叠(n = 637)、CheXlocalize上放射科医生掩码内的归因质量(n = 643)以及记录隐藏哪些区域会改变答案的16x16补丁遮挡图来审核忠实度。我们研究了三个MedGemma - 4B变体、LLaVA - RAD和Qwen3 - VL - 8B - Instruct上的跨家族探针以及专业的CheXagent - 2 - 3b,使用两个经过CXR训练的分类器(DenseNet121、ResNet50)作为阳性对照。只有当模型使用图像且注意力集中在遮挡会改变预测的区域时,热图才是忠实的。没有评估的VLM满足这两个标准。MedGemma和Qwen3 - VL使用图像,但注意力与补丁遮挡重要性呈反相关(rho < 0,95%自举置信区间低于零)。LLaVA - RAD的注意力呈正相关,但该模型几乎仅依赖文本(99.1%文本一致性,接近零因果质量),所以相关性将两个接近零的信号联系起来。注意力也错过了标注的解剖结构:与真实区域重叠从未超过偏移或随机对照,且没有方法将超过22%的质量置于放射科医生掩码内。两个CXR分类器通过所有指标,表明失败是VLM热图特有的而非评估问题。这些热图在视觉上令人安心但不忠实;临床解释需要可控的定位指标和因果扰动,而非仅靠视觉检查。
英文摘要
Attention and saliency heatmaps are widely used to explain medical Vision-Language Model (VLM) outputs on chest X-rays, yet whether they truly highlight the image evidence driving predictions has not been causally tested. We audit faithfulness via overlap with radiologist bounding boxes on PadChest (n=637), attribution mass within radiologist masks on CheXlocalize (n=643), and 16x16 patch-occlusion maps that record which regions, when hidden, change the answer. We study three MedGemma-4B variants, cross-family probes on LLaVA-RAD and Qwen3-VL-8B-Instruct, and the specialist CheXagent-2-3b, with two CXR-trained classifiers (DenseNet121, ResNet50) as positive controls. A heatmap is faithful only if the model uses the image and attention concentrates on regions whose occlusion alters the prediction. No evaluated VLM meets both criteria. MedGemma and Qwen3-VL use the image, but attention anti-correlates with patch-occlusion importance (rho < 0 with 95% bootstrap CIs below zero). LLaVA-RAD's attention correlates positively, but the model is almost text-only (99.1% text-only agreement, near-zero causal mass), so correlation ties two near-zero signals. Attention also misses annotated anatomy: overlap with true regions never beats shifted or random controls, and no method places more than 22% of its mass inside radiologist masks. The two CXR classifiers pass all metrics, indicating the failure is specific to VLM heatmaps, not the evaluation. These heatmaps are visually reassuring but not faithful; clinical explanations require controlled localization metrics and causal perturbation, not visual inspection alone.
CommentsiMIMIC Workshop 2026