发表机构
Pattern Recognition Revisited Lab (PURRlab); IT University of Copenhagen(模式识别重访实验室(PURRlab); 哥本哈根信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究MedCLIP模型中不同层的真实世界胸部X光捷径,发现局部捷径出现在较后层、弥散捷径出现在较早层,且两个数据集存在数据质量问题,凸显高质量数据集的必要性。
AI 中文摘要
基于对比语言-图像预训练(CLIP)的视觉语言模型在医学人工智能领域已达到了当前最优(SOTA)结果,但近期研究表明,这类模型仍易受捷径(shortcut)的影响。我们探究了真实世界的捷径如何在基于CLIP的医学模型MedCLIP及其视觉编码器(冻结的ResNet-50)的不同层中显现。我们在ResNet-50的中间层附加了17个线性分类探针,在三种不同的数据集配置和目标上进行训练:NIH-CXR14(气胸)和PadChest(心脏肥大和气胸)。该设置让我们能使用基于子组的校准和逐层置信度曲线观察模型在评估期间的行为。我们发现,最终的线性探针在模型中实现了高AUROC,但校准效果较差。逐层置信度分析表明,捷径在不同深度出现:与局部捷径(如引流管)一致的模式出现在较后层,而与弥散捷径(如扫描仪特有的噪声模式)一致的模式出现在较早层,这与之前的研究结果一致。最后,我们对图像进行了手动分析,发现NIH-CXR14和PadChest均存在数据质量问题。我们的研究结果强调,即使是SOTA模型仍易受捷径影响,因此需要高质量、标注完善的数据集以得出可靠结论。代码可在我们的GitHub上获取:this https URL。
英文摘要
Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder, a frozen ResNet-50. We attach 17 linear classification probes to the intermediate layers of the ResNet-50 and train them on three different dataset configurations and targets: NIH-CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows us to observe model behaviour during evaluation using subgroup-based calibration and layer-wise confidence curves. We find that the final linear probes achieve a high AUROC but poor calibration in the models. The layer-wise confidence analyses suggest that shortcuts emerge at different depths. Patterns consistent with localised shortcuts, such as drains, appear at later layers, while patterns consistent with diffuse shortcuts, such as scanner-specific noise patterns, emerge earlier, aligning with previous work. Finally, we conduct a manual analysis of the images, which reveals data quality issues in both NIH-CXR14 and PadChest. Our findings underscore that even SOTA models remain vulnerable to shortcuts, and the need for high-quality and well-annotated datasets to draw solid conclusions. Code can be found on our GitHub: https://github.com/nikodice4/MedCLIP_shortcuts.
Comments11 pages, 3 figure, poster presentation at the joint FAIMI, BRIDGE, and EPIMI workshop at MICCAI 2026 (Strasbourg, France) conference