专家引导的可解释少样本学习用于医学图像诊断
Expert-Guided Explainable Few-Shot Learning for Medical Image Diagnosis
- AI Research Lab, Department of Computer Science, University of South Dakota, Vermillion, SD 57069, USA(人工智能研究实验室,计算机科学系,南达科他大学,弗米利恩,SD 57069,美国)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对医学图像少样本学习中性能与可解释性难以兼顾的问题,提出融合放射科医生ROI和Grad-CAM的专家引导框架,通过Dice解释损失联合优化原型网络,在BraTS和VinDr-CXR上分别将准确率提升至83.61%和73.29%。
AI中文摘要:
医学图像分析常常因专家标注数据有限而面临重大挑战,这阻碍了模型的泛化能力和临床采用。我们提出了一种专家引导的可解释少样本学习框架,将放射科医生提供的感兴趣区域(ROIs)整合到模型训练中,以同时提升分类性能和可解释性。利用Grad-CAM进行空间注意力监督,我们引入了一种基于Dice相似度的解释损失,在训练期间将模型注意力与诊断相关区域对齐。该解释损失与标准原型网络目标联合优化,鼓励模型即使在数据有限的条件下也能聚焦于临床有意义的特征。我们在两个不同的数据集上评估了我们的框架:BraTS(MRI)和VinDr-CXR(胸部X射线),与未引导模型相比,在BraTS上准确率从77.09%显著提升至83.61%,在VinDr-CXR上从54.33%提升至73.29%。Grad-CAM可视化进一步证实,专家引导的训练始终将注意力与诊断区域对齐,提高了预测可靠性和临床可信度。我们的研究结果表明,整合专家引导的注意力监督能有效弥合少样本医学图像诊断中性能与可解释性之间的差距。
英文摘要:
Medical image analysis often faces significant challenges due to limited expert-annotated data, hindering both model generalization and clinical adoption. We propose an expert-guided explainable few-shot learning framework that integrates radiologist-provided regions of interest (ROIs) into model training to simultaneously enhance classification performance and interpretability. Leveraging Grad-CAM for spatial attention supervision, we introduce an explanation loss based on Dice similarity to align model attention with diagnostically relevant regions during training. This explanation loss is jointly optimized with a standard prototypical network objective, encouraging the model to focus on clinically meaningful features even under limited data conditions. We evaluate our framework on two distinct datasets: BraTS (MRI) and VinDr-CXR (Chest X-ray), achieving significant accuracy improvements from 77.09% to 83.61% on BraTS and from 54.33% to 73.29% on VinDr-CXR compared to non-guided models. Grad-CAM visualizations further confirm that expert-guided training consistently aligns attention with diagnostic regions, improving both predictive reliability and clinical trustworthiness. Our findings demonstrate the effectiveness of incorporating expert-guided attention supervision to bridge the gap between performance and interpretability in few-shot medical image diagnosis.