基于视觉-语言模型概念引导提示的可解释少样本视网膜疾病诊断
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
- Monash University(莫纳什大学)
- Centre for Eye Research Australia(澳大利亚眼科研究中心)
- Royal Victorian Eye and Ear Hospital(皇家维多利亚眼耳医院)
- The University of Melbourne(墨尔本大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Centre for Eye and Vision Research(眼与视觉研究中心)
- Indian Institute of Technology Bombay(印度理工学院孟买分校)
- Airdoc-Monash Research Lab(Airdoc-莫纳什研究实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种基于GPT知识库提取可解释概念并融入提示学习训练视觉-语言模型的方法,显著提升了视网膜疾病的少样本与零样本分类性能,同时增强了模型的可解释性。
AI中文摘要:
近期深度学习的进展表明,使用彩色眼底图像对视网膜疾病进行分类具有巨大潜力。然而,现有工作主要完全依赖图像数据,其诊断决策缺乏可解释性,且主要将医疗专业人员作为真实标签标注的标注者。为了填补这一空白,我们实施了两个关键策略:利用GPT模型的知识库提取视网膜疾病的可解释概念,并将这些概念作为语言组件纳入提示学习中,以使用眼底图像及其相关概念训练视觉-语言(VL)模型。我们的方法不仅改善了视网膜疾病分类,还丰富了少样本和零样本检测(新疾病检测),同时提供了基于概念的模型可解释性这一额外优势。我们在两个不同的视网膜眼底图像数据集上进行了广泛评估,结果表明通过我们的概念整合方法,基于VL模型的少样本方法获得了显著的性能提升,在16-shot学习和零样本(新类别)检测中,平均精度均值分别实现了约5.8%和2.7%的平均提升。我们的方法标志着向面向真实世界临床应用的可解释且高效的视网膜疾病识别迈出了关键一步。
英文摘要:
Recent advancements in deep learning have shown significant potential for classifying retinal diseases using color fundus images. However, existing works predominantly rely exclusively on image data, lack interpretability in their diagnostic decisions, and treat medical professionals primarily as annotators for ground truth labeling. To fill this gap, we implement two key strategies: extracting interpretable concepts of retinal diseases using the knowledge base of GPT models and incorporating these concepts as a language component in prompt-learning to train vision-language (VL) models with both fundus images and their associated concepts. Our method not only improves retinal disease classification but also enriches few-shot and zero-shot detection (novel disease detection), while offering the added benefit of concept-based model interpretability. Our extensive evaluation across two diverse retinal fundus image datasets illustrates substantial performance gains in VL-model based few-shot methodologies through our concept integration approach, demonstrating an average improvement of approximately 5.8\% and 2.7\% mean average precision for 16-shot learning and zero-shot (novel class) detection respectively. Our method marks a pivotal step towards interpretable and efficient retinal disease recognition for real-world clinical applications.