arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2307.01981eess.IVcs.CVcs.LG

ChatGPT辅助的可解释零样本医学图像诊断框架

A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis

  • Zhejiang University-University of Illinois at Urbana-Champaign Institute(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院)
  • Zhejiang University(浙江大学)
  • National University of Singapore(新加坡国立大学)
  • Angelalign Inc.(时代天使公司)

机构由 AI 辅助整理,请以论文原文为准。

Jiaxiang Liu, Tianxiang Hu, Yan Zhang, Xiaotang Gai, Yang Feng, Zuozhu Liu

更新

AI总结:

本文提出一种结合CLIP与ChatGPT的零样本医学图像诊断框架,通过LLM生成疾病症状等额外知识提升分类准确性与可解释性,并在多个数据集上验证了其有效性。

AI中文摘要:

零样本医学图像分类在现实场景中是一个关键过程,因为在这些场景中我们无法获得所有可能的疾病或大规模标注数据。该过程涉及计算查询医学图像与可能疾病类别之间的相似度分数,以确定诊断结果。近年来,以CLIP为代表的预训练视觉-语言模型(VLMs)在零样本自然图像识别方面表现出色,并在医学应用中展现出优势。然而,一个性能优异且可解释的零样本医学图像识别框架仍待开发。本文提出了一种新颖的基于CLIP的零样本医学图像分类框架,并辅以ChatGPT实现可解释诊断,模拟人类专家的诊断过程。其核心思想是利用类别名称查询大语言模型(LLMs),自动生成额外的线索和知识,例如疾病症状或除单一类别名称之外的描述,以帮助CLIP提供更准确且可解释的诊断。我们进一步设计了特定的提示,以提高ChatGPT生成的描述视觉医学特征的文本质量。在一个私有数据集和四个公开数据集上的大量实验结果以及详细分析证明了我们无需训练(免训练)的零样本诊断流程的有效性和可解释性,印证了VLMs和LLMs在医学应用中的巨大潜力。

英文摘要:

Zero-shot medical image classification is a critical process in real-world scenarios where we have limited access to all possible diseases or large-scale annotated data. It involves computing similarity scores between a query medical image and possible disease categories to determine the diagnostic result. Recent advances in pretrained vision-language models (VLMs) such as CLIP have shown great performance for zero-shot natural image recognition and exhibit benefits in medical applications. However, an explainable zero-shot medical image recognition framework with promising performance is yet under development. In this paper, we propose a novel CLIP-based zero-shot medical image classification framework supplemented with ChatGPT for explainable diagnosis, mimicking the diagnostic process performed by human experts. The key idea is to query large language models (LLMs) with category names to automatically generate additional cues and knowledge, such as disease symptoms or descriptions other than a single category name, to help provide more accurate and explainable diagnosis in CLIP. We further design specific prompts to enhance the quality of generated texts by ChatGPT that describe visual medical features. Extensive results on one private dataset and four public datasets along with detailed analysis demonstrate the effectiveness and explainability of our training-free zero-shot diagnosis pipeline, corroborating the great potential of VLMs and LLMs for medical applications.

补充信息

↑