GPT-4V(ision)能否服务于医疗应用?GPT-4V用于多模态医学诊断的案例研究
Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统评估了GPT-4V在17个人体系统、8种医学图像模态上的多模态诊断能力,发现其虽能识别模态和解剖结构,但在疾病诊断与报告生成上仍存在显著不足,尚不能有效支持真实临床决策。
AI中文摘要:
在大型基础模型的推动下,人工智能的发展近来取得了巨大进步,引发了公众的广泛关注。本研究旨在评估OpenAI最新模型GPT-4V(ision)在多模态医学诊断领域的表现。我们的评估涵盖17个人体系统,包括中枢神经系统、头颈部、心脏、胸部、血液学、肝胆、胃肠、泌尿生殖、妇科、产科、乳腺、肌肉骨骼、脊柱、血管、肿瘤、创伤、儿科,图像来自日常临床常规中使用的8种模态,例如X射线、计算机断层扫描(CT)、磁共振成像(MRI)、正电子发射断层扫描(PET)、数字减影血管造影(DSA)、乳腺X线摄影、超声和病理学。我们在提供或不提供患者病史的情况下,探查GPT-4V在多项临床任务上的能力,包括成像模态和解剖结构识别、疾病诊断、报告生成和疾病定位。我们的观察表明,尽管GPT-4V在区分医学图像模态和解剖结构方面表现出熟练能力,但它在疾病诊断和生成全面报告方面面临重大挑战。这些发现强调,尽管大型多模态模型在计算机视觉和自然语言处理方面取得了显著进展,但它距离被有效用于支持真实世界的医疗应用和临床决策仍有很大差距。本报告中使用的所有图像可在https://github.com/chaoyi-wu/GPT-4V_Medical_Evaluation中找到。
英文摘要:
Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's newest model, GPT-4V(ision), specifically in the realm of multimodal medical diagnosis. Our evaluation encompasses 17 human body systems, including Central Nervous System, Head and Neck, Cardiac, Chest, Hematology, Hepatobiliary, Gastrointestinal, Urogenital, Gynecology, Obstetrics, Breast, Musculoskeletal, Spine, Vascular, Oncology, Trauma, Pediatrics, with images taken from 8 modalities used in daily clinic routine, e.g., X-ray, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Digital Subtraction Angiography (DSA), Mammography, Ultrasound, and Pathology. We probe the GPT-4V's ability on multiple clinical tasks with or without patent history provided, including imaging modality and anatomy recognition, disease diagnosis, report generation, disease localisation. Our observation shows that, while GPT-4V demonstrates proficiency in distinguishing between medical image modalities and anatomy, it faces significant challenges in disease diagnosis and generating comprehensive reports. These findings underscore that while large multimodal models have made significant advancements in computer vision and natural language processing, it remains far from being used to effectively support real-world medical applications and clinical decision-making. All images used in this report can be found in https://github.com/chaoyi-wu/GPT-4V_Medical_Evaluation.