arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2406.06512cs.CVcs.AI

Merlin:一种计算断层扫描视觉-语言基础模型和数据集

Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset

  • Stanford University(斯坦福大学)
  • Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心)
  • Department of Electrical Engineering(电气工程系)
  • University of California Berkeley(加州大学伯克利分校)
  • Department of Radiology(放射科)
  • University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
  • Hospital Israelita Albert Einstein(以色列特医院阿尔伯特·爱因斯坦医院)

机构由 AI 辅助整理,请以论文原文为准。

Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen, Jiaming Liu, Longchao Liu, Dave Van Veen, Syed Jamal Safdar Gardezi, Hongkun Yu, Magdalini Paschali, Zhihong… 展开作者

Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen, Jiaming Liu, Longchao Liu, Dave Van Veen, Syed Jamal Safdar Gardezi, Hongkun Yu, Magdalini Paschali, Zhihong Chen, Jean-Benoit Delbrouck, Eduardo Reis, Robbie Holland, Cesar Truyts, Christian Bluethgen, Yufu Wu, Long Lian, Malte Engmann Kjeldskov Jensen, Sophie Ostmeier, Maya Varma, Jeya Maria Jose Valanarasu, Zhongnan Fang, Zepeng Huo, Zaid Nabulsi, Diego Ardila, Wei-Hung Weng, Edson Amaro Junior, Neera Ahuja, Jason Fries, Nigam H. Shah, Greg Zaharchuk, Marc Willis, Adam Yala, Andrew Johnston, Robert D. Boutin, Andrew Wentland, Curtis P. Langlotz, Jason Hom, Sergios Gatidis, Akshay S. Chaudhari

更新

中文总结 AI 辅助

Merlin是一种基于3D视觉-语言模型的医学影像分析工具,通过多阶段预训练框架,实现了对腹部CT扫描、电子健康记录和放射科报告的综合学习,提升了医学影像分析的自动化水平。

中文摘要 AI 辅助

大量腹部计算断层扫描(CT)扫描与放射科医生短缺加剧了对自动化医学图像分析工具的需求。以往最先进的自动化分析方法利用视觉-语言模型(VLMs)同时建模图像和放射科报告。然而,目前的医学VLMs通常仅限于2D图像和短报告。为克服这些不足以进行腹部CT解读,我们引入了Merlin,一种学习自体积CT扫描、电子健康记录数据和放射科报告的3D VLM。这一方法得益于一个多阶段预训练框架,无需额外的手动注释。我们使用高质量的临床数据集对Merlin进行训练,该数据集包含配对的CT扫描(>600万张图像来自15,331次CT扫描)、诊断代码(>180万代码)和放射科报告(>600万tokens)。我们对6种任务类型和752个个体任务进行了全面评估,涵盖了诊断、预后和质量相关任务。非适应任务包括零样本发现分类(30种发现)、表型分类(692种表型)和零样本跨模态检索(图像到发现和图像到印象)。模型适应任务包括5年慢性疾病预测(6种疾病)、放射科报告生成和3D语义分割(20个器官)。我们对Merlin进行了大规模验证,内部测试在5,137次CT扫描上进行,外部测试在3个独立站点和2个公开数据集上的44,098次CT扫描上进行。结果表明,Merlin在不同机构和解剖结构上具有高度泛化能力。Merlin在2D VLMs、CT基础模型和现成放射科模型上均表现更优。我们还发布了训练好的模型、代码和数据集,可在:https://github.com/StanfordMIMI/Merlin上获取。

英文摘要

The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previous state-of-the-art approaches for automated analysis leverage vision-language models (VLMs) that jointly model images and radiology reports. However, current medical VLMs are generally limited to 2D images and short reports. Here to overcome these shortcomings for abdominal CT interpretation, we introduce Merlin, a 3D VLM that learns from volumetric CT scans, electronic health record data and radiology reports. This approach is enabled by a multistage pretraining framework that does not require additional manual annotations. We trained Merlin using a high-quality clinical dataset of paired CT scans (>6 million images from 15,331 CT scans), diagnosis codes (>1.8 million codes) and radiology reports (>6 million tokens). We comprehensively evaluated Merlin on 6 task types and 752 individual tasks that covered diagnostic, prognostic and quality-related tasks. The non-adapted (off-the-shelf) tasks included zero-shot classification of findings (30 findings), phenotype classification (692 phenotypes) and zero-shot cross-modal retrieval (image-to-findings and image-to-impression). The model-adapted tasks included 5-year chronic disease prediction (6 diseases), radiology report generation and 3D semantic segmentation (20 organs). We validated Merlin at scale, with internal testing on 5,137 CT scans and external testing on 44,098 CT scans from 3 independent sites and 2 public datasets. The results demonstrated high generalization across institutions and anatomies. Merlin outperformed 2D VLMs, CT foundation models and off-the-shelf radiology models. We also release our trained models, code, and dataset, available at: https://github.com/StanfordMIMI/Merlin.

补充信息

↑