从密度到活检决策与恶性预测:多模态大语言模型与放射科医生在数字乳腺摄影和对比增强乳腺摄影中的基准研究
From Density to Biopsy Decisions and Malignancy Prediction: A Benchmark Study of Multimodal Large Language Models Against Radiologists in Digital and Contrast-Enhanced Mammography
浏览论文内容
中文总结 AI 辅助
本研究对比四种多模态大语言模型与放射科医生在乳腺密度、BI-RADS分类、活检决策及恶性概率估计上的表现,发现放射科医生在分类任务中更优,但带病灶掩膜的模型在连续恶性概率估计上接近人类水平,提示其可作为辅助工具。
中文摘要 AI 辅助
目的:比较四种多模态大语言模型(MLLMs)与不同专业水平的放射科医生在乳腺密度评估、BI-RADS评估、活检候选资格判定以及连续恶性概率估计方面的表现,使用数字乳腺摄影(DM)和对比增强乳腺摄影(CEM)。方法:本研究纳入179名女性,均接受配对DM/CEM检查并有参考标准。四种MLLMs(ChatGPT-5.2、Gemini-3.1 Pro、Sonnet-4.6、Muse Spark)在有无掩膜的情况下解读图像;三名放射科医生解读无掩膜图像。结果:在DM上的二分类密度分类中,放射科医生的准确率范围为55.81%至78.60%,超过大多数MLLM的值(62.33%-71.63%),而掩膜带来的益处有限。在DM(56.74-67.44%)和CEM(62.33-82.79%)上,放射科医生的五类BI-RADS准确率高于MLLMs(DM 31.16-45.12%;CEM 40.00-55.81%)。二分类活检候选资格准确率同样放射科医生更高(DM 85.12-89.77%;CEM 86.98-92.09%)优于MLLMs(DM 61.39-75.35%;CEM 69.30-82.79%),尽管CEM在所有读者中均提高了性能。病灶掩膜显著提高了MLLMs的连续恶性概率准确率,在DM上从64.65%-71.63%提升至72.56-78.60%,在CEM上从67.91%-77.21%提升至72.56%-81.86%,接近放射科医生的范围(DM 63.72-82.79%;CEM 81.86-88.84%)。相应的最佳掩膜模型的AUC与人类读者重叠。总体而言,Muse Spark表现最强,其次是Sonnet-4.6,在跨领域中MLLMs中表现最佳。结论:放射科医生在分类任务中普遍优于MLLMs,而选定的掩膜模型在连续恶性概率估计方面接近人类表现,提示其可能作为辅助角色。
英文摘要
Purpose: To compare four multimodal large language models (MLLMs) with radiologists of varying expertise in breast density assessment, BI-RADS assessment, biopsy candidacy determination, and continuous malignancy probability estimation using digital mammography (DM) and contrast-enhanced mammography (CEM). Methods: This study included 179 women with paired DM/CEM examinations and reference standards. Four MLLMs (ChatGPT-5.2, Gemini-3.1 Pro, Sonnet-4.6, Muse Spark) interpreted images with and without masks; three radiologists interpreted non-masked images. Results: For binary density classification on DM, radiologist accuracies ranged from 55.81% to 78.60%, exceeding most MLLM values (62.33%-71.63%), while masks added limited benefit. Five-category BI-RADS accuracies were higher for radiologists on DM (56.74-67.44%) and CEM (62.33-82.79%) compared with MLLMs (DM 31.16-45.12%; CEM 40.00-55.81%). Binary biopsy-candidacy accuracies were likewise higher for radiologists (DM 85.12-89.77%; CEM 86.98-92.09%) than for MLLMs (DM 61.39-75.35%; CEM 69.30-82.79%), although CEM improved performance across all readers. Lesion masks substantially improved MLLM continuous malignancy-probability accuracies from 64.65%-71.63% to 72.56-78.60% on DM and from 67.91%-77.21% to 72.56%-81.86% on CEM, approaching radiologist ranges (DM 63.72-82.79%; CEM 81.86-88.84%). The corresponding AUCs for the top masked models overlapped those of the human readers. Overall, Muse Spark, followed by Sonnet-4.6, demonstrated the strongest performance among the MLLMs across domains. Conclusion: Radiologists generally outperformed MLLMs in categorical tasks, while selected masked models approached human performance for continuous malignancy probability estimation, suggesting a potential adjunctive role.
发表机构
- Danube Private University(多瑙河私立大学)
- Urmia University of Medical Science(乌尔米耶医科大学)
- Yeditepe University(耶迪特佩大学)
- Kartal Dr. LÜtfi Kırdar City Hospital(卡尔塔尔·卢特菲·克尔达尔博士市立医院)
- Iran University of Medical Sciences(伊朗医科大学)
- University of Padova(帕多瓦大学)
- Bahçeşehir University(巴赫切谢希尔大学)
- Mooney's Bay Pain Clinic(穆尼湾疼痛诊所)
- Arden University(雅顿大学)
- University Hospital Cologne(科隆大学医院)
- University of Southern Queensland(南昆士兰大学)
- Austrian Center for Medical Innovation and Technology(奥地利医学创新与技术中心)
机构由 AI 辅助整理,请以论文原文为准。