arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

性能与一致性:评估基础模型在Lung-RADS筛查中的表现

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

Benjamin Renoust, Pierre Baudot, Tiffany Foriel, Yousra Haddou, Charles Voyton, Pierre-Henri Siot, Ezequiel Geremia, Danny Francis, Jean-Christophe Brisset, Valérie Bourdès, Sylvain Bodard, Benoit Huet

arXiv 2609.22281首次发表:更新:

发表机构

Median Technologies; The University of Osaka; Université de Paris Cité; AP-HP; Hôpital Universitaire Necker Enfants Malades; Memorial Sloan Kettering Cancer Center; Massachusetts General Hospital; Sorbonne Université(Median Technologies; 大阪大学; 巴黎西岱大学; 巴黎公立医院集团; 内克尔儿童医院; 纪念斯隆凯特琳癌症中心; 马萨诸塞总医院; 索邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估基础模型MedGemma在Lung-RADS筛查中的性能,发现微调后AUC达0.83,虽低于放射科医生平均0.90,但模型输出确定性高,可作为临床决策支持的互补工具。

AI 中文摘要

基础模型最近在广泛的医学影像任务中展示了强大的能力。然而,它们在结构化临床解读场景中的性能仍未得到充分探索。在肺癌筛查中,尽管存在如Lung-RADS等标准化框架,解读变异性仍然存在。在本研究中,我们评估了MedGemma,一个源自Gemini的医学通用基础模型及其针对肺癌检测与诊断进行微调的版本,并与放射科医生在NLST数据集上执行Lung-RADS v2022评估进行比较。十二名放射科医生在多读者设计中独立评估每个病例,从而能够量化读者间变异性。放射科医生实现了平均AUC为0.90,读者间存在显著变异性(范围:0.80-0.94)。原生基础模型实现了AUC为0.70,未能达到临床相关性能。相比之下,微调显著提升了性能至AUC为0.83,使模型处于个体放射科医生性能的较低范围内。这些发现凸显了峰值准确性与预测一致性之间的权衡。与放射科医生不同,在固定条件下,模型产生确定性输出,消除了相同输入下的运行间变异性,这与放射科医生中观察到的读者间变异性形成对比。这支持了微调基础模型作为临床决策支持潜在互补工具的作用,特别是在专业知识有限的场景中。然而,评估是在来自NLST的病例富集队列上进行的,未考虑现实世界患病率或外部验证,限制了直接的临床泛化。

英文摘要

Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as Lung-RADS. In this study, we evaluate MedGemma, a medical general-purpose foundation model derived from Gemini and its fine-tuned version adapted for lung cancer detection and diagnosis, compared against radiologists performing Lung-RADS v2022 assessment on the NLST dataset. Twelve radiologists independently evaluated each case in a multi-reader design, enabling quantification of inter-reader variability. Radiologists achieved a mean AUC of 0.90, with substantial variability across readers (range: 0.80-0.94). The native foundation model achieved an AUC of 0.70, failing to reach clinically relevant performance. In contrast, fine-tuning significantly improved performance to an AUC of 0.83, placing the model within the lower range of individual radiologists performance. These findings highlight a trade-off between peak accuracy and prediction consistency. Unlike radiologists, under fixed conditions, the model produces deterministic outputs, removing inter-run variability under identical inputs, in contrast to inter-reader variability observed among radiologists. This supports the role of fine-tuned foundation models potential complementary tools for clinical decision support, particularly in settings with limited expertise. However, evaluation is performed on a case-enriched cohort from NLST and does not account for real-world prevalence or external validation, limiting direct clinical generalization.

CommentsMICCAI CAPTION 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑