Pillar-0:放射学基础模型的新前沿
Pillar-0: A New Frontier for Radiology Foundation Models
AI总结:
Pillar-0通过预训练大量医学影像数据和RATE框架,在放射学任务中实现高性能表现,超越现有模型并扩展至新任务。
AI中文摘要:
放射学在现代医学中扮演着至关重要的角色,然而,随着成像量的增加,工作人员的增长速度远远跟不上。基础模型提供了一条通向协助完成放射学所有任务的途径,但现有的医学模型仍然有限:它们将体积CT和MRI处理为低保真度的2D切片,丢弃了关键的灰度对比信息,并缺乏能够反映真实临床实践的评估框架。我们引入了Pillar-0,这是一个预训练于42,990例腹部-盆腔CT、86,411例胸部CT、14,348例头部CT和11,543例乳腺MRI的放射学基础模型,这些数据来自一个大型学术中心,同时还有RATE,一个可扩展的框架,利用大语言模型提取366种放射学发现的结构化标签,其准确率接近完美。在内部测试集上,Pillar-0在14,230例腹部-盆腔CT、10,646例胸部CT、4,906例头部CT和1,585例乳腺MRI上建立了新的性能前沿,实现了86.4、88.0、90.1和82.9的平均AUROCs,比MedGemma(Google)、MedImageInsight(Microsoft)、Lingshu(Alibaba)和Merlin(Stanford)分别高出7.8-15.8 AUROC点,并在87.2%(319/366)的任务中排名第一。Pillar-0在斯坦福腹部CT数据集上的外部验证中也优于所有基线,包括Merlin(82.2 vs 80.6 AUROC)。Pillar-0还扩展到其预训练之外的任务,如长期肺癌风险预测,在NLST上比最先进的Sybil提高了3.0 C-index点,在MGH和CGMH上分别获得了5.9和1.9的提升。在脑出血检测中,Pillar-0在使用仅1/20的数据量时获得了超过95的AUROC。Pillar-0和RATE一起为构建高性能放射学系统提供了开放且临床严谨的基础,使之前由于计算、数据和评估限制而不可行的应用成为可能。
英文摘要:
Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full spectrum of radiology tasks, but existing medical models remain limited: they process volumetric CT and MRI as low-fidelity 2D slices, discard critical grayscale contrast information, and lack evaluation frameworks that reflect real clinical practice. We introduce Pillar-0, a radiology foundation model pretrained on 42,990 abdomen-pelvis CTs, 86,411 chest CTs, 14,348 head CTs, and 11,543 breast MRIs from a large academic center, together with RATE, a scalable framework that extracts structured labels for 366 radiologic findings with near-perfect accuracy using LLMs. Across internal test sets of 14,230 abdomen-pelvis CTs, 10,646 chest CTs, 4,906 head CTs, and 1,585 breast MRIs, Pillar-0 establishes a new performance frontier, achieving mean AUROCs of 86.4, 88.0, 90.1, and 82.9, outperforming MedGemma (Google), MedImageInsight (Microsoft), Lingshu (Alibaba), and Merlin (Stanford) by 7.8-15.8 AUROC points and ranking best in 87.2\% (319/366) tasks. Pillar-0 similarly outperforms all baselines in an external validation on the Stanford Abdominal CT dataset, including Merlin (82.2 vs 80.6 AUROC). Pillar-0 extends to tasks beyond its pretraining, such as long-horizon lung cancer risk prediction, where it improves upon the state-of-the-art Sybil by 3.0 C-index points on NLST, and generalizes with gains of 5.9 (MGH) and 1.9 (CGMH). In brain hemorrhage detection, Pillar-0 obtained a >95 AUROC when using only 1/20th of the data of the next most sample efficient baseline. Pillar-0 and RATE together provide an open, clinically rigorous foundation for building high-performance radiology systems, enabling applications that were previously infeasible due to computational, data, and evaluation constraints.