arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.17803cs.CVcs.AI

Pillar-0:放射学基础模型的新前沿

Pillar-0: A New Frontier for Radiology Foundation Models

Kumar Krishna Agrawal, Longchao Liu, Long Lian, Michael Nercessian, Natalia Harguindeguy, Yufu Wu, Peter Mikhael, Gigin Lin, Lecia V. Sequist, Florian Fintelman… 展开作者

Kumar Krishna Agrawal, Longchao Liu, Long Lian, Michael Nercessian, Natalia Harguindeguy, Yufu Wu, Peter Mikhael, Gigin Lin, Lecia V. Sequist, Florian Fintelmann, Trevor Darrell, Yutong Bai, Maggie Chung, Adam Yala

更新

AI总结:

Pillar-0通过预训练大量医学影像数据和RATE框架,在放射学任务中实现高性能表现,超越现有模型并扩展至新任务。

AI中文摘要:

放射学在现代医学中扮演着至关重要的角色,然而,随着成像量的增加,工作人员的增长速度远远跟不上。基础模型提供了一条通向协助完成放射学所有任务的途径,但现有的医学模型仍然有限:它们将体积CT和MRI处理为低保真度的2D切片,丢弃了关键的灰度对比信息,并缺乏能够反映真实临床实践的评估框架。我们引入了Pillar-0,这是一个预训练于42,990例腹部-盆腔CT、86,411例胸部CT、14,348例头部CT和11,543例乳腺MRI的放射学基础模型,这些数据来自一个大型学术中心,同时还有RATE,一个可扩展的框架,利用大语言模型提取366种放射学发现的结构化标签,其准确率接近完美。在内部测试集上,Pillar-0在14,230例腹部-盆腔CT、10,646例胸部CT、4,906例头部CT和1,585例乳腺MRI上建立了新的性能前沿,实现了86.4、88.0、90.1和82.9的平均AUROCs,比MedGemma(Google)、MedImageInsight(Microsoft)、Lingshu(Alibaba)和Merlin(Stanford)分别高出7.8-15.8 AUROC点,并在87.2%(319/366)的任务中排名第一。Pillar-0在斯坦福腹部CT数据集上的外部验证中也优于所有基线,包括Merlin(82.2 vs 80.6 AUROC)。Pillar-0还扩展到其预训练之外的任务,如长期肺癌风险预测,在NLST上比最先进的Sybil提高了3.0 C-index点,在MGH和CGMH上分别获得了5.9和1.9的提升。在脑出血检测中,Pillar-0在使用仅1/20的数据量时获得了超过95的AUROC。Pillar-0和RATE一起为构建高性能放射学系统提供了开放且临床严谨的基础,使之前由于计算、数据和评估限制而不可行的应用成为可能。

英文摘要:

Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full spectrum of radiology tasks, but existing medical models remain limited: they process volumetric CT and MRI as low-fidelity 2D slices, discard critical grayscale contrast information, and lack evaluation frameworks that reflect real clinical practice. We introduce Pillar-0, a radiology foundation model pretrained on 42,990 abdomen-pelvis CTs, 86,411 chest CTs, 14,348 head CTs, and 11,543 breast MRIs from a large academic center, together with RATE, a scalable framework that extracts structured labels for 366 radiologic findings with near-perfect accuracy using LLMs. Across internal test sets of 14,230 abdomen-pelvis CTs, 10,646 chest CTs, 4,906 head CTs, and 1,585 breast MRIs, Pillar-0 establishes a new performance frontier, achieving mean AUROCs of 86.4, 88.0, 90.1, and 82.9, outperforming MedGemma (Google), MedImageInsight (Microsoft), Lingshu (Alibaba), and Merlin (Stanford) by 7.8-15.8 AUROC points and ranking best in 87.2\% (319/366) tasks. Pillar-0 similarly outperforms all baselines in an external validation on the Stanford Abdominal CT dataset, including Merlin (82.2 vs 80.6 AUROC). Pillar-0 extends to tasks beyond its pretraining, such as long-horizon lung cancer risk prediction, where it improves upon the state-of-the-art Sybil by 3.0 C-index points on NLST, and generalizes with gains of 5.9 (MGH) and 1.9 (CGMH). In brain hemorrhage detection, Pillar-0 obtained a >95 AUROC when using only 1/20th of the data of the next most sample efficient baseline. Pillar-0 and RATE together provide an open, clinically rigorous foundation for building high-performance radiology systems, enabling applications that were previously infeasible due to computational, data, and evaluation constraints.

↑