发表机构
German Cancer Research Center (DKFZ); Helmholtz Imaging, German Cancer Research Center (DKFZ); Helmholtz Association; Heidelberg University(德国癌症研究中心; 亥姆霍兹影像,德国癌症研究中心; 亥姆霍兹联合会; 海德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
nnFoundation提出互补的卷积与Transformer3D放射学基础模型,在210万医学图像体上训练,经108项任务评估,优于现有方法,并强调架构与数据适配的协同作用。
AI 中文摘要
放射学人工智能已取得快速发展,但大多数系统仍然局限于特定任务、数据密集且在域偏移下表现脆弱。基础模型有望提供更具可迁移性和数据效率的解决方案,但现有方法在规模上有限、评估范围狭窄,且通常假设单个预训练模型能够支持多种下游任务。在此,我们提出nnFoundation,即互补的基于卷积和Transformer的3D放射学基础模型。在人类放射组计划(THRP)框架内开发,nnFoundation在来自125个机构和公共数据集的210万CT、MRI和PET图像体上进行了训练。我们在108项任务中对其进行评估,涵盖分割、检测、分类、报告生成和图像检索,包括在域偏移、外部合作伙伴评估以及低数据和低计算资源场景下的评估。在所有任务类型中,我们基于卷积和Transformer的nnFoundation模型均持续优于先前3D基础模型和从头训练,确立了放射学成像的最新性能。然而,性能遵循一致的任务依赖结构:卷积nnFoundation模型在空间局部化任务中占优,而基于Transformer的nnFoundation模型在需要全局语义推理和冻结特征设置中表现出色。事后将基础模型拓扑与数据集特征动态对齐,进一步改善了跨异构3D设置的迁移性能。这些结果表明,可迁移的3D放射学性能并非由单一通用模型决定,而是由可扩展预训练、互补架构和数据感知适应的相互作用所主导。我们发布了集成到nnU-Net和nnDetection中的nnFoundation模型,使其能够立即应用于既定的放射学工作流程。
英文摘要
Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.