发表机构
Johns Hopkins University; University Hospital Basel; Stanford University; University of California, San Francisco; Italian Institute of Technology; University of Bologna; École Polytechnique Fédérale de Lausanne; Johns Hopkins Medicine(约翰斯·霍普金斯大学; 巴塞尔大学医院; 斯坦福大学; 加利福尼亚大学旧金山分校; 意大利技术研究院; 博洛尼亚大学; 洛桑联邦理工学院; 约翰斯·霍普金斯医疗集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出报告监督(R-Super)框架,利用大量放射学报告补充稀缺的肿瘤掩码训练数据,在肾脏和胰腺肿瘤分割任务上显著提升了AI的检测与分割性能。
AI 中文摘要
分割模型在肿瘤检测中可超越放射科医生、分类模型及视觉-语言模型,重要的是,分割模型能勾勒出肿瘤轮廓,方便放射科医生更好地验证和信任AI输出。其主要局限在于肿瘤掩码的稀缺性:创建一个3D肿瘤掩码最长需30分钟,因此大多数公开CT数据集仅包含数百个掩码,即使是最大的私有数据集也仅包含数千个。肿瘤掩码并非临床常规产出,但放射学报告是常规产出的,公开数据集包含数万份CT-报告对,医院则拥有数十万份。这些报告详细描述了肿瘤,提供了大规模、信息丰富的训练数据。在此,我们提出报告监督(R-Super),一种利用报告直接监督和改进肿瘤分割的训练框架。R-Super引入了新的损失函数,指导分割模型分割出与报告中肿瘤数量、大小和位置描述匹配的肿瘤,报告仅用于训练。我们在肾脏和胰腺肿瘤分割任务上评估了R-Super,探索了不同的训练数据规模,最多使用了41418份CT-报告对加上3488个胰腺肿瘤CT-掩码对。在外部验证中,与仅使用掩码训练相比,R-Super使肿瘤检测F1分数和分割DSC最多提升了15%,还超越了CLIP、多任务学习等替代方法。利用大量现成的报告补充稀缺的掩码,R-Super在训练掩码极少(如50个)和掩码充足(如3488个)时均能显著提升AI性能,实现了肿瘤分割的规模化。
英文摘要
Segmentation models can surpass radiologists, classification models, and vision-language models in tumor detection. Importantly, segmentation models outline tumors, allowing radiologists to better verify and trust the AI output. Their main limitation is the scarcity of tumor masks: creating one 3D tumor mask takes up to 30 minutes, so most public CT datasets contain only a few hundred masks, and even the largest private datasets contain only a couple of thousand. Tumor masks are not produced in clinical routine, but radiology reports are. Public datasets contain tens of thousands of CT-Report pairs, and hospitals contain hundreds of thousands. These reports describe tumors in detail, providing large-scale, informative training data. Here, we introduce Report Supervision (R-Super), a training framework that uses reports to directly supervise and improve tumor segmentation. R-Super introduces new loss functions that teach segmentation models to segment tumors that match report descriptions of tumor count, sizes, and locations. Reports are only used for training. We evaluated R-Super on kidney and pancreatic tumor segmentation, exploring diverse training data sizes, up to 41,418 CT-Report plus 3,488 pancreatic tumor CT-Mask pairs. On external validation, R-Super increased tumor detection F1-Score and segmentation DSC by up to +15% with respect to mask-only training. It also surpassed alternative methods such as CLIP and multi-task learning. Leveraging numerous readily available reports to supplement scarce masks, R-Super strongly improves AI performance when very few training masks are available (e.g., 50), and when many masks are available (e.g., 3,488), unlocking scale in tumor segmentation.
CommentsPublished in Medical Image Analysis, 2026