发表机构
Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WILSON将多张全切片图像编码为多放大倍数复合图像,以病理报告为监督训练基础模型,实现患者级分析、诊断文本检索与生成,性能优于现有模型且计算成本更低。
AI 中文摘要
病理学家在诊断过程中会整合不同放大倍数下的形态学特征以及患者病例中多张切片的整体信息,而现有的病理学基础模型通常仅对单张切片中的数千个图像块进行编码并聚合其特征。为此,我们提出了WILSON,一个视觉-语言基础模型,它将全切片图像和多切片病例表示为单一的多放大倍数复合图像,并使用梅奥诊所约18.9万张切片(涵盖42个器官和829种诊断实体)的病理报告作为监督信号进行训练。在无需任务特定训练的情况下,WILSON在所有的内部队列中均超过了专门的病例级模型(宏F1分数0.52对0.38),并以272至2155倍更低的计算成本达到了比其大9.4倍的切片级模型的性能。在508例三阴性乳腺癌病例上进行端到端微调后,组织学亚型分类和间质肿瘤浸润淋巴细胞分级分别提升了0.16和0.11的宏F1分数。WILSON在检索匹配的诊断文本时达到了75.6%的召回率@1(PRISM为58.1%),并且在内部队列及大多数外部比较中,生成的文本描述比PRISM和PRISM2更接近报告派生的参考文本。因此,复合图像为病理学提供了一种紧凑且临床对齐的计算单元。
英文摘要
Pathologists integrate morphology across magnifications and across the slides of a patient case, whereas pathology foundation models encode thousands of tiles from single slides and aggregate their features. Here we present WILSON, a vision--language foundation model that represents whole-slide images and multi-slide cases as single multi-magnification composite images, trained on approximately 189k slides from Mayo Clinic spanning 42 organs and 829 diagnostic entities using pathology reports as supervision. Without task-specific training, WILSON exceeded a dedicated case-level model on all internal cohorts (macro-F1 0.52 versus 0.38) and matched slide-level models up to 9.4 times larger at 272- to 2,155-fold lower compute. End-to-end fine-tuning on 508 triple-negative breast cancer cases improved histologic subtyping and stromal tumor-infiltrating lymphocyte grading by 0.16 and 0.11 macro-F1. WILSON retrieved matching diagnostic text at 75.6% recall@1 (PRISM, 58.1%) and generated captions closer to report-derived references than PRISM and PRISM2 on the internal cohort and on most external comparisons. Composite images thus offer a compact, clinically aligned computational unit for pathology.
Comments56 pages, 6 main figures, with 11 additional figures and 28 tables in the appendices