全切片图像基础模型中的形态学信号可自动对切片进行分类
Morphology signal in whole slide image foundation models can automatically triage slides
浏览论文内容
中文总结 AI 辅助
该研究提出利用公开的全切片图像基础模型,通过零样本分类预测对患者的多张全切片图像排序,可准确识别肿瘤最多的切片,实现自动切片分类,且在多数据集上验证了有效性。
中文摘要 AI 辅助
癌症诊断与分期过程中的患者检查通常会生成多张全切片图像(WSI)。在WSI数据上训练模型的初始步骤之一,是识别包含肿瘤或其他下游预测任务(如评估复发风险或无进展生存期)所需诊断生物标志物的一张或多张切片。该步骤需要经验丰富的病理学家进行繁琐的手动筛选。许多已发布的数据集人为假设每位患者仅对应1张切片;另一种方式是将每位患者的所有切片用于模型训练,这可能会稀释少数包含肿瘤或其他相关信息的切片所携带的信号。本文中,我们提出了一种利用公开可用的WSI基础模型(FMs)克服这些挑战的流程。我们的评估表明,基于WSI FMs的零样本分类预测对WSI进行排序,可准确识别肿瘤最多的切片,说明WSI FMs包含足够的形态学信号以自动对切片进行分类。我们还提出了一种排序评估的形式化方法,用于基准测试FMs在切片分类任务中的性能。我们在多个数据集上证明,对于每位患者最多43张切片的情况,肿瘤切片可被识别为排名前2的切片。
英文摘要
Patient exams in the cancer diagnosis and staging process typically generate several whole slide images (WSIs). One of the initial steps in training models on WSI data is identifying one or a few slides containing tumor or other diagnostic biomarkers necessary for downstream prediction tasks such as estimating recurrence risk or progression-free survival. This step requires tedious manual curation by experienced pathologists. Many published datasets make the artificial assumption of 1 slide per patient. Alternatively, all slides per patient may be used for model training, which may dilute the signal from the few slides containing tumor or other relevant information. In this paper, we present a pipeline to overcome these challenges using publicly available WSI foundation models (FMs). Our evaluations show that ranking WSIs based on predictions from zero-shot classification using WSI FMs accurately identifies slides with the most tumor, indicating that WSI FMs contain sufficient morphology signal to automatically triage slides. We also present a formulation for ranked evaluation to benchmark FM performance in slide triage. We show, on multiple datasets, that tumor slides are identified in the top-2 ranked slides for patients with up to 43 slides.
发表机构
- Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。