现实小样本标注协议下的 whole-slide 图像分析
Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols
- Institute of Information and Communication Technologies, Electronics, and Applied Mathematics (ICTEAM), Université catholique de Louvain (UCLouvain)(天主教鲁汶大学信息与通信技术、电子与应用数学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现实 whole-slide 图像场景的小样本标注问题,提出 SlideCRF 模型,结合空间与生物线索优化预测,在四个数据集上宏观 F1 较现有转导方法更优,零样本预测提升显著。
AI中文摘要:
对 whole-slide 图像的分析实现自动化具有很高的临床价值,因为对癌症的表征需要对其进行详细检查。这类分析越来越依赖能提供补丁级零样本预测的视觉-语言模型,但这些预测仍存在噪声,必须通过少量标注进行优化。一种有前景的优化范式是小样本转导,这类方法不独立处理每个补丁,而是利用补丁间的关系,结合少量标注来共同优化所有预测。然而,当前的转导方法在评估时忽略了 whole-slide 图像的关键特性:(i)数据集由从多张切片中提取的独立补丁组成,未考虑复杂的组织结构;(ii)数据集大多是平衡的,而单张 whole-slide 图像存在严重的类别不平衡,且多个类别缺失;(iii)标注是随机采样的,未反映病理医生标注有限区域的方式。为使转导范式适配现实的 whole-slide 场景,本文做出以下贡献:首先,提出 SlideCRF,它通过结合空间和生物线索,并考虑给定切片中可能缺失的类别,将条件随机场适配到 whole-slide 图像;其次,提供一组基于空间定位的点击和涂鸦的现实标注协议,模拟病理医生的不同交互,如对模型错误的迭代修正。在四个数据集上的实验表明,SlideCRF 在宏观 F1 指标上优于当前的转导方法,在每个存在类别分别使用1次和16次点击时,相比零样本预测分别提升了24.2%和37.5%。
英文摘要:
Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis increasingly relies on vision-language models that provide patch-level zero-shot predictions. However, these predictions remain noisy and must be refined with a few annotations. A promising paradigm for this refinement is few-shot transduction. Rather than treating each patch independently, these methods leverage the relations between patches, together with a few annotations, to refine all predictions jointly. However, current transductive methods are evaluated under conditions that overlook key properties of whole-slide images: (i) datasets consist of independent patches extracted from multiple slides, ignoring the complex tissue organization; (ii) datasets are mostly balanced, whereas a single whole-slide image exhibits severe class imbalance, with several classes absent; and (iii) annotations are sampled at random, without reflecting how a pathologist annotates a limited number of regions. To align the transduction paradigm to realistic whole-slide settings, we introduce the following contributions. First, we propose SlideCRF, which adapts conditional random fields for whole-slide images by combining spatial and biological cues while accounting for classes that may be absent from a given slide. Second, we provide a set of realistic annotation protocols, based on spatially localized clicks and scribbles, modeling different pathologist interactions, such as the iterative correction of model errors. Across four datasets, we show that SlideCRF outperforms current transductive methods in macro F1, improving over the zero-shot predictions by +24.2% and +37.5% with one and 16 clicks per present class, respectively.