发表机构
Tsinghua Shenzhen International Graduate School, Tsinghua University; University of Washington; City University of Hong Kong (Dongguan); Medical Optical Technology R&D Center, Research Institute of Tsinghua; Jinfeng Laboratory(清华大学深圳国际研究生院; 华盛顿大学; 香港城市大学(东莞); 清华研究院医学光学技术研发中心; 金凤实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对计算病理学中多实例学习存在的问题,提出基于蒸馏的预训练框架,利用两个基础模型作为教师,引入角分散归一化蒸馏损失,将蒸馏权重用于下游适应,实验表明该方法在少样本场景中优势明显,能提升计算效率。
AI 中文摘要
多实例学习(MIL)已成为计算病理学中全切片图像(WSI)分析的主要范式。现有MIL聚合器通常针对每个下游任务从头开始训练,依赖有限的切片级标签同时学习聚合机制和下游判别表示,存在优化不稳定、过拟合和可迁移性有限等问题。本文提出基于蒸馏的MIL预训练框架,利用两个切片级基础模型TITAN和CARE作为教师,将其表示知识转移到多种MIL架构中。引入角分散归一化蒸馏损失平衡不同教师的监督,蒸馏权重用作下游适应的初始化。在15个基准数据集上进行系统评估,结果表明预训练总体上优于从头训练,尤其在少样本场景中,同时保持轻量级MIL模型的计算效率。
英文摘要
Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology. However, existing MIL aggregators are still typically trained from scratch for each downstream task, relying on limited slide-level labels to learn both aggregation mechanisms and downstream discriminative representations simultaneously. As a result, they often suffer from unstable optimization, overfitting, and limited transferability. Similar to pretrained ResNet and Vision Transformer models in natural image learning, MIL also requires reusable pretrained initialization. However, high-quality slide-level pretraining data remain scarce, and MIL models are usually lightweight and weakly supervised, making large-scale pretraining difficult in practice. To address this challenge, we propose a distillation-based pretraining framework for MIL, which leverages two slide-level foundation models, TITAN and CARE, as teachers to transfer their representational knowledge into a diverse set of MIL architectures. To effectively balance supervision from different teachers, we further introduce an angular dispersion normalized distillation loss. The distilled weights are then used as initialization for downstream adaptation. We conduct systematic evaluations on 15 benchmark datasets under both linear probing and full-parameter fine-tuning, and further validate its advantages in few-shot scenarios. Experimental results show that pretraining generally improves MIL aggregators over from scratch training, especially in linear-probing and few-shot settings, while maintaining the computational efficiency of lightweight MIL models. Code is available at https://github.com/fu0201/MIL_Pretrained.