发表机构
Artificial Intelligence Research Institute, Shenzhen MSU-BIT University(深圳北理莫斯科大学人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出角色引导的MoE-FFN框架,通过两阶段训练实现编码器级病理表示学习,在WSI分类中超越最强基线。
AI 中文摘要
全切片图像分类是计算病理学中的一项基础任务,其中图像块表示质量直接影响下游聚合和切片级判别能力。病理学基础模型被广泛用作WSI分类的冻结特征提取器;然而,其固定编码器可能产生对目标特定组织模式和判别线索适应性不足的表示。微调可以改善目标适应性,但引入了病理特异性表示能力与适应效率之间的权衡,特别是在数据稀缺场景下。为解决这一问题,我们提出了一种病理角色引导的专家混合前馈网络(MoE-FFN)框架,用于高效的编码器级表示学习。我们设计了一个两阶段训练范式来建立和适应病理感知的专家专业化。在源域专家初始化阶段,病理特异性先验从冻结的Virchow2教师模型中蒸馏到轻量级DINOv2-small学生模型中,同时角色原型作为弱病理锚点以鼓励不同的专家功能。MoE-FFN块被引入到选定的高层Transformer层中,为异质病理模式提供变换多样性。在目标域适应阶段,初始化专家通过非对称原型引导优化进行细化,增强任务相关的正证据并分离易混淆的硬负样本。由此产生的编码器提取离线图像块表示,可直接与标准MIL聚合器集成。在公共BRACS数据集和私有PAROTID WSI数据集上,跨五个代表性骨干网络的实验表明,相较于最强基线,该方法取得了一致的改进。
英文摘要
Whole slide image classification is a fundamental task in computational pathology, where patch representation quality directly affects downstream aggregation and slide-level discriminability. Pathology foundation models are widely adopted as frozen feature extractors for WSI classification; however, their fixed encoders may produce representations insufficiently adapted to target-specific tissue patterns and discriminative cues. Fine-tuning can improve target adaptation, but introduces a trade-off between pathology-specific representation capacity and adaptation efficiency, particularly in data-scarce settings. To address this, we propose a pathology role-guided mixture-of-experts feed-forward network (MoE-FFN) framework for efficient encoder-level representation learning. We design a two-stage training paradigm to establish and adapt pathology-aware expert specialization. In source-domain expert initialization, pathology-specific priors are distilled from a frozen Virchow2 teacher into a lightweight DINOv2-small student, while role prototypes serve as weak pathological anchors to encourage distinct expert functions. MoE-FFN blocks are introduced into selected high-level transformer layers to provide transformation diversity for heterogeneous pathological patterns. In target-domain adaptation, the initialized experts are refined through asymmetric prototype-guided optimization, enhancing task-relevant positive evidence and separating confusable hard negatives. The resulting encoder extracts offline patch representations that can be directly integrated with standard MIL aggregators. Experiments on the public BRACS dataset and a private PAROTID WSI dataset across five representative backbones demonstrate consistent improvements over the strongest baseline.