ADMIL:用于病理学选择性基础模型推理的注意力蒸馏多实例学习
ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology
浏览论文内容
中文总结 AI 辅助
ADMIL是将ABMIL教师注意力蒸馏为轻量级图块选择模型PriorNet的选择性计算框架,在多个病理数据集上可移除>98%的基础模型图块编码且不损失切片级性能,提升了部署效率。
中文摘要 AI 辅助
基于注意力的多实例学习(ABMIL)利用病理学基础模型嵌入对病理切片级任务有效,但穷尽式推理需对所有前景图块应用大型图像编码器,而后续注意力分布通常集中在少量信息丰富的区域。我们提出ADMIL(注意力蒸馏多实例学习),这是一种选择性计算框架,将ABMIL教师模型的注意力蒸馏为轻量级图块选择模型PriorNet。PriorNet采用EfficientNet架构,通过KL散度从原始图块像素学习教师注意力分布;推理时,它对前景池评分,选择前K个图块,仅对该子集调用昂贵的基础模型,之后由所选袋的ABMIL学生模型预测切片标签。在BRACS、PANDA和CAMELYON16数据集上,ADMIL分别在K=4、8、128个图块时匹配完整教师模型的 headline 性能,避免了>98%的基础模型(Virchow2)图块嵌入和模型推理FLOPs。随机和教师注意力 oracle 对照实验表明,该结果依赖于任务相关选择,而非仅图块数量减少。定量和定性分析显示,PriorNet高保真恢复教师的图块排序,同时聚焦于任务相关的形态区域。ADMIL表明,可在不牺牲切片级性能的情况下移除几乎所有昂贵的图块编码,为延迟和计算成本是关键考量的临床环境提供了更高效部署的潜在路径。
英文摘要
Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exhaustive inference requires applying a large image encoder to every foreground tile despite the subsequent attention distribution often concentrating over a small subset of informative regions. We introduce ADMIL (Attention-Distilled Multiple Instance Learning), a selective-compute framework that distills an ABMIL teacher's attention into a lightweight tile-selection model, PriorNet. Using an EfficientNet architecture, PriorNet learns the teacher attention distribution from raw tile pixels with KL divergence; at inference, it scores the foreground pool, selects the top-K tiles, and invokes the expensive foundation model only on that subset before a selected-bag ABMIL student predicts the slide label. Across BRACS, PANDA, and CAMELYON16, ADMIL matches full-teacher headline performance at K=4, 8, and 128 tiles, respectively, avoiding >98% of foundation model (Virchow2) tile embeddings and model inference FLOPs. Random and teacher-attention oracle controls show that this result depends on task-relevant selection rather than tile-count reduction alone. Quantitative and qualitative analyses suggest that PriorNet recovers the teacher's tile ordering with high fidelity while focusing on task-relevant morphological regions. ADMIL shows that nearly all expensive tile encodings can be removed without sacrificing slide-level performance, providing a potential path for more efficient deployment in clinical settings where latency and compute costs are key considerations.