arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17872cs.CV

DistillPath:一种高效的22M级蒸馏病理编码器,性能接近大型基础模型

DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance

Ramon Kaspar, Andrey Ignatov, Valentina Boeva

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出22M级蒸馏病理编码器DistillPath-KS16,通过蒸馏86M至1.1B的病理编码器得到,性能接近大型模型Virchow2,参数少29倍、速度快25倍以上,在EVA等基准上表现优异。

中文摘要 AI 辅助

许多高性能的病理切片编码器如今是拥有数亿至超过十亿参数的基础模型。使用这类模型对每张全切片图像中的数千个切片进行编码和存储,在普通硬件上成本高昂,因此,保留有用下游性能的紧凑编码器是极具价值的替代方案。我们提出了DistillPath-KS16,它从现有的22M kaiko ViT-S/16编码器出发,通过从已发布的病理编码器(作为冻结教师模型)中进行蒸馏来对其进行改进。该方法仅读取教师模型的最终类别标记和切片标记,并在6000张公开切片上进行训练,无需使用它们的DINO或iBOT预训练头,也无需使用十亿级切片语料库,因此它适用于任何公开的、能输出骨干网络标记的编码器。我们将四个参数规模从86M到1.1B的教师模型蒸馏到同一个学生模型中。所有变体在我们使用的三个基准(EVA、HEST和PLISM)上均优于kaiko基线,且最强教师模型取决于具体任务。在七任务EVA均值上,DistillPath-KS16-Virchow2达到0.795,与我们评估中得分最高的模型Virchow2仅相差0.015个点,但其参数数量约少29倍;在该聚合指标上,它的得分也高于H0-mini和GPFM,不过这一优势集中在特定任务而非均匀分布。由于它仍是22M级、具有384维特征的ViT-S/16,DistillPath-KS16的运行速度比Virchow2快25倍以上。代码可在此httpsURL获取,发布的模型权重可在此httpsURL获取。

英文摘要

Many high-performing pathology tile encoders are now foundation models with hundreds of millions to over a billion parameters. Encoding and storing the thousands of tiles in each whole-slide image with such models is costly on commodity hardware, so compact encoders that retain useful downstream performance are a valuable alternative. We present DistillPath-KS16, which starts from the existing 22M kaiko ViT-S/16 encoder and improves it by distilling from released pathology encoders used as frozen teachers. The recipe reads only the teachers' final class and patch tokens and trains on 6,000 public slides, needing neither their DINO nor iBOT pretraining heads nor a billion-tile corpus, so it applies to any released encoder that exposes backbone tokens. We distill four teachers spanning 86M to 1.1B parameters into the same student. Every variant improves the kaiko baseline on all three benchmarks we use, EVA, HEST, and PLISM, and the strongest teacher is task-dependent. On the seven-task EVA mean, DistillPath-KS16-Virchow2 reaches $0.795$, within $0.015$ points of Virchow2, the top-scoring model in our evaluation, at about $29\times$ fewer parameters; it also scores above H0-mini and GPFM on this aggregate metric, though that advantage is task-concentrated rather than uniform. Because it remains a 22M ViT-S/16 with 384-dimensional features, DistillPath-KS16 runs more than $25\times$ faster than Virchow2. Code is available at https://github.com/RamonKaspar/DistillPath, and released model weights are available at https://huggingface.co/collections/RamonK/distillpath.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑