arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09863cs.CV

预训练和蒸馏比架构家族对无标签单细胞分类更重要

Pretraining and Distillation Matter More Than Architecture Family for Label-Free Single-Cell Classification

Philip Graemer, Giuseppe Di Caprio

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过受控基准发现,预训练和知识蒸馏对无标签单细胞分类性能的影响远大于架构家族选择,蒸馏可显著提升部署效率。

中文摘要 AI 辅助

为无标签单细胞分类选择深度学习架构仍是一个开放问题,显微镜基准测试对CNN与Transformer的比较得出了相互矛盾的结论。我们在LIVECell相衬显微镜数据上提出了一个受控基准,使用源图像不重叠的训练/验证/测试划分以防止父图像泄漏,并在EfficientNet、Vision Transformer (ViT)和EVA-02模型上匹配优化、增强和评估协议。这使得架构、预训练、微调、令牌化和蒸馏的影响得以解耦。我们发现,先前报道的CNN优势主要由预训练而非架构解释:最小的预训练模型在参数少得多的情况下优于从零训练的最强模型。预训练将宏F1分数提高了3-4个百分点,而最佳预训练CNN与Transformer之间的差距低于0.5个百分点。然而,架构选择仍然重要:ViT-S/8优于ViT-S/16,并以四分之一的参数匹配四倍大的ViT-B/16,表明更细的令牌化有利于小细胞裁剪。相反,层间学习率衰减(EVA-02微调配方的核心)会降低性能,突显了从自然图像识别中迁移的启发式方法可能不适用于显微镜。最后,知识蒸馏显著改善了部署前沿:从教师委员会蒸馏的紧凑型EfficientNet-B0学生模型优于每个单独训练的主干网络,包括EfficientNet-B5和EVA-02教师模型。总体而言,我们的结果表明,严格控制预训练和评估对于解释生物医学架构基准至关重要,而蒸馏可能是比单独选择架构更有效的实用单细胞分类途径。

英文摘要

Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflicting conclusions about CNNs versus transformers. We present a controlled benchmark on LIVECell phase-contrast microscopy data using source-image-disjoint train/validation/test splits to prevent parent-image leakage and matched optimisation, augmentation, and evaluation protocols across EfficientNet, Vision Transformer (ViT), and EVA-02 models. This allows the effects of architecture, pretraining, fine-tuning, tokenisation, and distillation to be disentangled. We find that the previously reported CNN advantage is largely explained by pretraining rather than architecture: the smallest pretrained model outperforms the strongest model trained from scratch despite far fewer parameters. Pretraining improves macro-F1 by 3-4 points, while the gap between the best pretrained CNN and transformer is below 0.5 points. Architectural choices nevertheless matter: ViT-S/8 outperforms ViT-S/16 and matches the four-times-larger ViT-B/16 at a quarter of the parameters, showing that finer tokenisation benefits small cell crops. Conversely, layer-wise learning-rate decay, central to the EVA-02 fine-tuning recipe, degrades performance, highlighting that transfer heuristics from natural-image recognition may not generalise to microscopy. Finally, knowledge distillation substantially improves the deployment frontier: compact EfficientNet-B0 students distilled from teacher councils outperform every individually trained backbone, including the EfficientNet-B5 and EVA-02 teachers. Overall, our results show that rigorous control of pretraining and evaluation is essential for interpreting biomedical architecture benchmarks, while distillation may be a more effective route to practical single-cell classification than architecture choice alone.

发表机构

  • University of Strathclyde(思克莱德大学)
  • University of Glasgow(格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑