发表机构
CETYS Universidad(塞提斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于预训练视觉基础模型的两阶段检索增强框架,经多数据集测试,可实现跨异构数据集的稳健白血病细胞分类,为解决领域偏移问题提供低成本替代方案。
AI 中文摘要
白血病细胞图像分类面临来自采集、染色、光照和站点协议等真实领域偏移的挑战,导致单数据集模型在真实临床场景中泛化能力较差。本研究提出一种稳健框架,采用预训练视觉基础模型的两阶段流水线实现跨多个异构数据集的白血病分类。第一阶段执行白血病与非白血病的二分类,使用122167张单细胞图像进行训练;第二阶段仅在第一阶段的阳性样本上有条件应用,将亚型分类为急性淋巴细胞白血病(ALL)和急性髓系白血病(AML),使用69400张单细胞图像进行训练。为实现跨数据集训练,对五个异构数据集的标签进行了统一协调,并采用留出数据集协议评估性能以衡量领域偏移下的泛化能力。在该流水线中,对三种编码器进行了基准测试,分别是在单细胞图像上预训练的DinoBloom、在生物医学数据上预训练的BiomedCLIP,以及通用模型CLIP,测试场景包括线性探测、低秩适配(LoRA)和检索增强分类(RAC)模块,该模块会检索前k个最相似的细胞图像以提供细胞形态学依据。本研究的目标是量化特定领域预训练在领域偏移下对性能的贡献程度,以及高性价比的适配和检索是否可替代成本高昂的领域专用预训练。此外,留出协议还作为诊断工具,可揭示分类性能何时归因于数据集特定的人工制品而非细胞形态学特征。
英文摘要
Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios. This work presents a robust framework for leukemia classification across multiple heterogeneous datasets using a two-stage pipeline with a pretrained vision foundation model. Stage 1 performs binary classification (leukemia vs. non-leukemia) and is trained using 122,167 single-cell images. Stage 2 is conditionally applied to Stage 1 positives to perform subtype classification into Acute Lymphoblastic Leukemia (ALL) and Acute Myeloid Leukemia (AML), trained using 69,400 single-cell images. Labels are harmonized across five heterogeneous datasets to enable cross-dataset training, and performance is evaluated on a held-out dataset protocol to assess domain-shift generalization. Within this pipeline, three encoders are benchmarked (DinoBloom, pretrained on single-cell images; BiomedCLIP, pretrained on biomedical data; and CLIP as a general-purpose model) under linear probing, Low-Rank Adaptation (LoRA), and a Retrieval-Augmented Classification (RAC) module that retrieves the top-k most similar cell images to provide cytomorphological grounding. The objective is to quantify how much domain-specific pretraining contributes to performance under domain shift, and whether cost-effective adaptation and retrieval can be a viable alternative to expensive domain-specialized pretraining. The held-out protocol additionally serves as a diagnostic tool, revealing when classification performance is attributable to dataset-specific artifacts rather than to cytomorphological features.
CommentsAccepted at SPIE Optics + Photonics 2026 for oral presentation. 23 pages, 12 figures, 9 tables