arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMFE:用于疟疾病细胞分类的轻量级可解释机器学习框架

EMFE: A lightweight, explainable machine learning framework for malaria cell classification

Md Abdullah Al Kafi, Walayat Hussain, Mousumi Karmakar, Sumit Kumar Banshal, Ahmed Al Marouf

arXiv 2608.24793首次发表:更新:

发表机构

Daffodil International University; Peter Faber Business School; Australian Catholic University; Alliance University; University of Alberta(水仙国际大学; 彼得·费伯商学院; 澳大利亚天主教大学; 联盟大学; 阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出轻量级可解释机器学习框架 EMFE,结合经典机器学习方法,在疟疾细胞分类任务中取得高准确率,可替代深度学习模型并明确量化其局限性。

AI 中文摘要

stained blood-smear microscopy 的自动化疟疾诊断目前主要依赖深度卷积神经网络,这类模型准确率高但计算成本高、可解释性差,且很少以患者层面的严谨性进行验证。我们提出 EMFE(高效数学特征提取),这是一个包含5个特征的框架,用于将单个红细胞图像分类为受感染或未受感染,使用 Gray World 颜色归一化、自适应绿色通道阈值处理、形态斑点检测和经典机器学习方法。我们使用 NIH LHNCBC 疟疾数据集(来自200名患者的27558张图像),在患者分组的嵌套交叉验证(K_outer=20,K_inner=3)下评估随机森林、直方图梯度提升和支持向量机分类器,确保每位患者的细胞都保留在单个折中。优化后的随机森林实现了94.6%的合并折外准确率(95%置信区间[93.6, 95.7]),并通过未接触过的40名患者的保留测试(94.3%)和患者层面的置换检验(p<0.001,1000次置换)得到证实。消融实验量化了单个特征和管道阶段的贡献,与重新训练的 DenseNet121、ResNet50 和 MobileNetV2 模型进行硬件匹配比较,评估了准确率-效率权衡。合成扰动表征了三种失败模式,可解释性分析确定斑点饱和度为主要判别特征,患者层面的聚合进一步量化了灵敏度-特异性权衡和假阳性积累。这些结果证明了一种统计严谨、可解释且计算轻量的深度学习替代方案,同时明确量化了其局限性。

英文摘要

Automated malaria diagnosis from stained blood-smear microscopy is dominated by deep convolutional neural networks that are accurate but computationally expensive, poorly interpretable, and rarely validated with patient-level rigor. We present EMFE (Efficient Mathematical Feature Extraction), a five-feature framework for classifying single red-blood-cell images as parasitized or uninfected using Gray World color normalization, adaptive green-channel thresholding, morphological spot detection, and classical machine learning. Using the NIH LHNCBC malaria dataset (27,558 images from 200 patients), we evaluate Random Forest, Histogram Gradient Boosting, and Support Vector Machine classifiers under patient-grouped nested cross-validation (K_outer=20, K_inner=3), ensuring that cells from each patient remain within a single fold. The optimized Random Forest achieves 94.6% pooled out-of-fold accuracy (95% CI [93.6, 95.7]), corroborated by an untouched 40-patient holdout test (94.3%) and a patient-level permutation test (p<0.001, 1,000 permutations). Ablation experiments quantify the contribution of individual features and pipeline stages. Hardware-matched comparisons with retrained DenseNet121, ResNet50, and MobileNetV2 models assess the accuracy-efficiency trade-off. Synthetic perturbations characterize three failure modes, while explainability analysis identifies spot saturation as the dominant discriminative feature. Patient-level aggregation further quantifies sensitivity-specificity trade-offs and false-positive accumulation. These results demonstrate a statistically rigorous, interpretable, and computationally lightweight alternative to deep learning, while explicitly quantifying its limitations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑