发表机构
Binghamton University(宾汉姆顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出可靠性感知的混合K集成选择框架,结合判别力、校准与选择性预测,在宫颈细胞学分类中通过软投票集成Swin-Tiny和TinyViT-5M,显著降低AURC、NLL和WC-ECE,并展现出对度量权重选择的鲁棒性。
AI 中文摘要
仅凭高分类准确率不足以满足临床图像分析的需求,其中校准的置信度和可靠的确定性估计至关重要。本研究提出了一种可靠性感知的混合K集成选择框架,用于使用SIPaKMeD数据集进行多类别宫颈细胞学分类。使用固定的分层五折划分和三个训练种子评估了九种深度学习架构。在事后温度缩放后,使用宏F1、准确率、AUROC、期望校准误差(ECE)、最差类别ECE(WC-ECE)、风险覆盖曲线下面积(AURC)、Brier分数和负对数似然(NLL)对模型进行评估。使用等权复合分数对模型进行排名,并通过软投票从排名靠前的模型中形成混合K集成。使用5,000个Dirichlet采样的度量权重向量、留一度量分析和跨15次折乘种子评估的校正配对检验来检验鲁棒性。最终的混合2集成由Swin-Tiny和TinyViT-5M组成,相对于最佳单个模型,AURC降低了43%,NLL降低了17%,WC-ECE降低了36%。它在96.8%的随机加权场景中被选中,在所有留一度量分析中保持不变,并提高了完整复合分数。然而,在Holm-Bonferroni校正后,各度量的增益在统计上不显著(所有调整后p >= 0.168)。由于事后校准未使用完全独立的校准集,依赖校准的结果应视为探索性的内部估计。总体而言,该框架识别了一个紧凑的集成,对替代度量权重具有鲁棒性,并在单个数据集的内部验证下提高了可靠性点估计。
英文摘要
High classification accuracy alone is insufficient for clinical image analysis, where calibrated confidence and reliable uncertainty estimates are essential. This study proposes a reliability-aware Hybrid-K ensemble selection framework for multiclass cervical cytology classification using the SIPaKMeD dataset. Nine deep learning architectures were evaluated using a fixed stratified five-fold partition and three training seeds. After post-hoc temperature scaling, models were assessed using macro-F1, accuracy, AUROC, expected calibration error (ECE), worst-class ECE (WC-ECE), area under the risk-coverage curve (AURC), Brier score, and negative log-likelihood (NLL). Models were ranked using an equal-weight composite score, and Hybrid-K ensembles were formed from the top-ranked models using soft voting. Robustness was examined using 5,000 Dirichlet-sampled metric-weight vectors, leave-one-metric-out analysis, and corrected paired testing across 15 fold-by-seed evaluations. The final Hybrid-2 ensemble, comprising Swin-Tiny and TinyViT-5M, reduced AURC by 43%, NLL by 17%, and WC-ECE by 36% relative to the best individual model. It was selected in 96.8% of random weighting scenarios, remained unchanged across all leave-one-metric-out analyses, and improved the full composite score. However, per-metric gains were not statistically significant after Holm-Bonferroni correction (all adjusted p >= 0.168). Because post-hoc calibration did not use a fully independent calibration set, calibration-dependent results should be interpreted as exploratory internal estimates. Overall, the framework identified a compact ensemble robust to alternative metric weightings and improved reliability point estimates under internal validation on a single dataset.