AI 中文总结
针对液基宫颈细胞学样本类不平衡问题,研究采用加权随机采样训练三种轻量级架构,与两种软投票集成方法比较,通过事后温度缩放拟合校准子集,发现校准对提升可靠性更关键,集成规模优势不明显。
AI 中文摘要
宫颈细胞学分类模型通常在经过整理的、类平衡的基准上进行评估,但实际的液基细胞学(LBC)样本往往规模小且类不平衡。本文针对门捷列夫LBC数据集进行了类不平衡感知和校准感知的集成分类研究,使用其原生的四类贝塞斯达分类法(NILM、LSIL、HSIL、SCC)而非简化的二元形式。通过加权随机采样直接在门捷列夫LBC上训练三种轻量级架构(Swin-Tiny、TinyViT-5M、DenseNet121)以应对类不平衡,并与两种软投票集成方法(Hybrid-2、Hybrid-3)进行比较。事后温度缩放拟合在每个交叉验证折的训练部分划分出的保留校准子集上,该子集与用于拟合模型权重的训练数据和用于最终指标评估的折不同,避免了因同一数据用于两个目的而产生的乐观校准估计。校准显著降低了每个测试模型和集成配置的预期校准误差、布里尔分数和负对数似然,而判别指标(准确率、宏F1、宏AUROC)基本保持不变。一旦所有配置都经过适当校准,集成规模相对于最佳单个模型没有一致的额外可靠性优势。混淆矩阵表明,所有配置的分类错误都局限于高级别病变(HSIL)和癌(SCC)之间的边界;没有错误涉及阴性(NILM)或低级别(LSIL)类别。这些结果表明,对于该数据集,校准是可靠性的主要影响因素,而非集成规模,不过这一结论应结合数据集的适度规模来解读。
英文摘要
Cervical cytology classification models are typically evaluated on curated, class-balanced benchmarks, but real-world liquid-based cytology (LBC) collections are often small and class-imbalanced. This paper presents a class-imbalance-aware and calibration-aware ensemble classification study on the Mendeley LBC dataset, using its native four-class Bethesda taxonomy (NILM, LSIL, HSIL, SCC) rather than a collapsed binary formulation. Three lightweight architectures (Swin-Tiny, TinyViT-5M, DenseNet121) are trained directly on Mendeley LBC using weighted random sampling to counteract class imbalance, and compared against two soft-voting ensembles (Hybrid-2, Hybrid-3). Post-hoc temperature scaling is fit on a held-out calibration subset carved out of the training portion of each cross-validation fold, distinct from both the training data used to fit model weights and the evaluation fold used for final metrics, avoiding the optimistic calibration estimates that result when the same data is used for both purposes. Calibration substantially reduces expected calibration error, Brier score, and negative log-likelihood for every model and ensemble configuration tested, while discrimination metrics (accuracy, macro-F1, macro-AUROC) remain essentially unchanged. Ensemble size shows no consistent additional reliability benefit over the best individual model once all configurations are properly calibrated. Confusion matrices show that all classification errors, across every configuration, are confined to the boundary between high-grade lesions (HSIL) and carcinoma (SCC); no errors involve the negative (NILM) or low-grade (LSIL) categories. These results suggest that, for this dataset, calibration is the dominant lever for reliability, not ensemble size, though this conclusion should be read in light of the dataset's modest size.