在数据集偏移下使用深度集成进行基于ROI的甲状腺结节超声分类的校准选择性预测:回顾性评估
Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation
浏览论文内容
中文总结 AI 辅助
研究基于深度学习的甲状腺结节超声分类,开发校准的深度集成模型及选择性分诊框架,在TN5000上效果良好,在TN3K上表现变差,框架内部辨别力和校准能力强但外部阈值可转移性有限,选择性预测有帮助但部署前需更多评估。
中文摘要 AI 辅助
背景:深度学习模型可对甲状腺结节进行超声分类,但可靠的临床决策支持还需要校准概率、不确定性估计和选择性转诊,尤其是在数据集偏移的情况下。方法:我们开发了一种校准的确定性五成员深度集成模型,用于基于ROI的甲状腺结节分类和基于图像的选择性分诊。使用TN5000进行模型开发、五折交叉验证、成员级向量缩放校准和特定折叠阈值选择。TN3K用作独立的外部数据集偏移评估。该框架使用带有挤压和激励注意力的ConvNeXt-Tiny、集成平均恶性概率和互信息(MI)作为集成不一致分数。采用三层策略将图像分配为无需细针穿刺(No-FNA)建议、细针穿刺(FNA)推荐或放射科医生审查。结果:在汇总的TN5000交叉验证预测中,集成模型的AUC-ROC为0.9395,AP为0.9715,ECE为0.0088,Brier评分为0.0813。在50%的名义MI保留率下,7.2%的病例收到No-FNA建议,39.9%收到FNA推荐,52.9%由放射科医生审查,No-FNA的阴性预测值为98.3%,恶性肿瘤捕获率为99.83%。在TN3K上,AUC-ROC降至0.7870,AP降至0.7254,ECE升至0.1899,Brier评分升至0.2281。冻结的TN5000策略将83.7%分配给审查,1.0%分配给No-FNA,15.3%分配给FNA推荐。没有恶性图像进入No-FNA途径,但FNA推荐的阳性预测值降至76.6%。结论:该框架显示出强大的内部辨别力和校准能力,但外部阈值可转移性有限。选择性预测可能有助于识别不适合自动分诊的图像,但在部署前需要进行局部重新校准、阈值验证和前瞻性临床评估。
英文摘要
Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift. Methods: We developed a calibrated deterministic five-member deep ensemble for ROI-based thyroid nodule classification and selective image-based triage. TN5000 was used for model development, five-fold cross-validation, member-wise vector-scaling calibration, and fold-specific threshold selection. TN3K served as an independent external dataset-shift evaluation. The framework used ConvNeXt-Tiny with squeeze-and-excitation attention, ensemble-mean malignancy probability, and mutual information (MI) as an ensemble-disagreement score. A three-tier policy assigned images to No-FNA suggestion, FNA recommendation, or radiologist review. Results: On pooled out-of-fold TN5000 predictions, the ensemble achieved AUC-ROC 0.9395, AP 0.9715, ECE 0.0088, and Brier score 0.0813. At 50% nominal MI retention, 7.2% of cases received a No-FNA suggestion, 39.9% an FNA recommendation, and 52.9% radiologist review, with 98.3% No-FNA NPV and 99.83% malignancy capture. On TN3K, AUC-ROC decreased to 0.7870, AP to 0.7254, ECE increased to 0.1899, and Brier score to 0.2281. The frozen TN5000 policy assigned 83.7% to review, 1.0% to No-FNA, and 15.3% to FNA recommendation. No malignant image entered the No-FNA pathway, but FNA-recommendation PPV fell to 76.6%. Conclusion: The framework showed strong internal discrimination and calibration, but limited external threshold transportability. Selective prediction may help identify images unsuitable for automated triage, but local recalibration, threshold validation, and prospective clinical evaluation are required before deployment.
发表机构
- Daffodil International University(达福德国际大学)
- Birmingham City University(伯明翰城市大学)
机构由 AI 辅助整理,请以论文原文为准。