发表机构
American International University-Bangladesh (AIUB)(美国国际大学孟加拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估儿科肺炎分类的深度学习模型跨三个国家数据集的可迁移性,测试判别性、校准性等多维度,发现跨国偏移影响不同,有限标签可部分恢复灵敏度但存在特异性变异性。
AI 中文摘要
背景与目的:医学影像AI的外部评估常仅聚焦于判别性。本研究针对三个国家的儿科肺炎分类数据集,评估了一套计算方案,该方案分别测试判别性、概率校准性、固定操作点迁移、捷径关联信号及有限标签可恢复性。方法:经 exact-duplicate 去除后,5824 张广州X光片用于泄漏控制的源模型开发与内部测试;采用冻结的三种子 DenseNet121 双视图集成模型,在来自孟加拉国的 BDCXR-3257 数据集(n=3257)和来自越南的未接触过的 harmonized VinDr-PCXR/PediCXR 测试队列(n=1077)上进行零样本评估;匹配的 seed-42 变体用于测试架构鲁棒性。BDCXR 的二次分析采用固定的 651 张图像适应池和 2606 张保留集;163、326 和 651 个标签分别占 BDCXR 完整标签的 5%、10% 和 20%。结果:内部 AUROC 为 0.976,灵敏度为 95.1%;BDCXR 和 VinDr-PCXR 的 AUROC 分别为 0.798 和 0.742,而冻结阈值下的灵敏度降至 6.2% 和 0%;全图像基线(0.961 至 0.749)、 ungated 双视图模型(0.977 至 0.766)及 gated MixStyle 模型(0.966 至 0.789)的源到 BDCXR AUROC 均出现下降;使用 163 个 BDCXR 标签时,Platt 校准保留了 AUROC,同时将保留集灵敏度提升至 88.3%,但特异性为 47.9%,警报率为 78.5%;200 次重复的 163 标签拟合确认了灵敏度的恢复,但存在显著的特异性变异性。结论:跨国跨数据集偏移对排名、概率对齐和源定义决策行为的影响各异;迁移研究应分别评估这些组件,并量化表观恢复的操作负担。
英文摘要
Background and Objective: External evaluation of medical-imaging AI is often collapsed into discrimination. We evaluated a computational protocol that separately tests discrimination, probability calibration, fixed operatingpoint transport, shortcut-associated signal, and limited-label recoverability for pediatric pneumonia classification across datasets from three countries. Methods: After exact-duplicate removal, 5,824 Guangzhou radiographs supported leakage-controlled source development and internal testing. A frozen three-seed DenseNet121 dual-view ensemble was evaluated zero-shot on BDCXR-3257 from Bangladesh (n = 3, 257) and an untouched harmonized VinDr-PCXR/PediCXR test cohort from Vietnam (n = 1, 077). Matched seed-42 variants tested architectural robustness. Secondary BDCXR analyses used a fixed 651-image adaptation pool and 2,606-image hold-out; 163, 326, and 651 labels represented 5%, 10%, and 20% of complete BDCXR. Results: Internal AUROC was 0.976 with 95.1% sensitivity. BDCXR and VinDr-PCXR AUROC were 0.798 and 0.742, while frozen-threshold sensitivity fell to 6.2% and 0%. Source-to-BDCXR AUROC degradation occurred for a full-image baseline (0.961 to 0.749), ungated dual-view model (0.977 to 0.766), and gated MixStyle model (0.966 to 0.789). With 163 BDCXR labels, Platt recalibration preserved AUROC while increasing held-out sensitivity to 88.3%, but specificity was 47.9% and the alert rate was 78.5%. Two hundred repeated 163-label fits confirmed sensitivity recovery but substantial specificity variability. Conclusions: Cross-dataset shifts across countries affected ranking, probability alignment, and source-defined decision behavior differently. Transport studies should evaluate these components separately and quantify the operational burden of apparent recovery.