发表机构
Rajshahi University of Engineering & Technology (RUET); Elite Research Lab LLC; Multimedia University(拉杰沙希工程技术大学; 精英研究实验室有限责任公司; 多媒体大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对胸部X射线分类模型,提出DBCA-SegNet-MGAP框架,量化肺部归因包含性,发现诊断性能、校准度与肺部归因是独立属性,为域偏移下的模型评估提供新视角。
AI 中文摘要
深度学习模型可在胸部X射线(CXR)分类任务中取得优异性能,但无法确定其预测是否主要依赖肺部图像内容。本研究将肺部归因包含性作为与诊断性能不同的解剖学相关可靠性属性进行评估。我们提出DBCA-SegNet-MGAP,这是一种多任务解剖学引导的CNN-Transformer框架,通过双向跨骨干注意力结合互补特征表示,预测软肺部掩码,并通过掩码引导的自适应全局平均池化(MGAP)将此解剖学先验直接融入分类。采用解剖学局部能量比(ALR)和高强度累积ALR(cALR@0.9)量化肺部归因包含性。实验在三个训练随机种子下重复开展,使用COVID-19放射图像数据库进行四类内部测试,并采用锁定的深圳至蒙哥马利协议进行零样本外部肺结核测试。在COVID-19数据集上,所提模型的加权F1为$0.9615 \textpm 0.0015$,宏ROC-AUC为$0.9906 \textpm 0.0007$。在架构匹配的双分支对比中,将传统GAP替换为MGAP后,ALR从$0.3878 \textpm 0.0098$提升至$0.7086 \textpm 0.0104$,cALR@0.9从$0.5265 \textpm 0.0101$提升至$0.9905 \textpm 0.0018$,而加权F1基本保持不变($0.9618 \textpm 0.0015$ vs. $0.9615 \textpm 0.0015$)。在锁定的外部迁移至蒙哥马利数据集的情况下,ROC-AUC保持为$0.9080 \textpm 0.0043$,肺部ALR保持为$0.6466 \textpm 0.0081$,而加权F1降至$0.7528 \textpm 0.0080$,预期校准误差(ECE)升至$0.1683 \textpm 0.0055$。这些发现表明,诊断判别力、校准度和肺部归因包含性是模型的不同属性,支持在内部测试和外部域偏移下对其进行联合评估。
英文摘要
Deep-learning models can achieve strong chest X-ray (CXR) classification performance without establishing whether their predictions predominantly rely on pulmonary image content. This study evaluates pulmonary attribution containment as an anatomy-related reliability property distinct from diagnostic performance. We propose DBCA-SegNet-MGAP, a multi-task anatomy-guided CNN-Transformer framework that combines complementary feature representations through bidirectional cross-backbone attention, predicts a soft lung mask, and incorporates this anatomical prior directly into classification through Mask-Guided Adaptive Global Average Pooling (MGAP). Pulmonary attribution containment is quantified using the Anatomical Local Energy Ratio (ALR) and high-intensity cumulative ALR (cALR@0.9). Experiments were repeated across three training seeds using the COVID-19 Radiography Database for four-class internal testing and a locked Shenzhen-to-Montgomery protocol for zero-shot external tuberculosis testing. On COVID-19, the proposed model achieved a weighted F1 of $0.9615 \pm 0.0015$ and macro ROC-AUC of $0.9906 \pm 0.0007$. In an architecture-matched dual-bridge comparison, replacing conventional GAP with MGAP increased ALR from $0.3878 \pm 0.0098$ to $0.7086 \pm 0.0104$ and cALR@0.9 from $0.5265 \pm 0.0101$ to $0.9905 \pm 0.0018$, while weighted F1 remained essentially unchanged ($0.9618 \pm 0.0015$ vs. $0.9615 \pm 0.0015$). Under locked external transfer to Montgomery, ROC-AUC remained $0.9080 \pm 0.0043$ and pulmonary ALR remained $0.6466 \pm 0.0081$, whereas weighted F1 decreased to $0.7528 \pm 0.0080$ and ECE increased to $0.1683 \pm 0.0055$. These findings show that diagnostic discrimination, calibration, and pulmonary attribution containment are distinct model properties and support their joint evaluation under internal testing and external domain shift.
Comments34 pages, 8 figures, 7 tables. Code available at https://github.com/Abdullah-229/BeyondAccuracy