跨模态分诊网络:用于胸部X光片基于严重程度的分诊及视觉可解释性的多模态深度学习框架
Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
浏览论文内容
中文总结 AI 辅助
该研究提出跨模态分诊网络(CMTN),融合Swin Transformer V2与PubMedBERT,在胸部X光片分诊任务上性能优于BioViL,但与放射科医生判断一致性较低,凸显临床部署需专家标注数据。
中文摘要 AI 辅助
目的:胸部X光(CXR)扫描数量增多造成了分诊瓶颈,使紧急检查排队在常规检查之后。现有AI工具主要是缺乏严重程度感知的单模态二分类器,且多模态系统很少与放射科专家进行基准测试。为此,我们开发了一种多模态深度学习框架,用于联合严重程度分诊、病理检测及原生视觉解释。方法:我们提出跨模态分诊网络(CMTN),通过门控交叉注意力将Swin Transformer V2视觉编码器与PubMedBERT文本编码器融合。CMTN在来自MIMIC-CXR-JPG的34639组图像-文本对(12489名患者)上进行训练,优化序数焦点损失以进行四级严重程度分诊,优化二元交叉熵以检测14种病理。除定量基准测试外,还在两阶段临床审核中对注意力热图与盲法放射科专家进行评估:比较模型分诊输出与专家严重程度评估(100例),以及对空间-语义一致性进行分级(116张热图)。结果:CMTN与参考标签实现了强序数一致性(二次加权kappa[QWK]=0.9341,95%置信区间:0.9219至0.9449),14种病理的宏平均AUROC为0.9970,延迟为34毫秒,优于最先进的BioViL多模态基准(QWK=0.7679)。但盲法第一阶段临床审核显示其与真实放射科医生判断的一致性显著较低(QWK=0.1399)。第二阶段发现54.3%的热图达到临床可接受的空间定位。结论:CMTN展现了用于CXR分诊的高效多模态架构。算法与放射科医生判断的一致性差异表明,基于NLP衍生标签的基准测试性能不足,凸显临床部署前需放射科医生标注的真实值。
英文摘要
Purpose: Increased number of chest radiograph (CXR) scans create a triage bottleneck, queueing urgent examinations behind routine ones. Existing AI tools are predominantly unimodal binary classifiers lacking severity awareness, and multimodal systems are rarely benchmarked against expert radiologists. To this end, we developed a multimodal deep learning framework for joint severity triage, pathology detection, and native visual explanation. Approach: We propose the cross-modal triage network (CMTN), fusing a Swin Transformer V2 visual encoder with a PubMedBERT text encoder via gated cross-attention. The CMTN was trained on 34,639 image-text pairs (12,489 patients) from MIMIC-CXR-JPG, optimizing an ordinal focal loss for four-tier severity triage and binary cross-entropy for 14 pathologies. Beyond quantitative benchmarking, attention heatmaps were evaluated against a blinded expert radiologist in a two-phase clinical audit comparing model triage output to expert severity assessment (100 cases) and grading spatial-semantic concordance (116 heatmaps). Results: The CMTN achieved strong ordinal agreement with reference labels (quadratic weighted kappa [QWK] = 0.9341, 95\% CI: 0.9219 to 0.9449) and macro-AUROC of 0.9970 across 14 pathologies, with 34~ms latency, outperforming the state-of-the-art BioViL multimodal baseline (QWK = 0.7679). However, the blinded Phase I clinical audit revealed substantially lower agreement with genuine radiologist judgment (QWK = 0.1399). Phase II found 54.3\% of heatmaps achieved clinically acceptable spatial localization. Conclusions: The CMTN demonstrated an efficient multimodal architecture for CXR triage. The divergence between algorithmic and radiologist agreement demonstrates that benchmark performance against NLP-derived labels is insufficient, highlighting the need for radiologist-labeled ground truth before clinical deployment.
发表机构
- University of Guelph(圭尔夫大学)
- Guelph General Hospital(圭尔夫综合医院)
机构由 AI 辅助整理,请以论文原文为准。