发表机构
Ho Chi Minh City University of Technology; Vietnam National University Ho Chi Minh City; AK Technologies Company Limited(胡志明市理工大学; 越南国立大学胡志明市分校; AK科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出融合RAD-DINO单模态与BioViL-T视觉-语言表示的多标签胸部X射线分类框架,在MIMIC-CXR-JPG上实现AUROC 0.840,证明混合融合优于早期融合。
AI 中文摘要
多标签胸部X射线分类近年来引起了广泛关注,其中视觉表示和临床语义知识的有效利用发挥着重要作用。本研究提出一个框架,将来自RAD-DINO的单模态表示与来自BioViL-T的视觉-语言表示相结合,用于MIMIC-CXR-JPG数据集中14个标签的分类。RAD-DINO和BioViL-T的嵌入及其组合表示在潜在空间中分别进行细化,然后跨三个分支进行归一化和融合。除了提高分类性能外,本研究还旨在阐明每个嵌入源的作用及其互补程度。实验表明,RAD-DINO在独立使用时优于BioViL-T,而早期融合进一步改善了结果,表明这两个嵌入源包含互补信息。最佳模型实现了平均AUROC为0.840,mAP为0.467。消融分析表明,当每个嵌入源在潜在空间中被细化时,混合融合相对于早期融合提供了一致且统计显著的改进,这表明融合效果取决于每个分支提供的表示质量。然而,该研究仅在MIMIC-CXR-JPG上进行了内部评估,其对其他医疗机构数据的泛化性仍有待验证。源代码可在以下网址获取:this https URL。
英文摘要
Multi-label chest X-ray classification has attracted considerable attention in recent years, with the effective use of visual representations and clinical semantic knowledge playing an important role. This study proposes a framework that combines unimodal representations from RAD-DINO with vision--language representations from BioViL-T for the classification of 14 labels in the MIMIC-CXR-JPG dataset. The RAD-DINO and BioViL-T embeddings and their combined representation are refined separately in latent space before being normalized and fused across the three branches. In addition to improving classification performance, the study aims to clarify the role of each embedding source and the degree to which they complement one another. Experiments show that RAD-DINO outperforms BioViL-T when used independently, whereas early fusion further improves the results, indicating that the two embedding sources contain complementary information. The best-performing model achieves a mean AUROC of 0.840 and an mAP of 0.467. Ablation analysis shows that hybrid fusion provides consistent and statistically significant improvements over early fusion when each embedding source is refined in latent space, suggesting that fusion effectiveness depends on the quality of the representation supplied by each branch. However, the study has only been evaluated internally on MIMIC-CXR-JPG; its generalizability to data from other healthcare institutions therefore remains to be validated. The source code is available at: https://anonymous.4open.science/r/mimic-report-C210/.
Comments10 pages, 2 figures, 5 tables (main text); 12 pages, 1 figure, 13 tables (supplementary material)