发表机构
Medical Imaging and AI in Healthcare Research Group, Biosciences Center, SURA; Universidad de Antioquia; IBM Research; Neuroradiology Institute, Universitätsklinikum Erlangen(SURA生物科学中心医疗成像与医疗保健人工智能研究组; 安蒂奥基亚大学; IBM研究院; 埃尔兰根大学医院神经放射学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对胸部X光仓库标签非图像级真值的问题,提出RSA框架,以MIMIC-CXR心脏肥大数据为例发现其标签与专家评估一致性极低,用专家精选队列训练的DenseNet121模型达0.853的ROC-AUC,强调监督审计的重要性。
AI 中文摘要
公开胸部X光仓库被广泛用于训练医学AI系统,然而其标签通常从放射学报告中提取,而非直接在图像上验证。因此,仓库标签常被当作图像级真值,却未验证其是否反映X光片实际可见内容。我们提出仓库监督审计(Repository Supervision Auditing, RSA)框架,在模型开发前依据专家图像级注释评估仓库衍生标签。以MIMIC-CXR中的心脏肥大为例,RSA对比仓库标签与放射科医生审阅后的图像注释,刻画分歧来源,并构建用于部署导向评估的精选队列。仓库衍生的心脏肥大标签与专家图像级评估的一致性近乎为零,仅识别出1%的专家确认病例。多数分歧源于报告未提及而非明确否定,在被仓库衍生标签判定为“无发现”的研究中,近一半存在专家确认的心脏肥大。利用该专家精选队列,DenseNet121模型达到0.853的测试ROC-AUC。这些发现表明,仓库标签可能无法可靠代表图像级真值,凸显监督审计是开发可信医学影像AI的关键步骤。
英文摘要
Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather than verified directly on images. As a result, repository labels are often treated as image-level ground truth without validating whether they reflect what is actually visible in the radiograph. We introduce Repository Supervision Auditing (RSA), a framework that evaluates repository-derived labels against expert image-level annotations before model development. Using cardiomegaly in MIMIC-CXR as a case study, RSA compares repository labels with radiologist-reviewed image annotations, characterizes disagreement sources, and builds a curated cohort for deployment-oriented evaluation. Repository-derived cardiomegaly labels showed near-zero agreement with expert image-level assessment, identifying only 1% of expert-confirmed cases. Most discrepancies resulted from non-mention rather than explicit report negation, with expert-confirmed cardiomegaly identified in nearly half of studies assigned a repository-derived No Finding label. Using the resulting expert-curated cohort, a DenseNet121 model achieved a test ROC-AUC of 0.853. These findings show that repository labels may not reliably represent image-level truth and highlight supervision auditing as a critical step for developing trustworthy medical imaging AI.
CommentsAccepted (oral) at the 3rd MICCAI Student Board (MSB) EMERGE Workshop, MICCAI 2026. This is the authors' version; the final authenticated version will appear in Springer Lecture Notes in Computer Science (LNCS). 10 pages, 3 figures