发表机构
University of Cambridge; Medical University of Vienna; Cambridge University Hospitals NHS Foundation Trust; King's College London(剑桥大学; 维也纳医科大学; 剑桥大学医院 NHS 基金会信托; 伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对胸部X射线机器学习,探究评估参考选择对模型性能与排名的影响,发现常用指标与临床判断一致性差,强调需将评估参考选择作为临床有效性的核心部分。
AI 中文摘要
胸部X射线(CXR)机器学习严重依赖旨在近似临床判断的自动化评估,但常用的报告衍生病理分类标签或通用图像质量重建指标可能无法可靠反映临床判断。我们系统研究了病理分类和图像质量评估(IQA)中评估参考选择如何影响模型性能与排名。为实现评估参考间的可控比较,我们从剑桥大学医院(CUH)临床队列收集了配对的专家图像及报告衍生的胸部发现标签,整理了公开MIMIC-CXR数据集的子集,还收集了专家对诊断图像质量的评分。我们发现,对于监督图像分类器(ResNet、DenseNet)及多种零样本、微调的视觉语言模型(如MedKLIP、GLoRIA、ConVIRT),标签来源的变化不仅会导致性能估计出现显著差异,还会改变模型排名。同时,IQA指标与临床判断的一致性高度依赖指标选择,常用的IQA指标如SSIM和PSNR常无法与专家对诊断可用性的评估达成一致。我们的结果表明,评估选择至关重要:它们可决定哪些模型和方法表现最佳,进而被选中进行进一步开发或部署。因此,评估参考的选择应被视为CXR机器学习临床有效性的核心组成部分,并需结合病理、成像任务及预期下游临床使用场景进行论证。
英文摘要
Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.
Comments67 pages, including supplementary material