面向多中心乳腺MRI图像数据集质量保证的无监督异常检测
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
浏览论文内容
中文总结 AI 辅助
该研究针对多中心乳腺MRI数据集,构建含17种异常类型的无监督异常检测基准,评估四种方法,发现带位置编码的投影法性能最优,为医疗AI的可扩展无监督质量保证提供基础。
中文摘要 AI 辅助
损坏、不一致或异常数据会悄然威胁医疗AI的安全性与可靠性。尽管高风险医疗AI的数据集质量保证(QA)日益受到监管认可,可扩展的自动化检测仍有待开发。我们采用无监督异常检测(AD)与分布外(OOD)检测作为多中心动态增强乳腺MRI的自动化数据集QA机制。我们基于六个公开数据集构建了包含17种与QA相关的现实异常类型(包括协议违规、处理错误、解剖区域错误)的受控AD基准,并提出了基于人类视觉感知的放射学图像异常分类,实现对AD失效模式的细粒度分析。该基准包含近、中-远、远OOD样本,以及分布内和外部正常数据。我们评估了四种方法:一种扩展了领域特定特征提取器与新型位置编码的基于投影的方法,一种扩展至完整3D体积并采用增强训练目标的基于重构的方法,以及两种未修改的混合OOD检测方法。中-远与远OOD样本被可靠检测,而近OOD样本与来自未见过机构的外部正常数据则暴露了方法间的特定差异。基于3D重构的方法在检测性能(AUROC:0.936)与对未见过机构的泛化间实现了最佳平衡;带位置编码的基于投影的方法取得了最高的整体检测性能(AUROC:0.954)。两种混合方法均表现出关键失效模式,证实针对某一模态或解剖结构验证的方法若不进行领域特定适配则无法泛化。所有方法均对植入物与乳房切除术仍存在挑战。我们的研究结果为医疗AI流程中可扩展的无监督QA奠定了基础并提供了实用指导。
英文摘要
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.
发表机构
- Cambridge University Hospitals(剑桥大学医院)
- Mitera Hospital(米特拉医院)
- Radboud University Medical Center(拉德堡德大学医学中心)
- University Hospital Aachen(亚琛大学医院)
- University Medical Center Utrecht(乌得勒支大学医学中心)
- Ribera Hospital(里贝拉医院)
- Duke Breast Cancer MRI(杜克乳腺癌磁共振成像项目)
机构由 AI 辅助整理,请以论文原文为准。