arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进黑盒放射学AI的校准:使用测试时增强

Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation

Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang, David B. Larson, Akshay S. Chaudhari, Camila Gonzalez

arXiv 2609.29931首次发表:更新:

发表机构

Stanford University; Medical University of Vienna(斯坦福大学; 维也纳医科大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种模型无关的测试时增强框架,仅通过输入输出访问改进黑盒放射学AI的校准,在肺栓塞和颅内出血任务中显著降低校准误差,优于需访问模型内部的方法。

AI 中文摘要

放射学AI系统越来越多地为临床决策提供信息,如分诊、后续影像检查和治疗计划。为了安全地做出这些决策,模型输出必须得到良好校准,即预测概率准确反映真实风险。许多标准的校准改进技术,如MC Dropout和深度集成,需要访问模型参数或重新训练。然而,专有的临床AI系统以黑盒形式运行,阻止访问模型内部。为此,我们提出了一个模型无关的框架,使用临床基础的测试时增强(TTA)来改进黑盒模型的校准。我们的框架应用几何和物理启发的3D CT扰动,并在不访问模型内部或原始训练数据的情况下学习概率级聚合策略。在肺栓塞和颅内出血检测任务中,DualTTA在TTA方法中实现了最强的整体校准,分别将预期校准误差降低了54%(0.239 -> 0.109)和43%(0.051 -> 0.029),同时仅需要输入输出访问。此外,DualTTA在大多数校准指标上优于需要访问模型内部的不确定性估计技术,如温度缩放、MC Dropout和深度集成。这些结果表明,学习的TTA聚合可以改进临床AI系统的校准,为改进黑盒医学AI的可靠性提供了一种实用方法。

英文摘要

Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.

Comments11 pages, 3 figures, 1 table. Accepted at the MICCAI 2026 Workshop on Uncertainty for Safe Utilization of Machine Learning in Medical Imaging (UNSURE 2026). Code: https://github.com/stanfordaide/TTA_Calibration

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑