arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20582cs.LGcs.AIcs.MA

贝叶斯不确定性估计改善医学人工智能代理中的临床决策

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

Frederik Hauke, Patrick Wienholt, Christiane Kuhl, Dyke Ferber, Jakob Nikolas Kather, Sven Nebelung, Daniel Truhn

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对医学图像分析机器学习模型缺乏置信度度量问题,采用蒙特卡洛随机失活获取认知不确定性信号,添加到点预测提高错误检测AUROC,在临床决策支持代理实验中,以特定形式传达该信号可降低误诊率,凸显其对决策的价值。

中文摘要 AI 辅助

用于医学图像分析的机器学习模型通常缺乏可靠的置信度度量,限制了它们在模糊或非典型病例中的应用。本文表明,应用于多任务胸部X光分类器(八个胸部发现,137,593张训练图像)的蒙特卡洛随机失活提供了一个认知不确定性信号,该信号跟踪训练集规模上的泛化情况,并标记出自信但容易出错的预测。将此信号添加到点预测中,使错误检测的AUROC从0.74提高到0.77(ΔAUROC +0.023,95% CI [+0.014, +0.033])。在一个受控的2x2析因实验中,一个临床决策支持代理仅在以二元错误风险标志而非原始分数形式提供不确定性时利用了该不确定性,将对不可靠发现的自信误诊从8.5%降至2.7%。因此,认知不确定性估计携带了超出点预测的与决策相关的信息,但其对下游代理的价值取决于其传达方式。

英文摘要

Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 ($Δ$AUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.

发表机构

  • University Hospital RWTH Aachen(亚琛工业大学附属医院)
  • Heidelberg University Hospital(海德堡大学医院)
  • TU Dresden(德累斯顿工业大学)
  • TUD Dresden University of Technology(德累斯顿工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑