发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在真实医学影像数据上验证两阶段延迟学习,并提出带引导的L2D扩展决策空间,通过更简单的损失函数在多个数据集上超越经典方法及人类/AI基线。
AI 中文摘要
医学图像解读工作量大且耗时,虽然人工智能解读可以减轻工作量,但完全自主部署存在潜在的安全隐患,且低特异性在实践中可能导致临床医生工作量增加。延迟学习(L2D)通过从输入特征、AI模型和人类表现中学习,在自主预测和人类专家之间选择性路由病例来解决这一问题。尽管L2D已有理论保证,但其性能尚未在带有人类阅片者标注的真实世界医学数据集上得到验证。我们评估了两阶段L2D的预测器-拒绝器公式,其中AI预测器模型固定,与可训练的路由或拒绝器模型分离,并在Collab-CXR(一个每个病例具有多个人类标注的多标签胸部X射线数据集)上进行评估。这是首个在带有真实世界医学影像数据和人类标注的背景下研究L2D的工作。我们进一步引入了一种新设置,即带引导的L2D,其中决策空间扩展为三个选择:自主预测、延迟给人类专家、或延迟给人类专家并提供AI引导。我们比较了多种拒绝器架构、损失函数以及不同的输入特征可用性。这在两个更大的数据集VinDr-CXR和CheXpert上进行了复现。我们的结果表明,带引导的两阶段L2D优于经典的两阶段延迟学习,以及仅人类、仅AI和AI引导的人类基线。值得注意的是,与当前文献中正式定义的L2D替代损失函数相比,这一性能是通过更简单的损失函数实现的。
英文摘要
Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries potential safety concerns and low specificity may in practice lead to increased clinician workload. Learning to Defer (L2D) addresses this by selectively routing cases between autonomous prediction and human experts by learning from input features and AI model and human performance. While theoretical guarantees have been proven for L2D, its performance has not been validated on real-world medical datasets with human reader annotations. We evaluate the predictor-rejector formulation of two-stage L2D, where the AI predictor model is fixed and separate from the trainable routing or rejector model, on Collab-CXR, a multilabel chest X-ray dataset with multiple human annotations per case. This is the first work to look at L2D in the context of real-world medical imaging data with human annotations. We further introduce a new setup, L2D with Guidance, where the decision space is extended to three choices: predict autonomously, defer to a human expert, or defer to a human expert and provide AI guidance. We compare multiple rejector architectures and loss functions, and different input feature availabilities. This is reproduced on two larger datasets, VinDr-CXR and CheXpert. Our results show that two-stage L2D with Guidance outperforms classic two-stage learning to defer, as well as human-alone, AI-alone and AI-guided human baselines. Notably, this performance is achieved with simpler loss functions compared to formally defined L2D surrogate loss functions in current literature.
CommentsAccepted at HAIC workshop, MICCAI 2026