arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15180cs.LG

重新思考临床预测中视觉-语言模型不确定性估计的正确性

Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models

Mingcheng Zhu, Jinning Liang, Tingting Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出两轴框架评估临床预测中视觉-语言模型不确定性估计的正确性标准,发现精确匹配最优,且标准选择显著影响UE评估结果。

中文摘要 AI 辅助

视觉-语言模型越来越多地被用于从电子健康记录和医学图像中进行临床预测,在这些场景中,识别不可靠的预测对于安全部署至关重要。不确定性估计(UE)能够检测此类预测,但其评估依赖于一个正确性标准,该标准决定每个模型输出是否正确。如果该标准与人类判断不一致或扭曲了下游UE性能,那么关于模型可靠性的结论可能会产生误导。我们引入了一个两轴框架,通过正确性标准与人类判断的一致性及其对人类参考的UE性能的保真度来评估这些标准。我们使用由两位评审者标注的450个预测,在三个临床预测任务和三个模型上评估了八个标准。在审计的任务中,规范精确匹配(EM)达到了最高的人类一致性和最低的UE失真,而基于BERT的匹配(BEM)和LLM评判者也显示出较强的人类一致性。在四种UE方法和23,254个临床预测中,标准的选择使错误检测AUROC最多变化0.146,并逆转了UE方法的相对排名。LLM评判者还选择性地接受了无效或不确定的输出,接受了30个人类识别错误中的16个。这些结果表明,正确性评估是临床UE评估的一个组成部分,在比较UE方法之前应进行验证。

英文摘要

Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion that determines whether each model output is correct. If this criterion disagrees with human judgement or distorts downstream UE performance, conclusions about model reliability can be misleading. We introduce a two-axis framework that evaluates correctness criteria by their agreement with human judgements and fidelity to human-referenced UE performance. We assess eight criteria across three clinical prediction tasks and three models using 450 predictions annotated by two reviewers. Across the audited tasks, canonical exact matching (EM) achieved the highest observed human agreement and lowest UE distortion, while the BERT-based matching (BEM) and LLM-judge also showed strong human agreement. Across four UE methods and 23,254 clinical predictions, criterion choice changed error-detection AUROC by up to 0.146 and reversed the relative ranking of UE methods. The LLM-judge also selectively accepted invalid or uncertain outputs, accepting 16 of 30 such human-identified errors. These results demonstrate that correctness assessment is an integral component of clinical UE evaluation and should be validated before UE methods are compared.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑