发表机构
School of Computing, Macquarie University(麦考瑞大学计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对临床信息不确定性下的LLMs开展行为分析,提出基于MedMCQA的评估框架,发现模型置信度与准确率不匹配、弃权能力差异大等关键局限,凸显需采用不确定性感知评估方法。
AI 中文摘要
大语言模型(LLMs)在医学问答与临床推理任务中已展现出优异性能,但其在不确定性场景下的可靠性仍未得到充分理解,这引发了对其在高风险临床场景中部署的关键担忧。在这类环境中,错误预测本身就存在风险,而带有高置信度的错误预测危害尤甚,可能会误导临床决策。本文对临床信息不确定性下的LLMs开展系统性行为分析,提出了基于MedMCQA数据集的评估框架,包含两种互补的不确定性设置:其一,通过提示修改引入语言不确定性线索,以模拟模糊的临床场景;其二,构建答案移除设置,即刻意排除正确选项,要求模型识别信息不足并弃权(不执行)。我们针对500道医学问题,采用校准间隙、预期校准误差(ECE)、不安全置信错误率(UCER)等多种校准指标,分析模型的准确率与置信度行为。结果显示存在一致的失效模式:尽管准确率随不确定性增加而下降,但模型置信度与准确率仍不匹配,导致不安全置信错误大幅增加,表明模型置信度对临床有意义的信息缺失基本不敏感。此外,我们观察到不同模型在正确答案不可用时的弃权能力存在显著差异,部分模型仍持续生成高置信度的幻觉答案。这些发现暴露了当前LLMs认知可靠性的关键局限,强调需在其部署至临床工作流前采用不确定性感知的评估方法。
英文摘要
Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.