arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在医学推理中表现出元认知敏感性

Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

Ahmad Nazzal

arXiv 2608.14552首次发表:更新:

AI 中文总结

本研究开发临床基准测试医学LLM,发现其在AT-NCD与DRCI鉴别中具部分元认知敏感性,错误集中于冲突病例,建立了医学LLM相关评估框架。

AI 中文摘要

大型语言模型(LLM)在医学领域的评估与应用日益增多,但其临床实用性取决于答案准确性,以及模型的置信度是否能匹配证据质量与不确定性。我们开发了一种受心理物理学启发的可控临床基准,用于测试医学LLM的诊断选择与置信度行为,该基准聚焦于阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关认知障碍(DRCI)的鉴别。我们生成了45个合成 vignette(情景案例),其证据强度、冲突证据及缺失信息各不相同,每个情景以三种提示变体呈现,共产生135次试验。在使用gpt-4.1-nano的预实验中,所有试验均生成了有效的结构化输出。在强制选择试验中,诊断准确率为93.5%,平均置信度为78.4%,AUROC2为0.876。置信度随证据距诊断边界的距离增加而升高,随信息缺失而降低,且在调整证据强度与提示格式后,正确试验的置信度仍高于错误试验。这些发现表明模型存在部分元认知敏感性,而非全局无意义的置信度;不过错误集中在中等、存在冲突的AT-NCD病例中,此时模型偏向DRCI,且置信度高于经验准确性应有的水平。模型比较显示,置信度质量应直接测量,而非仅从基准准确率或模型能力推断。本研究建立了可重复的框架,用于评估医学LLM的证据敏感性、元认知敏感性及局部校准失效问题。

英文摘要

Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑