arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.15645cs.CL

迈向可靠的医疗大语言模型:在真实医疗咨询中评估和增强大语言模型信心估计的基准测试

Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation

  • Nanyang Technological University(南洋理工大学)
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhiyao Ren, Yibing Zhan, Siyuan Liang, Guozheng Ma, Baosheng Yu, Dacheng Tao

AI总结:

本文提出MedConf框架,通过证据引导的语言自评估方法,在医疗咨询中提升大语言模型的信心估计精度与可靠性。

AI中文摘要:

大规模语言模型(LLMs)通常基于不完整的信息提供临床判断,增加了误诊的风险。现有研究主要在单轮、静态环境中评估信心,忽略了在真实咨询中随着临床证据积累,信心与正确性之间的耦合关系,这限制了它们对可靠决策的支持。我们提出了第一个评估多轮交互中真实医疗咨询中信心的基准测试。我们的基准测试统一了三种医疗数据以进行开放性诊断生成,并引入了信息充分性梯度来表征随着证据增加时信心与正确性动态变化。我们在此基准测试上实现了并比较了27种代表性方法;两项关键见解出现:(1)医疗数据放大了token级和一致性级信心方法的固有局限性;(2)医疗推理必须同时评估诊断准确性和信息完整性。基于这些见解,我们提出了MedConf,一个基于证据的语言自我评估框架,通过检索增强生成构建症状档案,将患者信息与支持、缺失和矛盾关系对齐,并通过加权整合将它们整合成可解释的信心估计。在两个LLM和三个医疗数据集中,MedConf在AUROC和皮尔逊相关系数指标上均优于现有最先进方法,在信息不足和多病共存条件下保持稳定性能。这些结果表明,信息充分性是可信医疗信心建模的关键决定因素,为构建更可靠和可解释的大医疗模型提供了新的途径。

英文摘要:

Large-scale language models (LLMs) often offer clinical judgments based on incomplete information, increasing the risk of misdiagnosis. Existing studies have primarily evaluated confidence in single-turn, static settings, overlooking the coupling between confidence and correctness as clinical evidence accumulates during real consultations, which limits their support for reliable decision-making. We propose the first benchmark for assessing confidence in multi-turn interaction during realistic medical consultations. Our benchmark unifies three types of medical data for open-ended diagnostic generation and introduces an information sufficiency gradient to characterize the confidence-correctness dynamics as evidence increases. We implement and compare 27 representative methods on this benchmark; two key insights emerge: (1) medical data amplifies the inherent limitations of token-level and consistency-level confidence methods, and (2) medical reasoning must be evaluated for both diagnostic accuracy and information completeness. Based on these insights, we present MedConf, an evidence-grounded linguistic self-assessment framework that constructs symptom profiles via retrieval-augmented generation, aligns patient information with supporting, missing, and contradictory relations, and aggregates them into an interpretable confidence estimate through weighted integration. Across two LLMs and three medical datasets, MedConf consistently outperforms state-of-the-art methods on both AUROC and Pearson correlation coefficient metrics, maintaining stable performance under conditions of information insufficiency and multimorbidity. These results demonstrate that information adequacy is a key determinant of credible medical confidence modeling, providing a new pathway toward building more reliable and interpretable large medical models.

↑