arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.10405cs.CLcs.AIcs.ETcs.HC

大型语言模型在支持医疗诊断与治疗中的性能

Performance of Large Language Models in Supporting Medical Diagnosis and Treatment

  • FEUP - Faculty of Engineering of the University of Porto(波尔图大学工程学院(FEUP))
  • INEGI – Institute of Science and Innovation in Mechanical and Industrial Engineering(机械与工业工程科学与创新研究所(INEGI))

机构由 AI 辅助整理,请以论文原文为准。

Diogo Sousa, Guilherme Barbosa, Catarina Rocha, Dulce Oliveira

更新

AI总结:

本研究评估多款开源和闭源大型语言模型(LLMs)在2024年葡萄牙国家医学专科准入考试(PNA)中的性能,发现部分模型准确性超医学生基准,还分析了模型的成本效益及思维链等推理方法的影响,为LLMs辅助医疗决策提供参考。

AI中文摘要:

大型语言模型(LLMs)与医疗保健的融合在提升诊断准确性和支持治疗规划方面具有重大潜力。这些AI驱动的系统可分析海量数据集,协助临床医生识别疾病、推荐治疗方案并预测患者结局。本研究评估了一系列当代LLMs(包括开源和闭源模型)在2024年葡萄牙国家医学专科准入考试(PNA)——一项标准化医学知识评估——中的性能。结果显示,模型在准确性和成本效益上存在显著差异,部分模型在该特定任务上的表现超过了医学生的人类基准。我们基于准确性和成本的综合评分确定了领先模型,讨论了思维链(Chain-of-Thought)等推理方法的影响,并强调LLMs有望成为辅助医疗专业人员进行复杂临床决策的宝贵补充工具。

英文摘要:

The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making.

补充信息

↑