ClinicalGPT-R1:利用大语言模型提升全科疾病诊断的推理能力
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ClinicalGPT-R1,经20000条临床记录训练,在MedBench-Hard基准测试中,其中文诊断性能优于GPT-4o、英文性能与GPT-4相当,提升了疾病诊断的推理能力。
AI中文摘要:
近期大语言模型(LLMs)在数学、编码等领域展现出出色的推理能力,但其在临床诊断中的应用仍未得到充分探索。本文提出ClinicalGPT-R1,这是一款用于疾病诊断的推理增强型全科大语言模型,基于20000条真实临床记录的数据集训练,采用多种训练策略提升诊断推理能力。为评估性能,本文构建了涵盖7个主要医学专科及代表性疾病的挑战性数据集MedBench-Hard。实验结果表明,ClinicalGPT-R1在中文诊断任务中表现优于GPT-4o,在英文场景中性能与GPT-4相当,该对比研究有效验证了ClinicalGPT-R1在疾病诊断任务中的优越性能,相关资源可通过指定URL获取。
英文摘要:
Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here, we introduce ClinicalGPT-R1, a reasoning enhanced generalist large language model for disease diagnosis. Trained on a dataset of 20,000 real-world clinical records, ClinicalGPT-R1 leverages diverse training strategies to enhance diagnostic reasoning. To benchmark performance, we curated MedBench-Hard, a challenging dataset spanning seven major medical specialties and representative diseases. Experimental results demonstrate that ClinicalGPT-R1 outperforms GPT-4o in Chinese diagnostic tasks and achieves comparable performance to GPT-4 in English settings. This comparative study effectively validates the superior performance of ClinicalGPT-R1 in disease diagnosis tasks. Resources are available at https://github.com/medfound/medfound.