arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.14107cs.CLcs.AI

DiagnosisArena:面向大型语言模型的诊断推理基准测试

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

  • Shanghai Jiao Tong University(上海交通大学)
  • SII
  • SPIRAL Lab(SPIRAL实验室)
  • Generative AI Research Lab (GAIR)(生成式AI研究实验室)
  • Shanghai Chest Hospital(上海胸科医院)
  • Beijing Anzhen Hospital, Capital Medical University(北京安贞医院,首都医科大学)

机构由 AI 辅助整理,请以论文原文为准。

Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang

AI总结:

该研究提出DiagnosisArena基准,评估大型语言模型的临床诊断推理能力,发现顶尖模型准确率仅最高51.12%,凸显其泛化瓶颈,旨在推动AI诊断推理进步。

AI中文摘要:

具备复杂推理能力的突破性大型语言模型的出现,为解决各类科学挑战(包括复杂临床场景中的挑战)带来了重大希望。为了使其能够在现实医疗环境中安全、有效地部署,迫切需要系统地对当前模型的诊断能力进行基准测试。鉴于现有医疗基准在评估高级诊断推理方面存在局限性,我们提出了DiagnosisArena,这是一个旨在严格评估专业水平诊断能力的综合性、高难度基准。DiagnosisArena包含1113组分段患者病例及对应诊断,覆盖28个医学专科,源自10种顶级医学期刊发表的临床病例报告。该基准通过精细的构建流程开发,涉及AI系统与人类专家的多轮筛选和审查,并进行了全面检查以防止数据泄露。我们的研究显示,即使是最先进的推理模型o3、o1和DeepSeek-R1,准确率也仅分别达到51.12%、31.09%和17.79%。这一发现凸显了当前大型语言模型在面对临床诊断推理挑战时存在显著的泛化瓶颈。我们希望通过DiagnosisArena推动AI诊断推理能力的进一步提升,为解决现实世界的临床诊断挑战提供更有效的解决方案。我们提供该基准及评估工具以支持进一步的研究与开发,网址为this https URL。

英文摘要:

The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable their safe and effective deployment in real-world healthcare settings, it is urgently necessary to benchmark the diagnostic capabilities of current models systematically. Given the limitations of existing medical benchmarks in evaluating advanced diagnostic reasoning, we present DiagnosisArena, a comprehensive and challenging benchmark designed to rigorously assess professional-level diagnostic competence. DiagnosisArena consists of 1,113 pairs of segmented patient cases and corresponding diagnoses, spanning 28 medical specialties, deriving from clinical case reports published in 10 top-tier medical journals. The benchmark is developed through a meticulous construction pipeline, involving multiple rounds of screening and review by both AI systems and human experts, with thorough checks conducted to prevent data leakage. Our study reveals that even the most advanced reasoning models, o3, o1, and DeepSeek-R1, achieve only 51.12%, 31.09%, and 17.79% accuracy, respectively. This finding highlights a significant generalization bottleneck in current large language models when faced with clinical diagnostic reasoning challenges. Through DiagnosisArena, we aim to drive further advancements in AI's diagnostic reasoning capabilities, enabling more effective solutions for real-world clinical diagnostic challenges. We provide the benchmark and evaluation tools for further research and development https://github.com/SPIRAL-MED/DiagnosisArena.

补充信息

↑