AI Hospital:在多智能体医疗交互模拟器中评测大语言模型
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
- Alibaba Inc(阿里巴巴公司)
- Huazhong University of Science Technology(华中科技大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出多智能体医疗交互模拟框架 AI Hospital 和 MVME 基准,用于评估 LLMs 在症状采集、检查建议和诊断中的多轮临床交互能力,并引入争议解决协作机制提升诊断准确性。
AI中文摘要:
人工智能已显著推动医疗健康领域发展,尤其是大语言模型(LLMs)在医学问答基准上表现优异。然而,由于医患交互的复杂性,其真实临床落地仍受限。为此,我们提出 AI Hospital,一个多智能体框架,模拟以医生(Doctor)为玩家、患者(Patient)、考官(Examiner)、主任医师(Chief Physician)等 NPC 参与的动态医疗交互,从而在临床场景中对 LLMs 进行更真实的评估。我们构建了多视角医学评估(Multi-View Medical Evaluation,MVME)基准,利用高质量中文病历和 NPC 评估 LLMs 在症状采集、检查建议和诊断方面的表现。此外,提出一种争议解决协作机制,通过迭代讨论提升诊断准确性。尽管有所改进,当前 LLMs 在多轮交互中相比单步方法仍存在显著性能差距。我们的发现表明需要进一步研究以弥合这些差距并提升 LLMs 的临床诊断能力。我们的数据、代码和实验结果已在 https://github.com/LibertFan/AI_Hospital 全部开源。
英文摘要:
Artificial intelligence has significantly advanced healthcare, particularly through large language models (LLMs) that excel in medical question answering benchmarks. However, their real-world clinical application remains limited due to the complexities of doctor-patient interactions. To address this, we introduce \textbf{AI Hospital}, a multi-agent framework simulating dynamic medical interactions between \emph{Doctor} as player and NPCs including \emph{Patient}, \emph{Examiner}, \emph{Chief Physician}. This setup allows for realistic assessments of LLMs in clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and NPCs to evaluate LLMs' performance in symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance diagnostic accuracy through iterative discussions. Despite improvements, current LLMs exhibit significant performance gaps in multi-turn interactions compared to one-step approaches. Our findings highlight the need for further research to bridge these gaps and improve LLMs' clinical diagnostic capabilities. Our data, code, and experimental results are all open-sourced at \url{https://github.com/LibertFan/AI_Hospital}.