日本卒中大语言模型评估:使用大语言模型进行日本安全卒中护理的多轮对话基准
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
- National Hospital Organization Kagoshima Medical Center(国立医院机构鹿儿岛医疗中心)
- Imamura General Hospital(今村综合医院)
- Iwamoto Neurosurgery(岩本神经外科)
- Kagoshima University(鹿儿岛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对日本卒中护理,构建多轮对话基准,评估18个大语言模型的安全性与性能,发现顶尖模型达标,但多个模型存在关键错误,并指出病史采集问题数与得分正相关。
AI中文摘要:
背景:大语言模型(LLMs)在多项选择医学知识考试中已达到与医生相当的性能,但其在临床病史采集、紧急程度评估和安全性方面的能力仍未得到充分评估。我们提出了日本卒中大语言模型评估(Japanese Stroke LLM Evaluation),这是一个针对日本卒中护理的多轮对话基准,并在实践导向条件下评估了大语言模型的性能和安全性。方法:我们创建了10个卒中及相关病症病例,并在多轮日语对话中评估大语言模型。大语言模型扮演医生角色,而一名委员会认证的神经外科医生扮演模拟患者和评估者。每个病例包含病史采集和行动阶段,使用预先指定的标准进行评分。可能直接威胁生命的错误被定义为关键错误。安全阈值为总体得分至少80%且零关键错误。2025年10月和2026年6月共评估了18个模型。结果:Claude Fable 5获得最高分(87.4%),且零关键错误,其次是Claude Opus 4.7(80.3%)和GLM-5.2(75.6%)。两个领先模型达到了安全阈值。11个模型出现了17个关键错误,包括在t-PA给药前未确认实验室结果或血糖、在气道稳定前进行手术、遗漏颈血管评估,以及在t-PA适应症外使用。病史采集问题数量与病史采集得分相关(r = 0.648,p = 0.007)。结论:日本卒中大语言模型评估为实践导向条件下的大语言模型性能提供了一个基准,包括对病史采集问题的数量上限。病例和评估由神经外科专家创建,而非使用大语言模型作为评判者。2026年,云端和本地部署模型的性能均有所提升,部分模型超过了安全阈值。需要使用真实世界病例进行进一步评估。
英文摘要:
Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions. Methods: We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator. Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes. The safety threshold was at least 80% overall with zero critical mistakes. Eighteen models were evaluated in October 2025 and June 2026. Results: Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%). Two leaders met the safety threshold. Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication. History-taking question count correlated with history-taking score (r = 0.648, p = 0.007). Conclusions: Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions. Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach. Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold. Further evaluation using real-world cases is required.