发表机构
University of Missouri; Community Health Service Center Shanghai Pudong New Area; Jinsha County Chinese Medicine Hospital; Shanghai Seventh People's Hospital; The First People’s Hospital of Lanzhou City; The First Affiliated Hospital of Yunnan University of Chinese Medicine; Northwestern University; Linyi Traditional Chinese Medicine Hospital; Shanghai University of Traditional Chinese Medicine(密苏里大学; 上海市浦东新区社区卫生服务中心; 金沙县中医院; 上海市第七人民医院; 兰州市第一人民医院; 云南中医药大学第一附属医院; 西北大学; 临沂市中医院; 上海中医药大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建349例真实临床病例库,对比16个大语言模型与60名中医医师,发现LLM在诊断建议上评分更高,但处方存在差异和安全隐患,需医师监督。
AI 中文摘要
大语言模型(LLMs)正越来越多地被探索用于临床应用中,然而对其在真实世界传统中医(TCM)实践中的评估仍然有限。我们构建了一个包含来自62家医院的349例去标识化门诊病例的临床病例库,并从中选取60个代表性病例,对16个大语言模型和由60名执业中医医师组成的对照队列进行了评估。模型输出和医师报告均经过匿名化处理,并由五位资深中医专家在九个诊断和治疗维度上进行评分。前沿通用大语言模型获得的专家评分高于医师对照组,尤其是在医疗建议、治疗原则和部分诊断任务方面。然而,处方层面的分析揭示了在草药选择、剂量和治疗策略上的差异,定性安全审查则发现了幻觉和不良的模板驱动输出。这些发现凸显了大语言模型在中医决策支持方面的潜力,同时也强调了医师监督、安全约束和前瞻性临床评估的必要性。
英文摘要
Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.