TCMClinicalReason-Bench:语言模型能否从真实临床病例的发病机制推理到处方?
TCMClinicalReason-Bench: Can Language Models Reason from Pathogenesis to Prescription over Real-World Clinical Cases?
浏览论文内容
中文总结 AI 辅助
该研究提出TCMClinicalReason-Bench基准,基于2000例真实电子病历评估LLM从病因病机到处方的中医推理能力,发现内容质量不足且跨模块一致性中等,定位了处方生成等关键缺陷。
中文摘要 AI 辅助
大型语言模型(LLM)生成的临床叙述往往缺乏以患者具体证据为基础的充分支撑。在中医(TCM)中,错误可能从病因和发病机制开始,经由证候诊断和治疗原则,一直传播到处方生成。我们利用2,000例多中心电子健康记录病例开发了TCMClinicalReason-Bench,以区分基于病例的回应与流畅但缺乏支持的诊断和治疗结论。在零样本设置下评估了五个通用型LLM和两个中医专用LLM。一个受证据约束的评分标准评估了七个诊断和治疗组成部分及三个跨模块关系,允许基于病例的替代答案。Qwen3.7-Plus结合中医检索作为自动评判者,同时由五位资深中医临床医生对600例子集进行平行盲评。结构完整性几乎饱和(99.3%-100.0%),但标准化内容得分在40.7%至54.1%之间。五个通用型模型平均得分为50.0%,而两个较小的中医专用模型平均得分为41.3%。跨模块逻辑一致性范围为60.8%至66.8%,并与跨病例的内容呈中等相关(Pearson r = 0.515-0.656)。缺陷最大之处在于处方生成、处方分析和基于症状的调整。在100个独立病例的评判者压力测试中,三个关系的扰动检测率分别为57%、56%和31%,矛盾检测比遗漏检测更可靠。将组成部分质量与跨模块一致性分开,可以定位终点和完整性指标所遗漏的失败,并识别出临床医生监督仍然必要的环节。
英文摘要
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinicalReason-Bench using 2,000 multicenter electronic health record cases to distinguish case-grounded responses from fluent but unsupported diagnostic and therapeutic conclusions. Five general-purpose and two TCM-specific LLMs were evaluated in zero-shot settings. An evidence-constrained rubric assessed seven diagnostic and therapeutic components and three cross-block relations, allowing case-supported alternatives. Qwen3.7-Plus with TCM retrieval served as the automated judge, alongside parallel blinded ratings by five senior TCM clinicians on a 600-case subset. Structural completeness was nearly saturated (99.3-100.0%), but normalized content scores ranged from 40.7% to 54.1%. The five general-purpose models averaged 50.0%, versus 41.3% for the two smaller TCM-specific models. Cross-block logic consistency ranged from 60.8% to 66.8% and correlated moderately with content across cases (Pearson's r = 0.515-0.656). Deficits were greatest in prescription generation, prescription analysis, and symptom-guided modification. In judge stress testing on 100 independent cases, perturbation detection rates across the three relations were 57%, 56%, and 31%, with contradictions detected more reliably than omissions. Separating component quality from cross-block consistency localizes failures missed by endpoint and completeness metrics and identifies where clinician oversight remains necessary.
发表机构
- Nanjing University of Chinese Medicine(南京中医药大学)
- Beijing University of Chinese Medicine(北京中医药大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Huzhou University(湖州大学)
机构由 AI 辅助整理,请以论文原文为准。