用于评估大型语言模型在临床登记处抽象任务中性能的歧义分类:一项多中心前瞻性研究
An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study
- Carta Healthcare(卡塔医疗)
- Stanford University School of Medicine(斯坦福大学医学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本多中心前瞻性研究评估LLM在临床登记处抽象任务中的性能,发现其准确率远低于人类抽象人员,且随问题歧义度和临床推理要求升高而降低。
AI中文摘要:
目的:评估大型语言模型(LLM)在未处理的电子病历(EMR)数据上完成临床登记处抽象任务的性能。方法:我们评估了LLM回答美国心脏病学会国家心血管数据登记处(ACC NCDR)问题的性能。在一所学术医疗中心的试点研究中,该模型为每个登记处问题识别候选数据源,经验丰富的抽象人员利用这些结果定义问题特定的文档集。在第二所中心针对另一项ACC NCDR登记处开展的验证研究中,LLM使用这些问题特定的文档集回答问题。在查看任何输出之前,两名抽象人员独立建立了真实值,并将每个问题分配到6个类别之一,这些类别按解决问题所需的歧义程度和临床推理量排序:药物/事件标记、二元临床存在、行政类、定量实验室/生理指标、临床解读、事件时间。结果:分析样本包含9430个经协调至4715个共识答案的抽象人员答案(501个来自试点,4214个来自验证)。在试点中,每个问题的候选数据源平均值在人口统计学类别的14.6个(标准差13.9)到病史和风险因素类别的89.2个(标准差56.1)之间。在验证中,人类评分者间一致性约为98%,而LLM答案中87%与共识完全匹配,2%部分匹配,9%不匹配。在157个至少有20个答案的问题中,问题级平均准确率为91.5%(标准差13.4%),且准确率随歧义程度增加而下降,从药物/事件标记类的96%降至事件时间类的62%。结论:LLM在未处理EMR数据上回答临床登记处问题的准确率远低于人类抽象人员;LLM的准确率随歧义程度和所需临床推理水平的提高而稳步下降。
英文摘要:
Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% (SD 13.4\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.