arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLMs 锚定主诉,未能在顺序临床分诊中整合证据

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan

arXiv 2609.22904首次发表:更新:

发表机构

University of New South Wales; Campbelltown Hospital(新南威尔士大学; 坎贝尔敦医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出顺序分诊评估方法,发现LLMs在急诊分诊中锚定主诉而忽略后续证据,性能从完整记录的中度一致降至顺序检查点的公平至中度一致,远低于临床专家,提示离线基准不足以支持临床部署。

AI 中文摘要

急诊科(ED)的分诊是一个逐轮展开的顺序决策过程。现有对大型语言模型(LLMs)分诊能力的评估使用完整的回顾性记录,并报告其性能接近医生水平。我们提出了一种在顺序分诊中评估LLMs的方法,该任务要求从护士与患者对话的不断增长的前缀中预测分诊 acuity 标签。我们在两个语料库的五个顺序检查点上评估了六个LLMs:425个LLM生成的(SIMULATED)对话和50个医生撰写的(CLINICIAN)对话,两者均根据急诊严重程度指数(ESI)进行标注。每个模型,通过二次加权kappa(QWK)衡量,在完整记录上从中度至实质性一致性,在每个顺序检查点下降到公平至中度一致性。受控扰动实验表明,每个检查点的标签都锚定在主诉交流上,而提示干预未能突破这一平台期。模型从后续轮次中提取了临床相关内容,但真实标签的意外性(surprisal)在各个检查点上升,因此模型未能整合证据。三位临床专家在同一对话上的QWK达到0.887-0.929,而最佳模型仅达到0.295。预测集中在ESI-2和ESI-3,且模型之间的一致性高于与真实标签的一致性,因此集成反而加剧了失败。仅基于离线基准部署LLMs进行急诊分诊会忽略这一顺序失败。

英文摘要

Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑