发表机构
NYU Langone Health; Johns Hopkins University School of Medicine; Washington University School of Medicine in St. Louis; University of Alabama at Birmingham Heersink School of Medicine; Perelman School of Medicine at the University of Pennsylvania; New York University(纽约大学朗格尼健康中心; 约翰霍普金斯大学医学院; 圣路易斯华盛顿大学医学院; 阿拉巴马大学伯明翰分校希尔斯金医学院; 宾夕法尼亚大学佩雷尔曼医学院; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示大型语言模型在临床病历生成中易受附带信息污染,导致闲聊插入和错误归属,并提出双编码假说,强调临床使用前需评估抗干扰能力。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于支持环境文档记录和临床推理。在此,我们通过评估模型对患者就诊过程中附带信息的敏感性,考察了这两种应用共有的失败模式的影响。在576段患者-临床医生对话中,我们发现前沿模型在35%的病历中插入了闲聊交流,而平均质量评分在五分制量表上最多变化0.20分。在3.7%的前沿模型病历中,模型错误归属了这些题外话或将其用于临床。在57次模拟录音咨询中,来自另一次患者就诊的背景语音在-10分贝下泄漏到48.2%的转录文本中,在四个开放权重模型生成的下游病历中检测到5.3%的污染。我们提出了LLM中临床推理和分心的双编码假说,初步证据表明,与附带信息干扰相关的LLM组件也支持临床推理。这些发现支持在临床使用前评估对附带信息的抵抗力,并采取防止污染同时保留临床推理的保障措施。
英文摘要
Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.