arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在患者证据演变过程中表现出不可靠的临床判断更新

Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

Min Zeng, Rui Zhang

arXiv 2610.02684首次发表:更新:

发表机构

University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过重症监护轨迹评估发现,LLMs在临床判断更新中不可靠,存在对恶化证据过度反应及先前信念偏差,提示无法通过提示修复,并引入EVLU指标揭示可靠性-覆盖率权衡。

AI 中文摘要

大型语言模型(LLMs)越来越多地被探索用于临床推理,但它们是否能在患者证据演变时适当地修正判断仍不清楚。我们利用电子健康记录中的匹配重症监护轨迹评估了纵向信念更新。在多种LLMs中,当估计值发生变化时,以前一判断为条件相比减少预测误差更常增加预测误差,这一现象在第二个终点上得到重复验证。受控干预揭示了两种失败模式。首先,在前一评估固定时,模型对恶化的呼吸证据的反应比对匹配的改善证据更强烈;这种不对称性在中度和强证据水平下经净空归一化后仍然存在。其次,在当前证据固定时,将先前风险从10%增加到90%使估计值偏移了26.2个百分点,表明先前模型信念的因果影响。提示未能恢复可靠的更新。证据验证的纵向更新(EVLU)识别出更少但更可靠的修订,揭示了可靠性-覆盖率的权衡。这些发现确立了纵向信念更新作为LLM可靠性一个独立维度。

英文摘要

Large language models (LLMs) are increasingly explored for clinical reasoning, but whether they appropriately revise judgments as patient evidence evolves remains unclear. We evaluated longitudinal belief updating using matched intensive-care trajectories from electronic health records. Across diverse LLMs, conditioning on a preceding judgment more often increased than reduced prediction error when estimates changed, replicated for a second endpoint. Controlled interventions revealed two failure modes. First, with preceding assessment fixed, models responded more strongly to worsening than matched improving respiratory evidence; this asymmetry persisted after headroom normalization at moderate and strong evidence levels. Second, with current evidence fixed, increasing prior risk from 10% to 90% shifted estimates by 26.2 percentage points, demonstrating causal influence of prior model beliefs. Prompting did not restore reliable updating. Evidence-Validated Longitudinal Update (EVLU) identified fewer, more reliable revisions, revealing a reliability-coverage trade-off. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑