发表机构
Case School of Engineering, Case Western Reserve University; National Library of Medicine, National Institutes of Health; Rice University(凯斯西储大学凯斯工程学院; 美国国立卫生研究院国家医学图书馆; 莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床语言模型评估中因未来数据暴露导致的后见之明偏差,构建配对基准,通过时间掩蔽减少偏差且不损准确率。
AI 中文摘要
临床决策是前瞻性的,但临床语言模型通常在回顾性记录上进行评估,这些记录揭示了最终诊断、治疗反应和结果。此类评估可能奖励使用未来信息,而非在决策点存在的不确定性下进行推理。我们引入了一个配对基准,用于测量与临床时间推理中后见之明偏差一致的结果条件性偏移。该基准包含来自PubMed Central开放获取子集的171份病例报告——40份脓毒症病例和131份GLP-1/糖尿病病例——以文本叙述以及人工标注和LLM生成的文本时间序列(TTS)两种形式呈现。对于每个病例,问题与临床上有意义的截断点相关联,并配对一个前瞻性参考答案和一个与结果一致的“后见之明陷阱”。模型使用在截断点截断的TTS或完整时间线回答每个问题;附加条件改变叙述来源(原始或合成)和TTS标注来源(人工或LLM)。我们评估准确率(Acc)、后见之明陷阱率(HTR)、答案不稳定率(AIR)和后见之明偏差率(HBR),每个指标捕捉后见之明偏差的不同信号。在GPT 5.6 Sol、Gemma 4、GLM 5.2和Opus 5上,完整时间线暴露产生一致的后见之明敏感偏移,而时间掩蔽在不降低准确率的情况下减少了偏差。
英文摘要
Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduce a paired benchmark for measuring outcome-conditioned shifts consistent with hindsight bias in clinical temporal reasoning. It contains 171 case reports from the PubMed Central Open Access Subset---40 sepsis and 131 GLP-1/diabetes cases---represented as both textual narratives and human-annotated and LLM-generated textual time series (TTS). For each case, questions are tied to a clinically meaningful cutoff and paired with a prospective reference answer and an outcome-consistent \emph{hindsight trap}. Models answer each question using either a TTS truncated at the cutoff or the complete timeline; additional conditions vary the narrative source (original or synthetic) and TTS annotation source (human or LLM). We evaluate accuracy (Acc), hindsight trap rate (HTR), answer instability rate (AIR), and hindsight bias rate (HBR), each of which captures different signals of hindsight bias. Across GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5, full timeline exposure produces consistent hindsight-sensitive shifts, while temporal masking reduces bias without lowering accuracy.
CommentsMachine Learning for Health Symposium (ML4H 2026)