发表机构
University of California San Francisco; Waymark(加利福尼亚大学旧金山分校; 韦马克公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对电子病历处理中的临床中间信息丢失问题,提出查询条件化临床抑制方法,经实验验证其在中间位置指令及整体任务上的表现均优于多种基线检索与重排序方法。
AI 中文摘要
如今,每位患者的电子病历(EHR)通常已超过10万个token,但大语言模型存在中间信息丢失(Lost-in-the-Middle, LitM)效应:长上下文中间位置的信息检索可靠性低于边缘位置,在临床场景中这并非无关紧要——病历中最关键的事实可能位于中间,我们将此称为临床中间信息丢失(Clinical Lost-in-the-Middle, CLitM)问题,首次使用MedAlign对其进行系统表征,并比较上下文选择策略作为解决方案。在2196条指令-响应对和6种语言模型中,我们观察到峰值准确率(59.5%,95%置信区间[46.3,71.0],对应20-30分位数)与谷值准确率(37.6%,[23.2,52.5],对应70-80分位数)之间存在21.9个百分点的差距;67.8%的参考答案位于EHR时间线的10-90分位数区间,即CLitM谷值区域内。我们引入查询条件化临床抑制(Query-Conditioned Clinical Suppression, QCCS),这是一种轻量级的查询条件化选择门,并针对BM25、带章节标题过滤的BM25、密集检索、交叉编码器重排序(N=83个保留指令)对其进行评估。在Qwen2.5-7B-Instruct(16k上下文)模型下,通过LLM-as-judge评分,QCCS在所有5种对比方法中表现更优:对于中间位置指令,QCCS准确率达16.7%,而BM25为3.3%、交叉编码器为0.0%、密集检索为0.0%、全上下文为6.7%;整体上QCCS准确率达25.3%,而仅检索的对比方法最高仅为3.6%。这一优势并非由检索召回率解释:当k=20时,BM25在98.8%的指令中检索到了正确证据句(QCCS为34.9%),但检索分支即使检索到该证据句,准确率最多仅为2.6%,而QCCS即使未检索到也能达到25.0%的准确率。在本次概念验证评估中,与正确证据句检索召回率相比,查询对齐的上下文选择能更好地预测EHR指令遵循准确率。
英文摘要
Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use this is not benign: the single most consequential fact in a note can sit at its center. We term this the clinical lost-in-the-middle (CLitM) problem, give its first systematic characterization using MedAlign, and compare context-selection strategies as remedies. Across 2,196 instruction-response pairs and six language models, we observe a 21.9 percentage-point gap between peak accuracy (59.5%, 95% CI [46.3, 71.0], 20-30% decile) and trough accuracy (37.6% [23.2, 52.5] at 70-80%); 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline, inside the CLitM trough. We introduce Query-Conditioned Clinical Suppression (QCCS), a lightweight query-conditioned selection gate, and evaluate it against BM25, BM25 with section-header filtering, dense retrieval, and cross-encoder reranking (N=83 held-out instructions). With Qwen2.5-7B-Instruct (16k context), QCCS outperforms all five comparators under LLM-as-judge scoring: for middle-position instructions QCCS reaches 16.7% versus BM25 3.3%, cross-encoder 0.0%, dense 0.0%, and full context 6.7%; overall QCCS reaches 25.3% versus at most 3.6% for retrieval-only comparators. This advantage is not explained by retrieval recall: at k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS 34.9%), yet retrieval arms stay at most 2.6% accurate even when they retrieve it, whereas QCCS reaches 25.0% even when it does not. In this proof-of-concept evaluation, query-aligned context selection predicts EHR instruction-following accuracy better than gold-sentence retrieval recall.
Comments29 pages, 5 figures. Code: https://github.com/sanjaybasu/inhibitory-attention-ehr