arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ObGynLongBench:揭示纵向电子健康记录决策中的证据到电子健康记录差距

ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

Jun Xiang, Zhijie Bao, Rong Hu, Kaizhou Qin, Wei Chen, Zhongyu Wei

arXiv 2609.07601首次发表:更新:

发表机构

School of Data Science, Fudan University; Obstetrics & Gynecology Hospital of Fudan University; Shanghai Innovation Institute; School of Software Engineering, Huazhong University of Science and Technology(复旦大学数据科学学院; 复旦大学附属妇产科医院; 上海创新研究院; 华中科技大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ObGynLongBench是一个基于规则的产科和妇科长期上下文EHR基准,包含1,500个来自真实妊娠病史的决策案例,评估17个LLMs发现证据到EHR差距,并表明主动搜索智能体表现最佳,突显证据利用是关键挑战。

AI 中文摘要

大型语言模型(LLMs)在个性化医疗助手中的应用日益受到关注。然而,现有医学基准大多依赖于带有预选证据的静态问答,尚不清楚LLMs能否从真实的纵向电子健康记录(EHRs)中做出可靠的临床决策。为弥合这一差距,我们引入了ObGynLongBench,一个基于规则的、面向产科和妇科决策的长期上下文EHR基准,包含来自976条真实妊娠EHR病史的1,500个临床决策点案例及可追溯规则。每个案例锚定于一位患者、一个妊娠时间线节点和一个决策前信息边界,从而支持仅证据、就诊级EHR和病史级EHR三种评估模式。对17个LLMs的评估揭示了一个显著的证据到EHR差距:当直接提供证据时模型表现良好,但当证据必须从同日记录或完整的决策前EHR病史中提取时,准确率下降。进一步分析确定证据利用是关键瓶颈:随着EHR上下文变长和证据需求复杂化,性能下降,且同一患者病史中较早的失败往往预示较晚的失败。最后,主动搜索智能体在EHR访问策略中表现最佳,突显了患者特定证据利用作为可靠个性化医疗助手核心挑战的重要性。资源可在该https URL获取。

英文摘要

The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic health records (EHRs). To bridge this gap, we introduce ObGynLongBench, a rule-grounded long-context EHR benchmark for obstetric and gynecologic decision-making, comprising 1,500 clinical decision-point cases from 976 real pregnancy EHR histories and traceable rules. Each case is anchored to a patient, a pregnancy-timeline point, and a pre-decision information boundary, enabling Evidence-only, Visit-level EHR, and History-level EHR evaluation. Evaluating 17 LLMs reveals a substantial Evidence-to-EHR Gap: models perform well when evidence is directly provided, but accuracy drops when evidence must be extracted from same-day records or full pre-decision EHR histories. Further analyses identify evidence utilization as a key bottleneck: performance decreases with longer EHR contexts and more complex evidence requirements, and earlier failures often predict later failures within the same patient history. Finally, active-search agents perform best among EHR access strategies, highlighting patient-specific evidence utilization as a central challenge for reliable personalized medical assistants. Resources are available at https://github.com/xiangjun2003/ObgynLongbench.

Comments25 pages, including references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑